Image processing method and apparatus, and electronic device and storage medium

By detecting and segmenting the face images in naked-eye 3D display technology, and fusion coordinates determine the position of the human eye, the problem of insufficient human eye tracking operation speed and accuracy in the prior art is solved, and the effect of naked-eye 3D display is improved.

WO2025129551A1PCT designated stage expired Publication Date: 2025-06-26BOE TECHNOLOGY GROUP CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2023/140539
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-21
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

In the existing naked-eye 3D display technology, the computing speed and accuracy of human eye tracking need to be improved to ensure viewing effect.

Method used

By obtaining face image frames, key point detection and segmentation are performed, key point coordinates are integrated to determine the position of the human eye, and the accuracy and computing speed of key point positioning are improved.

Benefits of technology

It achieves higher key point positioning accuracy and computing speed, improving the effect of naked-eye 3D display.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023140539_26062025_PF_FP_ABST
    Figure CN2023140539_26062025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are an image processing method and apparatus, and an electronic device and a storage medium. The method comprises: acquiring a facial image in a first image frame; performing key point detection on the facial image to obtain first coordinates of key points; performing facial image segmentation on the facial image to obtain local facial images; performing key point detection on a first local image to obtain second coordinates of a first key point, and performing key point detection on a second local image to obtain second coordinates of a second key point; fusing the first coordinates of the first key point and the second coordinates of the first key point to obtain third coordinates of the first key point, and fusing the first coordinates of the second key point and the second coordinates of the second key point to obtain third coordinates of the second key point; and on the basis of the third coordinates of the first key point, the third coordinates of the second key point and a preset facial image, obtaining eye position coordinates. Results of coarse localization and fine localization are fused, so that the operation speed and accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing method, device, electronic device and storage medium Technical Field

[0001] At least one embodiment of the present disclosure relates to an image processing method, apparatus, electronic device, and storage medium. Background Art

[0002] Glasses-free 3D display, also known as autostereoscopic display, utilizes the parallax of the human eye, allowing viewers to see realistic three-dimensional images with space and depth without the need for auxiliary equipment (such as 3D glasses or 3D helmets). To ensure this viewing experience, the speed and accuracy of eye tracking algorithms need to be improved.

[0003] Summary of the Invention

[0004] At least one embodiment of the present disclosure provides an image processing method, apparatus, electronic device, and storage medium.

[0005] At least one embodiment of the present disclosure provides an image processing method, including: acquiring a facial image in a first image frame; performing key point detection on the facial image to obtain first coordinates of the key points; the first coordinates of the key points include the first coordinates of the first key points and the first coordinates of the second key points; performing facial image segmentation on the facial image to obtain partial facial images; the partial facial images include the first partial image and the second partial image; performing key point detection on the first partial image to obtain the second coordinates of the first key points; performing key point detection on the second partial image to obtain the second coordinates of the second key points; fusing the first coordinates of the first key points and the second coordinates of the first key points to obtain the third coordinates of the first key points; fusing the first coordinates of the second key points and the second coordinates of the second key points to obtain the third coordinates of the second key points; and obtaining eye position coordinates based on the third coordinates of the first key points, the third coordinates of the second key points and a preset facial image.

[0006] For example, according to at least one embodiment of the present disclosure, the facial image is segmented to obtain partial facial images, including: segmenting the facial image according to the area where the eyes are located and the area where the mouth and nose are located to obtain the first partial image and the second partial image; the first partial image includes an eye area image, and the second partial image includes a mouth and nose area image.

[0007] For example, according to at least one embodiment of the present disclosure, the area where the eyes are located is determined based on the first coordinates of the first key point, and the area where the mouth and nose are located is determined based on the first coordinates of the second key point.

[0008] For example, according to at least one embodiment of the present disclosure, the first key points include eye key points, and the second key points include mouth and nose key points.

[0009] For example, according to at least one embodiment of the present disclosure, the first coordinate of the first key point and the second coordinate of the first key point are fused to obtain the third coordinate of the first key point, including: fusing the first coordinate of the eye key point and the second coordinate of the eye key point to obtain the third coordinate of the eye key point; fusing the first coordinate of the second key point and the second coordinate of the second key point to obtain the third coordinate of the second key point, including: fusing the first coordinate of the mouth and nose key point and the second coordinate of the mouth and nose key point to obtain the third coordinate of the mouth and nose key point.

[0010] For example, according to at least one embodiment of the present disclosure, the eye key points include the left eye center key point and the right eye center key point; the mouth and nose key points include the left mouth corner key point, the right mouth corner key point and the nose tip key point.

[0011] For example, according to at least one embodiment of the present disclosure, before performing key point detection on the facial image to obtain the first coordinates of the key points, it also includes: performing feature extraction on the facial image through a neural network model to obtain a first feature map; downsampling the first feature map to obtain a second feature map; performing key point detection on the facial image to obtain the first coordinates of the key points, including: performing the key point detection on the second feature map to obtain the first coordinates of the key points.

[0012] For example, according to at least one embodiment of the present disclosure, performing facial image segmentation on the facial image to obtain a partial facial image includes: performing facial image segmentation on the first feature map to obtain the partial facial image.

[0013] For example, according to at least one embodiment of the present disclosure, the neural network model is a quantized neural network model.

[0014] For example, according to at least one embodiment of the present disclosure, the method further includes: quantizing the initial neural network model to obtain a current neural network model; determining whether the current neural network model satisfies the operation restriction condition; in response to the current neural network model satisfying the operation restriction condition, using the current neural network model as the quantized neural network model; in response to the current neural network model not satisfying the operation restriction condition, obtaining the output error of each processing layer in the current neural network model, and correcting the processing layer with the largest output error until the current neural network model satisfies the operation restriction condition.

[0015] For example, according to at least one embodiment of the present disclosure, the operation restriction condition includes a restriction condition on at least one of operation speed and model accuracy.

[0016] For example, according to at least one embodiment of the present disclosure, the method further includes: acquiring a second image frame, wherein the first image frame and the second image frame are adjacent in the time domain, and the first image frame is a forward frame of the second image frame; determining whether the score values ​​corresponding to the third coordinate of the first key point and the third coordinate of the second key point obtained based on the first image frame both meet preset requirements; in response to the score value of the third coordinate of the first key point and the score value of the third coordinate of the second key point both meeting the preset requirements, acquiring a facial image in the second image frame based on the third coordinate of the first key point and the third coordinate of the second key point obtained based on the first image frame.

[0017] For example, according to at least one embodiment of the present disclosure, obtaining a facial image in the second image frame based on the third coordinate of the first key point and the third coordinate of the second key point obtained in the first image frame includes: obtaining an affine transformation matrix based on the third coordinate of the first key point and the third coordinate of the second key point; and obtaining the facial image in the second image frame based on the affine transformation matrix.

[0018] For example, according to at least one embodiment of the present disclosure, the preset requirement includes that the score value of the third coordinate of the first key point and the score value of the third coordinate of the second key point are both greater than a preset score threshold.

[0019] For example, according to at least one embodiment of the present disclosure, the method further includes preprocessing the first image frame, and the preprocessing includes: performing region detection on the first image frame to obtain a face region; performing face extraction on the face region to obtain the face image.

[0020] For example, according to at least one embodiment of the present disclosure, a first thread and a second thread are executed in parallel, the first thread includes an image acquisition thread, and the second thread includes an image analysis thread; the image acquisition thread includes acquiring the first image frame and storing it in a cache space; the image analysis thread includes at least one of the following operations: extracting the first image frame in the cache space, obtaining the face image in the first image frame, performing key point detection on the face image, performing face image segmentation on the face image, performing key point detection on the first local image, performing key point detection on the second local image, fusing the first coordinate of the first key point and the second coordinate of the first key point, fusing the first coordinate of the second key point and the second coordinate of the second key point, and obtaining the eye position coordinates based on the third coordinate of the first key point, the third coordinate of the second key point and the preset face image.

[0021] For example, according to at least one embodiment of the present disclosure, the method further includes: generating a synchronization signal based on a pixel rearrangement operation; wherein the second thread includes: extracting the first image frame in the cache space in response to the synchronization signal.

[0022] For example, according to at least one embodiment of the present disclosure, the eye position coordinates are obtained based on the third coordinate of the first key point, the third coordinate of the second key point and a preset facial image, including: obtaining the rotation angle of the facial orientation in the facial image relative to the facial orientation of the preset facial image based on the third coordinate of the first key point, the third coordinate of the second key point and the coordinates of the facial key points of the preset facial image; obtaining the pupil distance of the human eye; and obtaining the eye position coordinates based on the rotation angle and the pupil distance.

[0023] For example, according to at least one embodiment of the present disclosure, obtaining the eye position coordinates based on the rotation angle and the pupil distance includes: obtaining the component of the eye position coordinates on the z-axis using the following formula: Among them, X c [2] represents the coordinates of the eye position, d represents the pupil distance, [0] represents the component on the x-axis, [1] represents the component on the y-axis, [2] represents the component on the z-axis, and f x 、f y , x0, y0 are preset constants, C y =cosθ y , C z =cosθ z , S y = sinθ y , where θ y and θ zare the components of the rotation angle on the y-axis and the z-axis, respectively, Δx=x r -x l , The first key point includes the eye key point, x r Indicates the third coordinate of the left eye center key point in the eye key points, x l Indicates the third coordinate of the right eye center key point among the eye key points.

[0024] At least one embodiment of the present disclosure provides an image processing device, comprising: an acquisition module configured to acquire a facial image in a first image frame; a first detection module configured to perform key point detection on the facial image to obtain first coordinates of the key points; the first coordinates of the key points include the first coordinates of the first key point and the first coordinates of the second key point; a segmentation module configured to perform facial image segmentation on the facial image to obtain partial facial images; the partial facial images include the first partial image and the second partial image; a second detection module configured to perform key point detection on the first partial image to obtain the second coordinates of the first key point, and to perform key point detection on the second partial image to obtain the second coordinates of the second key point; a fusion module configured to fuse the first coordinate of the first key point and the second coordinate of the first key point to obtain the third coordinate of the first key point, and to fuse the first coordinate of the second key point and the second coordinate of the second key point to obtain the third coordinate of the second key point; and a calculation module configured to obtain eye position coordinates based on the third coordinate of the first key point, the third coordinate of the second key point, and a preset facial image.

[0025] At least one embodiment of the present disclosure provides an electronic device, comprising: a shooting device configured to shoot a first image frame; an image processing device configured to receive the first image frame and execute any of the aforementioned image processing methods based on the first image frame.

[0026] At least one embodiment of the present disclosure provides an electronic device, comprising: a memory, which non-transitorily stores computer-executable instructions; and a processor, configured to execute the computer-executable instructions, wherein the computer-executable instructions implement the image processing method according to any of the aforementioned items when executed by the processor.

[0027] At least one embodiment of the present disclosure provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, the image processing method according to any of the above items is implemented.

[0028] According to the image processing method, device, electronic device and storage medium described in the embodiments of the present disclosure, key point detection is performed on the full face image of the face image to obtain the first coordinate of the first key point and the first coordinate of the second key point, which can obtain rough key point coordinates with a lower amount of calculation. Key point detection is performed on the segmented partial face images respectively to obtain the second coordinate of the first key point in the first partial image and the second coordinate of the second key point in the second partial image, which can obtain key point coordinates with higher accuracy. The third coordinates of the first key point and the third coordinates of the second key point are obtained by fusion, which can not only improve the calculation speed, but also improve the accuracy of key point positioning. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly introduced below. Obviously, the drawings described below only relate to some embodiments of the present disclosure, but do not limit the present disclosure.

[0030] FIG1 shows a schematic diagram of an image processing method.

[0031] FIG2 is a flowchart of an image processing method provided by at least one embodiment of the present disclosure.

[0032] FIG3 is a process diagram of an image processing method provided by at least one embodiment of the present disclosure.

[0033] FIG4 is a flowchart of an image processing method provided by at least one embodiment of the present disclosure.

[0034] FIG5 is a schematic diagram of key points provided by at least one embodiment of the present disclosure.

[0035] FIG6 is a schematic diagram of the rotation angle of a facial posture provided by at least one embodiment of the present disclosure.

[0036] FIG7 is a schematic diagram of a coordinate system provided by at least one embodiment of the present disclosure.

[0037] FIG8 is a flowchart of an image processing method provided by at least one embodiment of the present disclosure.

[0038] FIG9 is a flowchart of an image processing method provided by at least one embodiment of the present disclosure.

[0039] FIG10 is a flowchart of an image processing method provided by at least one embodiment of the present disclosure.

[0040] FIG11 is a flow chart of an image processing method provided by at least one embodiment of the present disclosure.

[0041] FIG12 is a schematic diagram of region detection provided by at least one embodiment of the present disclosure.

[0042] FIG13 is a flowchart of an image processing method provided by at least one embodiment of the present disclosure.

[0043] FIG14 shows a schematic diagram of an image processing method.

[0044] FIG15 is a schematic diagram of an image processing method provided by at least one embodiment of the present disclosure.

[0045] FIG16 is a schematic diagram of an image processing method provided by at least one embodiment of the present disclosure.

[0046] FIG17 is a schematic block diagram of an image processing apparatus provided by at least one embodiment of the present disclosure.

[0047] FIG18 is a schematic block diagram of an electronic device provided by at least one embodiment of the present disclosure.

[0048] FIG19 is a schematic block diagram of another electronic device provided by at least one embodiment of the present disclosure.

[0049] FIG20 is a schematic diagram of a non-transitory computer-readable storage medium provided by at least one embodiment of the present disclosure.

[0050] FIG21 is a schematic diagram of a hardware environment provided by at least one embodiment of the present disclosure. DETAILED DESCRIPTION

[0051] To make the purpose, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the described embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.

[0052] Unless otherwise defined, technical or scientific terms used in this disclosure should have the ordinary meanings understood by people with ordinary skills in the field to which this disclosure belongs. The words "first", "second" and similar terms used in this disclosure do not indicate any order, quantity or importance, but are simply used to distinguish different components. The words "include" or "comprising" and similar terms mean that the elements or objects preceding the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects.

[0053] Glasses-free 3D technology allows viewing stereoscopic images without the need for 3D glasses. Eye tracking is a key component of glasses-free 3D display systems. Eye tracking accurately tracks the viewer's eye position and adjusts the displayed content accordingly, maintaining a good 3D effect even when viewed from different angles. Eye tracking can be achieved using infrared sensors, cameras, or other sensor devices that monitor the viewer's eye position and movement in real time, then feed this information back to the display device, allowing it to adjust the displayed content to suit the viewer's viewing angle.

[0054] During their research, the inventors of this application discovered that achieving better naked-eye 3D display effects requires ensuring the accuracy of eye position coordinates tracked by eye tracking technology, and the accuracy of facial landmark location results has a direct impact on the accuracy of eye tracking. Furthermore, naked-eye 3D display systems place high demands on the speed, accuracy, stability, and synchronization of eye tracking and eye position measurement on terminal devices. Only when these multiple aspects reach high levels can a good naked-eye 3D display effect be achieved.

[0055] FIG1 shows a schematic diagram of an image processing method.

[0056] Referring to Figure 1, in order to obtain more accurate eye position coordinates, a convolutional neural network (CNN) can be used to extract features from a facial image 10 to obtain a feature map 20. A multilayer perceptron (MLP) is then used to regress and predict the key point coordinates of the feature map 20 to achieve detection and positioning of key points on the entire face. For example, the size of the facial image 10 can be 256*256*3, and the size of the feature map 20 can be 64*64*32. However, the method of directly performing MLP full-face key point regression positioning after feature extraction can obtain results that meet accuracy requirements, but the computational complexity is large and the calculation speed is slow.

[0057] At least one embodiment of the present disclosure provides an image processing method, including: acquiring a facial image in a first image frame; performing key point detection on the facial image to obtain first coordinates of the key points; the first coordinates of the key points include the first coordinates of the first key points and the first coordinates of the second key points; performing facial image segmentation on the facial image to obtain partial facial images; the partial facial images include the first partial image and the second partial image; performing key point detection on the first partial image to obtain the second coordinates of the first key points; performing key point detection on the second partial image to obtain the second coordinates of the second key points; fusing the first coordinates of the first key points and the second coordinates of the first key points to obtain the third coordinates of the first key points; fusing the first coordinates of the second key points and the second coordinates of the second key points to obtain the third coordinates of the second key points; and obtaining eye position coordinates based on the third coordinates of the first key points, the third coordinates of the second key points and a preset facial image.

[0058] According to the image processing method, device, electronic device and storage medium described in the embodiments of the present disclosure, key point detection is performed on the full face image of the face image to obtain the first coordinate of the first key point and the first coordinate of the second key point, which can obtain rough key point coordinates with a lower amount of calculation. Key point detection is performed on the segmented partial face images respectively to obtain the second coordinate of the first key point in the first partial image and the second coordinate of the second key point in the second partial image, which can obtain key point coordinates with higher accuracy. The third coordinates of the first key point and the third coordinates of the second key point are obtained by fusion, which can not only improve the calculation speed, but also improve the accuracy of key point positioning.

[0059] The image processing method provided in the embodiments of the present disclosure can be applied to the image processing device provided in the embodiments of the present disclosure, which can be configured on an electronic device. The electronic device can be a personal computer, a mobile terminal, etc. The mobile terminal can be a hardware device such as a mobile phone, a tablet computer, a laptop computer, or XR glasses.

[0060] The image processing method, device, electronic device and storage medium are described below with reference to the accompanying drawings and through some embodiments.

[0061] Figure 2 is a flow chart of an image processing method provided by at least one embodiment of the present disclosure. Figure 3 is a process chart of an image processing method provided by at least one embodiment of the present disclosure.

[0062] As shown in FIG. 2 and FIG. 3 , the image processing method provided by at least one embodiment of the present disclosure includes the following steps S110 to S160 .

[0063] Step S110 : Acquire a face image in the first image frame 110 .

[0064] Step S120: Detect key points on the face image to obtain first coordinates of the key points. The first coordinates of the key points include the first coordinates of the first key point and the first coordinates of the second key point.

[0065] Step S130 : performing facial image segmentation on the facial image to obtain partial facial images; the partial facial images include a first partial image 121 and a second partial image 122 .

[0066] Step S140: performing key point detection on the first partial image 121 to obtain the second coordinates of the first key point; performing key point detection on the second partial image 122 to obtain the second coordinates of the second key point.

[0067] Step S150: Fusing the first coordinate of the first key point and the second coordinate of the first key point to obtain the third coordinate of the first key point; fusing the first coordinate of the second key point and the second coordinate of the second key point to obtain the third coordinate of the second key point.

[0068] Step S160: obtaining eye position coordinates based on the third coordinate of the first key point, the third coordinate of the second key point, and the preset face image.

[0069] As shown in Figures 2 and 3, the image processing method provided by the embodiment of the present disclosure performs key point detection on the full face image of the face image to obtain the first coordinates of the first key point and the first coordinates of the second key point, which can obtain rough key point coordinates with a lower amount of computation. Key point detection is performed on the segmented partial face images respectively to obtain the second coordinates of the first key point in the first partial image 121 and the second coordinates of the second key point in the second partial image 122, which can obtain key point coordinates with higher precision. The third coordinates of the first key point and the third coordinates of the second key point are obtained by fusion, which can not only improve the computing speed but also improve the accuracy of key point positioning. Compared with the method shown in Figure 1 that directly performs MLP full-face key point regression positioning after feature extraction, the image processing method provided by the embodiment of the present disclosure can reduce the computational complexity by approximately 50% while ensuring the accuracy of the human eye position coordinates, thereby meeting the requirements of real-time operation on low-computing power terminal devices.

[0070] As shown in FIG. 2 and FIG. 3 , for example, in step S110 , the first image frame 110 may include a photo of a human face captured by a camera of a smartphone, a camera of a tablet computer, a lens of a digital camera, or the like.

[0071] As shown in Figures 2 and 3, for example, in step S120, the key points can be some key points with strong characterization capabilities on the face, such as key points such as the eyes, eye corners, eyebrows, the highest point of the cheekbones, nose, mouth, chin, and the outer contour of the face. For example, the key points can be points in the area of ​​the face where the facial organs (eyes, nose, and mouth) are located. For example, the first coordinates of the key points are the first coordinates obtained by roughly locating the entire face image. For example, the first key points include the eye key points, and the second key points include the mouth and nose key points. For example, the first coordinates of the first key point can be the first coordinates obtained by roughly locating the eye key points. For example, the first coordinates of the second key point can be the first coordinates obtained by roughly locating the mouth and nose key points. For example, the eye key points can include the left eye center key point and the right eye center key point. For example, the mouth and nose key points can include the left corner of the mouth key point, the right corner of the mouth key point, and the nose tip key point.

[0072] As shown in Figures 2 and 3, for example, in step S130, the facial image can be segmented according to the areas where different facial organs are located on the face, to obtain partial facial images including different facial organs. For example, the area surrounded by the outer contour of the partial facial image and the area where the facial organs are located do not necessarily completely overlap. The area surrounded by the outer contour of the partial facial image only needs to roughly surround the corresponding facial organs. For example, the first partial image 121 can be a partial image including the eyes, and the second partial image 122 can be a partial image including the nose and mouth. For example, there may or may not be overlapping portions between the first partial image 121 and the second partial image 122.

[0073] As shown in Figures 2 and 3, in some examples, performing facial image segmentation on a facial image to obtain partial facial images, that is, step S130, may include: segmenting the facial image according to the area where the eyes are located and the area where the mouth and nose are located to obtain a first partial image 121 and a second partial image 122. The first partial image 121 includes an eye area image, and the second partial image 122 includes an mouth and nose area image. For example, the eye area may be an area including the left eye and the right eye. For example, the first partial image may include a left eye image and a right eye image. For example, the first partial image may be an image including both the left eye and the right eye. For example, the mouth and nose area may be an area including the mouth and the nose. For example, the second partial image may include a mouth image and a nose image. For example, the second partial image may be an image including both the mouth and the nose. For example, the second partial image may include a complete nose and a complete mouth. For example, the second partial image may include the tip of the nose and the complete mouth.

[0074] As shown in Figures 2 and 3, in some examples, the eye region is determined based on the first coordinate of the first key point, and the mouth and nose region is determined based on the first coordinate of the second key point. By performing key point detection on a full-face image, the eye region can be determined based on the first coordinate of the first key point, and the mouth and nose region can be determined based on the first coordinate of the second key point. Thus, the first coordinates of the first key point and the first coordinates of the second key point can provide a reference for segmenting the facial image, thereby improving the accuracy of region division.

[0075] As shown in Figures 2 and 3, for example, in step S140, the first partial image 121 includes an eye. By performing key point detection on the first partial image 121, the second coordinates of the first key point can be obtained. For example, the second coordinates of the first key point can be obtained after fine-tuning the positioning of the eye key point. For example, the second partial image 122 includes a mouth and nose. By performing key point detection on the second partial image 122, the second coordinates of the second key point can be obtained. For example, the second coordinates of the second key point can be obtained after fine-tuning the positioning of the mouth and nose key points.

[0076] As shown in Figures 2 and 3, for example, in step S150, by fusing the first coordinate of the first key point obtained by coarse positioning with the second coordinate of the first key point obtained by fine positioning, a more accurate third coordinate of the first key point can be obtained. By fusing the first coordinate of the second key point obtained by coarse positioning with the second coordinate of the second key point obtained by fine positioning, a more accurate third coordinate of the second key point can be obtained. For example, more accurate third coordinates of the eye key point and the mouth and nose key point can be obtained.

[0077] As shown in Figures 2 and 3, in some examples, the first coordinate of the first key point and the second coordinate of the first key point are fused to obtain the third coordinate of the first key point, that is, step S150, including fusing the first coordinate of the eye key point and the second coordinate of the eye key point to obtain the third coordinate of the eye key point. By fusing the first coordinate of the eye key point obtained by coarse positioning and the second coordinate of the eye key point obtained by fine positioning, the third coordinate of the eye key point can be obtained with higher accuracy.

[0078] As shown in Figures 2 and 3, for example, the third coordinates of the eye key points may include the coordinates of the 16 contour points around the left eye and the coordinates of the 16 contour points around the right eye. For example, the third coordinates of the eye key points may include the third coordinates of the left eye center key point and the third coordinates of the right eye center key point. For example, the third coordinates of the left eye center key point may be obtained by taking a weighted average of the coordinates of the 16 contour points around the left eye. For example, the third coordinates of the right eye center key point may be obtained by taking a weighted average of the coordinates of the 16 contour points around the right eye.

[0079] As shown in Figures 2 and 3, in some examples, the first coordinate of the second key point and the second coordinate of the second key point are fused to obtain the third coordinate of the second key point, that is, step S150, including: fusing the first coordinate of the mouth and nose key point and the second coordinate of the mouth and nose key point to obtain the third coordinate of the mouth and nose key point. By fusing the first coordinate of the mouth and nose key point obtained by coarse positioning and the second coordinate of the mouth and nose key point obtained by fine positioning, the third coordinate of the mouth and nose key point with higher accuracy can be obtained. For example, the third coordinate of the mouth and nose key point includes the third coordinate of the mouth key point and the third coordinate of the nose key point. For example, the third coordinate of the mouth key point includes the third coordinate of the left corner of the mouth key point and the third coordinate of the right corner of the mouth key point. For example, the third coordinate of the nose key point includes the third coordinate of the nose tip key point.

[0080] As shown in Figures 2 and 3, for example, in step S160, the preset facial image is a standard facial image. For example, the standard face is a pre-collected facial image. For example, the standard face may include a male face and a female face, and the standard face is saved in a standard face database for easy query and retrieval. The information of the standard face represents the relative position relationship of the facial features of the standard face, and can be saved in the standard face database. For example, before executing the image processing method, the above-mentioned standard face database is pre-established. Each time the image processing method is executed, the standard face database can be queried to obtain at least one required standard face and its information. Based on the third coordinate of the first key point and the third coordinate of the second key point obtained in step S150, combined with the preset facial image for calculation, it is possible to improve the accuracy of the human eye position coordinates obtained by positioning while reducing the amount of calculation and increasing the operation speed.

[0081] FIG4 is a flowchart of an image processing method provided by at least one embodiment of the present disclosure.

[0082] As shown in Figures 2 to 4, in some examples, the eye position coordinates are obtained based on the third coordinate of the first key point, the third coordinate of the second key point and the preset face image, that is, step S160, which may include steps S161 to S163.

[0083] Step S161: obtaining a rotation angle of a face orientation in the face image relative to a face orientation in the preset face image based on the third coordinate of the first key point, the third coordinate of the second key point, and the coordinates of the face key points in the preset face image;

[0084] Step S162: obtaining the pupil distance of the human eye;

[0085] Step S163: Obtain the eye position coordinates based on the rotation angle and pupil distance.

[0086] As shown in Figures 2 to 4, for example, in step S161, the face posture rotation matrix and the local coordinate system rotation matrix can be calculated based on the third coordinate of the first key point, the third coordinate of the second key point and the coordinates of the facial key points of the preset face image, and the rotation angle of the face posture is calculated based on the face posture rotation matrix and the local coordinate system rotation matrix.

[0087] Figure 5 is a schematic diagram of key points provided by at least one embodiment of the present disclosure. Figure 6 is a schematic diagram of the rotation angle of a facial posture provided by at least one embodiment of the present disclosure. Figure 7 is a schematic diagram of a coordinate system provided by at least one embodiment of the present disclosure.

[0088] As shown in Figure 5, the present disclosure selects the coordinates of 5 facial key points for facial posture estimation. The key points include the left eye center key point, the right eye center key point, the nose tip key point, the left mouth corner key point and the right mouth corner key point. Referring to Figure 5, the present disclosure can measure the eye position coordinates of the 3D key points based on the third coordinates of the above 5 key points. In the process of measuring the 3D eye position, the measurement can be performed through a binocular camera or a monocular camera. Compared with the 3D eye position measurement method of a binocular camera, or the 3D eye position measurement method based on 17 key points of a monocular camera, the method of calculating the 3D eye position coordinates based on 5 2D image key points of a monocular camera disclosed in the present disclosure greatly reduces the amount of calculation. On the basis of meeting the requirements of accuracy and stability, the calculation time on a low-computing-power terminal device is only 0.5 milliseconds.

[0089] As shown in Figure 6, the facial posture consists of pitch, yaw, and roll, which correspond to the rotation angles around the X-axis, Y-axis, and Z-axis, respectively. Taking the preset face image as the standard face image as an example, the average position of the 3D face key points and the corresponding position of the 3D face projected onto the 2D camera image can be obtained. The coordinates of the 3D key points correspond to the coordinates of the 2D key points one by one, and are represented by Xi and xi, respectively, to estimate the camera projection matrix PA. Given N correspondences from the camera global coordinate system coordinates to the key points on the 2D camera image Among them, N ≥ 4, the maximum likelihood estimate of the camera projection matrix PA can be determined. The face pose rotation matrix can be obtained by decomposing the camera projection matrix PA into scaling, rotation and translation transformations

[0090] As shown in Figures 6 and 7, the face pose rotation matrix is the face pose rotation matrix in the camera's local coordinate system, which needs to be converted to the face pose rotation matrix R in the camera's global coordinate system h For example, the two-dimensional coordinates x of the center of the left eye and the center of the right eye can be obtained according to the third coordinate of the left eye center key point and the third coordinate of the right eye center key point. c , the camera local coordinate system rotation matrix R is calculated based on the following formula loc . Z loc =M -1 x c Z loc =Z loc / ‖Z loc ‖ X loc =cross([0,1,0] T ,Z loc ) Y loc =cross(Z loc ,X loc ) R loc =[X loc ; Y loc ; Z loc ]

[0091] Among them, X loc 、Y loc , Z loc They are the X-axis vector, Y-axis vector, and Z-axis vector of the camera's local coordinate system, M is the camera's internal parameter matrix, and cross() is a vector cross product operation.

[0092] The face posture rotation matrix R in the camera global coordinate system is obtained based on the following formula h .

[0093] Based on the following formula, R h Decomposed into three different axial rotations.

[0094] Based on the following formula, θ can be obtained x ,θ y and θ z .

[0095] When sy ≥ 1e-6, θ x ,θ y and θz It can be obtained based on the following formula. x =atan2(R h [2,1],R h [2,2]) θ y =atan2(-R h [2,0],sy) θ z =atan2(R h [1,0],R h [0,0])

[0096] When sy<1e-6, θ x ,θ y and θ z It can be obtained based on the following formula. x =atan2(-R h [1,2],R h [1,1]) θ y =atan2(-R h [2,0],sy) θ z =0

[0097] Among them, θ x ,θ y and θ z They represent the rotation angles of the face around the X-axis, Y-axis, and Z-axis respectively, and atan2() is the inverse tangent function.

[0098] The pupil distance d of the human eye is known. According to the pupil distance d and the face posture rotation matrix R in the camera global coordinate system h , based on the following formula, the 3D human eye position coordinates can be obtained. l =R h [-d / 2,1,0] T +X c X r =R h [d / 2,1,0] T +X c

[0099] Among them, X l is the 3D coordinate of the left eye, X r is the 3D coordinate of the right eye, X c is the 3D coordinate of the center of the left eye and the right eye. Based on the above formula, X l With X r Use X respectively c 、R h and d, thus reducing the number of unknown variables from 6 to 3.

[0100] X cAs the independent variable, the following equation is established based on the projection formula of the pinhole camera. h [-d / 2,1,0] T +X c )=k l [x l [0],x l [1],1] T M(R h [d / 2,1,0] T +X c )=k r [x r [0],x r [1],1] T

[0101] Among them, X c is the independent variable, x l is the 2D coordinate of the left eye, x r is the 2D coordinate of the right eye, k l and k r is the preset coefficient, f x 、f y , x0, y0 are all elements of the camera internal parameter matrix M. The inventors have found that the preset coefficient k l and k r The impact on the solution result is small. In order to simplify the equation, the inventors of this disclosure change k l and k r After deletion, the number of valid equations is 4, which is greater than the number of independent variables.

[0102] Based on the coordinates obtained by the detection of the first image frame as two-dimensional coordinates, in the process of converting the two-dimensional coordinates into three-dimensional coordinates, in order to obtain the depth coordinate value in the Z-axis direction in the three-dimensional coordinate system, the inventors of the present disclosure further proposed a robust solution formula. c Estimation is performed to obtain more stable and accurate three-dimensional human eye position coordinates. Based on the rotation angle and pupil distance, the human eye position coordinates are obtained, including: obtaining the human eye position coordinates through the following formula:

[0103] Among them, X c [2] represents the component of the eye position coordinate on the z axis, d represents the pupil distance, [0] represents the component on the x axis, [1] represents the component on the y axis, [2] represents the component on the z axis, and f x 、f y , x0, y0 are preset constants, C y =cosθ y , C z =cosθ z, S y = sinθ y , where θ y and θ z are the components of the rotation angle on the y-axis and z-axis, respectively, Δx = x r -x l , Among them, the first key point includes the eye key point, x r Indicates the third coordinate of the left eye center key point in the eye key point, x l Indicates the third coordinate of the right eye center keypoint in the eye keypoint.

[0104] Based on the following formula, the Z-axis coordinates of the left eye center point and the right eye center point in the human eye position coordinates can be obtained. l [2]=[R h [2,0],R h [2,1],R h [2,1],R h [2,2]][-d / 2,1,0] T +X c [2] X r [2]=[R h [2,0],R h [2,1],R h [2,1],R h [2,2]][d / 2,1,0] T +X c [2]

[0105] Based on the following formula, the position coordinates of the left eye center point and the right eye center point in the human eye position coordinates can be obtained.

[0106] FIG8 is a flowchart of an image processing method provided by at least one embodiment of the present disclosure.

[0107] As shown in FIG2 , FIG3 and FIG8 , in some examples, before performing key point detection on a facial image and obtaining the first coordinates of the key points, that is, before step S120 , steps S101 to S102 are also included.

[0108] Step S101: extracting features from a facial image using a neural network model to obtain a first feature map 120;

[0109] Step S102 : down-sample the first feature map 120 to obtain a second feature map 130 .

[0110] 8 , for example, in step S101 , the neural network model may be a convolutional neural network (CNN). For example, the size of the face image may be 256*256*3, the size of the first feature map 120 may be 64*64*32, and the size of the second feature map 130 may be 32*32*32.

[0111] As shown in Figures 2, 3, and 8, in some examples, key point detection is performed on a facial image to obtain the first coordinates of the key points, that is, step S120, which includes: performing key point detection on the second feature map 130 to obtain the first coordinates of the key points. After extracting the first feature map 120 through the convolutional neural network, the image size can be reduced by downsampling the first feature map 120 to obtain a smaller second feature map 130. As a result, the CNN network is more lightweight, and when performing key point detection on the second feature map 130, the amount of computation can be reduced, the computation speed can be increased, and the first coordinates of the key points can be obtained more quickly.

[0112] As shown in Figures 2, 3, and 8, in some examples, performing facial image segmentation on a facial image to obtain partial facial images, i.e., step S130, includes performing facial image segmentation on the first feature map 120 to obtain partial facial images. By segmenting the first feature map 120 before downsampling, the clarity of the first partial image 121 and the second partial image 122 can be increased, thereby improving the accuracy of the second coordinates of the first key point and the second coordinates of the second key point.

[0113] FIG9 is a flowchart of an image processing method provided by at least one embodiment of the present disclosure.

[0114] In some examples, the neural network model is a quantized neural network model. By quantizing the neural network model, the model can be optimized, thereby improving the inference speed of the neural network model while maintaining its accuracy. Compared to a non-quantized neural network model, a quantized neural network model can increase the inference speed of the model by 2 to 3 times while maintaining accuracy.

[0115] As shown in FIG. 2 , FIG. 3 and FIG. 9 , in some examples, the image processing method further includes steps S210 to S240 .

[0116] Step S210: quantizing the initial neural network model to obtain a current neural network model;

[0117] Step S220: determining whether the current neural network model satisfies the computational constraints;

[0118] Step S230: In response to the current neural network model satisfying the computational constraint condition, using the current neural network model as a quantized neural network model;

[0119] Step S240: In response to the current neural network model not meeting the computational constraints, the output errors of the processing layers in the current neural network model are obtained, and the processing layer with the largest output error is corrected until the current neural network model meets the computational constraints.

[0120] As shown in Figures 2, 3, and 9, for example, in step S210, the initial neural network model can be quantized using the Int8 quantization method. Int8 quantization is a technique for converting neural network weights and activation values ​​into 8-bit integers (int8). This technique can be used to reduce the storage space and computing resources required by the neural network while maintaining accuracy.

[0121] As shown in Figures 2, 3, and 9, for example, in step S220, the computational constraints may include constraints on at least one of computational speed and model accuracy. For example, the computational speed and model accuracy may be evaluated separately to determine whether both meet the requirements.

[0122] As shown in Figures 2, 3 and 9, for example, in step S230, if the computing speed and model accuracy of the current neural network model meet the requirements, it is determined that the quantization of the current neural network model is completed, and the current neural network model is used as the quantized neural network model for calculation.

[0123] As shown in Figures 2, 3, and 9, for example, in step S240, if at least one of the computational speed and model accuracy of the current neural network model does not meet the requirements, an error analysis is performed on each processing layer of the current neural network model to obtain the computational output error of each processing layer. By selecting the processing layer with the largest output error for correction, the accuracy of the processing layer with the largest output error can be improved, thereby completing the quantization processing of the model at a faster quantization speed.

[0124] In some embodiments, before each acquisition of a facial image in the current image frame, the process also includes performing region detection on the current image frame to obtain a facial region, tracking the face in the facial region, and then extracting the facial image from the current image frame to obtain the key point positioning results. In research, the inventors found that when a user views 3D display content, the user's movement frequency is low and the range of movement is small. For most of the time during the viewing process, the face is located in an area near the center of the camera's field of view. Based on this, the face tracking step can be omitted during the user's viewing process. In addition, the inventors of the present disclosure also found that due to the low level of facial movement, the coordinate changes of the key points obtained during eye tracking are relatively small. To this end, the inventors added a determination of whether the key point positioning results in the previous frame are stable. When the positioning results are stable, the facial region detection step and the facial image extraction step can be omitted, thereby improving the speed, stability, and accuracy of eye tracking.

[0125] FIG10 is a flowchart of an image processing method provided by at least one embodiment of the present disclosure.

[0126] Referring to FIG. 10 , in some examples, steps S310 to S330 are also included.

[0127] Step S310: Acquire a second image frame, wherein the first image frame and the second image frame are adjacent in the time domain, and the first image frame is a forward frame of the second image frame;

[0128] Step S320: determining whether the score values ​​corresponding to the third coordinate of the first key point and the third coordinate of the second key point obtained based on the first image frame respectively meet preset requirements;

[0129] Step S330: In response to the score value of the third coordinate of the first key point and the score value of the third coordinate of the second key point both reaching the preset requirements, the face image in the second image frame is obtained based on the third coordinate of the first key point and the third coordinate of the second key point obtained in the first image frame.

[0130] It should be noted that the "first image frame" and the "second image frame" are used to refer to any two temporally continuous or adjacent frames in an image frame sequence. The "first image frame" is used to refer to the previous frame of two temporally adjacent frames, and the "second image frame" is used to refer to the next frame of two temporally adjacent frames. Neither the "first image frame" nor the "second image frame" is limited to a specific frame of image, nor is it limited to a specific order. It should also be noted that the embodiments of the present disclosure use the forward frame of two adjacent frames as a reference, and may also use the backward frame of two adjacent frames as a reference, as long as the entire image processing method remains consistent.

[0131] As shown in Figures 2, 3 and 10, for example, in step S320, the third coordinate of the first key point and the third coordinate of the second key point can be scored respectively. For example, the preset requirement includes that the score value of the third coordinate of the first key point and the score value of the third coordinate of the second key point are both greater than a preset score threshold. For example, the preset score threshold can be 0 (for example, last_face_score>0). For example, if no face is detected in the first image frame, the size of the face in the first image frame is too small, the rotation angle of the face in the first image frame is too large, the quality of the face in the first image frame is too low, etc., it is judged that the score value is less than or equal to 0. For example, when the score value does not meet the preset requirement, the second image frame is used as the first image frame to obtain a face image.

[0132] As shown in Figures 2, 3, and 10, for example, in step S330, when the score value reaches the preset requirement, it can be determined that the previous frame of image (e.g., the first image frame) can stably obtain the coordinates of the key points, and based on the third coordinates of the first key point and the third coordinates of the second key point, the facial image is obtained in the next frame of image (e.g., the second image frame). It can be understood that since the third coordinates of the first key point and the third coordinates of the second key point obtained in the first image frame are obtained by fusing the roughly positioned first coordinates and the finely positioned second coordinates of the first image frame, the accuracy is relatively high. Therefore, the result of obtaining the facial image in the second image frame based on the third coordinates also has better accuracy. In addition, the step of pre-processing the second image frame can be omitted, which reduces the amount of calculation and improves the calculation speed.

[0133] As shown in Figures 2, 3, and 10, in some examples, obtaining a facial image in the second image frame based on the third coordinates of the first key point and the third coordinates of the second key point obtained in the first image frame, i.e., step S330, includes: obtaining an affine transformation matrix based on the third coordinates of the first key point and the third coordinates of the second key point; and obtaining the facial image in the second image frame based on the affine transformation matrix. For example, after obtaining the third coordinates of the first key point and the third coordinates of the second key point, the affine transformation matrix can be calculated using the least squares method. By performing an affine transformation on the second image frame, the areas corresponding to the key points can be extracted, thereby obtaining a facial image.

[0134] FIG11 is a flow chart of an image processing method provided by at least one embodiment of the present disclosure.

[0135] Referring to Figure 11, in some examples, the first image frame can be preprocessed. The preprocessing can include steps S11 and S12. Step S11: Performing region detection on the first image frame to obtain a face region; Step S12: Performing face extraction on the face region to obtain a face image. By performing region detection and face extraction on the first image frame, more accurate keypoint location results can be obtained.

[0136] FIG12 is a schematic diagram of region detection provided by at least one embodiment of the present disclosure.

[0137] Referring to FIG12 , in some examples, a common target detection algorithm can be used to perform region detection on the first image frame. For example, the face region in the first image frame can be detected using a YOLOv5 model. For example, the face region of the first image frame can be detected using a target detection model. For example, the target detection model can simultaneously detect the head region of the first image frame based on detecting the face region of the first image frame, thereby correcting the face region detection result through the detection of the head region. By improving the single-task face region detection to a multi-task model architecture of face region detection and head region detection, a more accurate face region can be obtained through head region detection when the face is occluded, thereby improving the accuracy and stability of face region detection. For example, referring to FIG12 , the target detection model includes an input layer (Input), a backbone layer (Backbone), a fusion layer (Neck), and a processing layer (Head). For example, an image (e.g., the first image frame) can be input through the input layer. The backbone layer is the backbone network of the target detection model, capable of performing multi-level feature extraction on the input image. The fusion layer bridges the backbone layer and the processing layer, fusing feature maps from different levels extracted by the backbone layer to obtain richer semantic information. For example, the fusion layer can fuse feature maps from different levels extracted by the backbone layer through top-down and bottom-up fusion to generate feature maps with multi-scale information. The processing layer processes the feature maps extracted by the backbone layer to detect objects. For example, the processing layer can divide the feature map into multiple regions and perform detection on each region.

[0138] In some examples, to improve the detection speed of facial regions, the number of channels in the backbone network of the target detection model can be reduced, and the size of the first image frame can also be reduced. For example, the number of channels in the backbone network of the target detection model can be reduced by 40% to 60%. For example, the number of channels in the backbone network can be reduced by 45% to 55%. For example, the number of channels in the backbone network can be reduced by 50%. For example, the size of the first image frame can be reduced by 40% to 60%. For example, the size of the first image frame can be reduced by 45% to 55%. For example, the size of the first image frame can be reduced by 50%.

[0139] In some embodiments, the image acquisition thread and the image analysis thread can be executed serially. For example, an ultra-high-speed camera with an acquisition speed of up to 1000 frames per second (FPS) can be used to acquire images, and the image analysis thread can be executed serially to analyze the acquired images and obtain the coordinates of the human eye position. However, the inventors of the present disclosure have found in their research that image acquisition under the serial mechanism is time-consuming, the overall processing speed is difficult to reduce, and the cost of ultra-high-speed cameras is high.

[0140] FIG13 is a flowchart of an image processing method provided by at least one embodiment of the present disclosure.

[0141] Referring to FIG13 , in some examples, a first thread and a second thread are executed in parallel, the first thread including an image acquisition thread, and the second thread including an image analysis thread. For example, the image acquisition thread includes acquiring a first image frame and storing it in a cache space. For example, a cache unit may be provided, and the acquired first image frame may be stored in the cache space through the cache unit; for example, the cache unit may also be implemented by a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), a field programmable gate array (FPGA), or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and corresponding computer instructions.

[0142] For example, the image analysis thread includes at least one of the following operations: extracting the first image frame in the cache space, obtaining the face image in the first image frame, performing key point detection on the face image, performing face image segmentation on the face image, performing key point detection on the first partial image, performing key point detection on the second partial image, fusing the first coordinate of the first key point and the second coordinate of the first key point, fusing the first coordinate of the second key point and the second coordinate of the second key point, and obtaining the eye position coordinates based on the third coordinate of the first key point, the third coordinate of the second key point and the preset face image.

[0143] By executing the first thread and the second thread in parallel, image acquisition and image analysis can be performed in parallel, thereby reducing the image acquisition time by 9 milliseconds, improving the overall processing speed, and being applicable to cameras with lower acquisition speeds, thereby reducing costs.

[0144] Figure 14 is a schematic diagram of an image processing method. Figure 15 is a schematic diagram of an image processing method provided by at least one embodiment of the present disclosure. Figure 16 is a schematic diagram of an image processing method provided by at least one embodiment of the present disclosure.

[0145] During the study, the inventors of the present application found that, in some embodiments, the FPGA in the naked-eye 3D display system performs pixel rearrangement processing with a fixed duration. When analyzing the image, since there are differences in each frame of image acquired, the duration of the image analysis and processing is correspondingly different. For example, referring to FIG14 , the FPGA performs pixel rearrangement processing with a fixed duration of 16.7 milliseconds. When the image processing method is applied to a terminal device, the duration of the terminal device outputting the human eye position coordinates is not fixed. For example, referring to FIG14 , the output time interval of the human eye position coordinates may be 13.3 milliseconds, 15 milliseconds, etc. Therefore, when synchronization processing is not performed, the timing of the human eye position coordinates output by the terminal device is not synchronized with the FPGA, which will cause flickering problems.

[0146] Referring to FIG. 15 , in some examples, the process further includes generating a synchronization signal based on a pixel rearrangement operation. For example, the pixel rearrangement operation is performed by an FPGA in a naked-eye 3D display system. The second thread includes extracting a first image frame from a cache in response to the synchronization signal. For example, by waiting for the synchronization signal sent by the FPGA, the terminal device extracts the first image frame from the cache and performs image analysis, thereby preventing flicker in the display.

[0147] Referring to Figure 16, a terminal device that applies the image processing method provided by an embodiment of the present disclosure can obtain a first image frame through a camera. For example, the camera can be a high-speed camera. The terminal device receives a synchronization signal, performs key point detection on the first image frame when the synchronization condition is met, and sends the detected human eye position coordinates in the display coordinate system in the physical space to the PC (computer end) and FPGA. The PC generates 3D content based on the human eye position coordinates and outputs a video stream to the FPGA. The FPGA performs a pixel rearrangement operation based on the video stream and the received human eye position coordinates, and sends the pixel rearrangement signal to the display for 3D display. For example, the display can be an 8K display. For example, the display can be a 3D display.

[0148] The image processing method provided by the disclosed embodiments can achieve a frame rate greater than 65 frames per second (FPS), with the X-axis error of the eye position coordinates less than 0.2° and the Y-axis error less than 0.2°. Compared with the actual three-dimensional coordinates, the Z-axis error estimated based on the two-dimensional coordinates is less than 4%. In terms of stability, the coordinate fluctuations on the X-axis, Y-axis, and Z-axis are less than 0.04°, less than 0.04°, and less than 0.2%, respectively. Synchronization can reach 100%, laying a solid foundation for achieving excellent display effects.

[0149] FIG17 is a schematic block diagram of an image processing apparatus provided by at least one embodiment of the present disclosure.

[0150] At least one embodiment of the present disclosure further provides an image processing device 400. As shown in FIG17 , the image processing device 400 may include an acquisition module 401, a first detection module 402, a segmentation module 403, a second detection module 404, a fusion module 405, and a calculation module 406. These components are interconnected by a bus system and / or other forms of connection mechanisms (not shown). For example, these modules can be implemented by hardware (e.g., circuit) modules, software modules, or any combination thereof. The following embodiments are the same and will not be described in detail. For example, these units can be implemented by a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), a field programmable gate array (FPGA), or other forms of processing units with data processing capabilities and / or instruction execution capabilities, as well as corresponding computer instructions. It should be noted that the components and structures of the image processing device 400 shown in FIG17 are exemplary only and not restrictive. As needed, the image processing device 400 may also have other components and structures.

[0151] For example, the acquisition module 401 is configured to acquire a face image in the first image frame.

[0152] For example, the first detection module 402 is configured to perform key point detection on the face image to obtain first coordinates of the key points, wherein the first coordinates of the key points include the first coordinates of the first key point and the first coordinates of the second key point.

[0153] For example, the segmentation module 403 is configured to perform facial image segmentation on the facial image to obtain partial facial images; the partial facial images include a first partial image and a second partial image.

[0154] For example, the second detection module 404 is configured to perform key point detection on the first partial image to obtain the second coordinates of the first key point, and is configured to perform key point detection on the second partial image to obtain the second coordinates of the second key point.

[0155] For example, the fusion module 405 is configured to fuse the first coordinate of the first key point and the second coordinate of the first key point to obtain the third coordinate of the first key point, and is configured to fuse the first coordinate of the second key point and the second coordinate of the second key point to obtain the third coordinate of the second key point.

[0156] For example, the calculation module 406 is configured to obtain the eye position coordinates based on the third coordinate of the first key point, the third coordinate of the second key point, and the preset face image.

[0157] For example, the acquisition module 401, the first detection module 402, the segmentation module 403, the second detection module 404, the fusion module 405, and the calculation module 406 may include code and programs stored in a memory; the processor may execute the code and programs to implement some or all of the functions of the acquisition module 401, the first detection module 402, the segmentation module 403, the second detection module 404, the fusion module 405, and the calculation module 406 described above. For example, the acquisition module 401, the first detection module 402, the segmentation module 403, the second detection module 404, the fusion module 405, and the calculation module 406 may be dedicated hardware devices to implement some or all of the functions of the acquisition module 401, the first detection module 402, the segmentation module 403, the second detection module 404, the fusion module 405, and the calculation module 406 described above. For example, the acquisition module 401, the first detection module 402, the segmentation module 403, the second detection module 404, the fusion module 405, and the calculation module 406 may be a single circuit board or a combination of multiple circuit boards to implement the functions described above. In an embodiment of the present application, the circuit board or the combination of multiple circuit boards may include: (1) one or more processors; (2) one or more non-temporary memories connected to the processors; and (3) firmware stored in the memories that is executable by the processors.

[0158] It should be noted that the acquisition module 401 can be used to implement step S110 shown in Figure 2, the first detection module 402 can be used to implement step S120 shown in Figure 2, the segmentation module 403 can be used to implement step S130 shown in Figure 2, the second detection module 404 can be used to implement step S140 shown in Figure 2, the fusion module 405 can be used to implement step S150 shown in Figure 2, and the calculation module 406 can be used to implement step S160 shown in Figure 2. Therefore, for a detailed description of the functions that can be implemented by the acquisition module 401, the first detection module 402, the segmentation module 403, the second detection module 404, the fusion module 405, and the calculation module 406, reference can be made to the relevant descriptions of steps S110 to S160 in the embodiment of the above-mentioned image processing method, and any repetitions will not be repeated here. In addition, the image processing device 400 can achieve technical effects similar to those of the above-mentioned image processing method, which will not be repeated here.

[0159] It should be noted that in the embodiments of the present disclosure, the image processing device 400 may include more or fewer circuits or units, and the connection relationship between the various circuits or units is not limited and can be determined according to actual needs. The specific configuration of each circuit or unit is not limited and can be composed of analog devices according to circuit principles, or can be composed of digital chips, or constructed in other applicable ways.

[0160] For example, in some examples, the image processing device 400 may further include a quantization module, which is configured to: quantize the initial neural network model to obtain a current neural network model; determine whether the current neural network model satisfies the operation restriction conditions; in response to the current neural network model satisfying the operation restriction conditions, use the current neural network model as the quantized neural network model; in response to the current neural network model not satisfying the operation restriction conditions, obtain the output error of each processing layer in the current neural network model, and correct the processing layer with the largest output error until the current neural network model satisfies the operation restriction conditions.

[0161] For example, in some examples, the image processing module may also include a judgment module, which is configured to: obtain a second image frame, wherein the first image frame and the second image frame are adjacent in the time domain, and the first image frame is the forward frame of the second image frame; determine whether the score values ​​corresponding to the third coordinates of the first key point and the third coordinates of the second key point obtained based on the first image frame both meet preset requirements; in response to the score value of the third coordinate of the first key point and the score value of the third coordinate of the second key point both meeting the preset requirements, obtain a facial image in the second image frame based on the third coordinates of the first key point and the third coordinates of the second key point obtained in the first image frame.

[0162] For example, in some examples, the image processing module may further include a preprocessing module, which is configured to: perform region detection on the first image frame to obtain a face region; and perform face extraction on the face region to obtain a face image.

[0163] For example, the acquisition module 401, the first detection module 402, the segmentation module 403, the second detection module 404, the fusion module 405, the calculation module 406, the quantification module, the judgment module and the preprocessing module may also implement more or further functions.

[0164] For example, the segmentation module 403 can also be configured to: segment the facial image according to the eye area and the mouth and nose area to obtain a first partial image and a second partial image; the first partial image includes the eye area image, and the second partial image includes the mouth and nose area image.

[0165] For example, the fusion module 405 may be further configured to: fuse the first coordinate of the eye key point with the second coordinate of the eye key point to obtain the third coordinate of the eye key point; and fuse the first coordinate of the mouth and nose key point with the second coordinate of the mouth and nose key point to obtain the third coordinate of the mouth and nose key point.

[0166] For example, the acquisition module 401 can also be configured to: extract features from the facial image using a neural network model to obtain a first feature map; and downsample the first feature map to obtain a second feature map. For example, the first detection module 402 can also be configured to: perform key point detection on the second feature map to obtain first coordinates of the key points.

[0167] For example, the segmentation module 403 may also be configured to perform facial image segmentation on the first feature map to obtain a partial facial image.

[0168] For example, the judgment module can also be configured to: obtain an affine transformation matrix based on the third coordinate of the first key point and the third coordinate of the second key point; and obtain a facial image in the second image frame based on the affine transformation matrix.

[0169] For example, the calculation module 406 can also be configured to: obtain the rotation angle of the face orientation in the face image relative to the face orientation of the preset face image based on the third coordinate of the first key point, the third coordinate of the second key point and the coordinates of the face key point of the preset face image; obtain the pupil distance of the human eye; and obtain the coordinates of the human eye position based on the rotation angle and the pupil distance.

[0170] FIG18 is a schematic block diagram of an electronic device provided by at least one embodiment of the present disclosure.

[0171] Some embodiments of the present disclosure further provide an electronic device 500. For example, as shown in FIG18 , the electronic device 500 includes a photographing device 501 and an image processing device 502.

[0172] For example, the photographing device 501 is configured to photograph a first image frame.

[0173] For example, the image processing device 502 is configured to receive a first image frame and perform any one of the above-mentioned image processing methods based on the first image frame.

[0174] For example, the photographing device 501 may be a rear camera of the electronic device 500 , or a front camera and a reflection device of the electronic device 500 .

[0175] For example, the image processing device 502 can be implemented as a central processing unit, a dedicated processing chip, a digital signal processor, etc., and this disclosure does not impose any specific limitations on this.

[0176] For example, the electronic device 500 can be a terminal device such as a naked-eye 3D display, and can also be provided with a display unit (such as a touch screen), etc. For example, the display unit can provide a corresponding human-computer interaction interface for displaying responses to interactive operations, interactive action prompt information, etc. The present disclosure does not impose specific restrictions on this.

[0177] For example, for a detailed description of the process of the electronic device 500 executing the image processing method, reference may be made to the relevant description in the above-mentioned embodiment of the image processing method, and repeated parts will not be repeated here.

[0178] FIG19 is a schematic block diagram of another electronic device provided by at least one embodiment of the present disclosure.

[0179] Some embodiments of the present disclosure also provide another electronic device. For example, as shown in FIG19 , electronic device 600 includes a memory 601 and a processor 602. It should be noted that the components of electronic device 600 shown in FIG19 are merely exemplary and non-limiting. Depending on actual application requirements, electronic device 600 may also have other components.

[0180] For example, the memory 601 non-transiently stores computer-executable instructions, and the processor 602 is configured to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor 602, the image processing method according to any of the above-described embodiments is implemented. The specific implementation and related explanations of each step of the image processing method can be found in the above-described embodiments of the image processing method, and any repetitive details are not repeated here.

[0181] For example, the memory 601 may include any combination of one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), a hard disk, an erasable programmable read-only memory (EPROM), a portable compact disk read-only memory (CD-ROM), a USB memory, a flash memory, etc. One or more computer-readable instructions may be stored on the computer-readable storage medium, and the processor 602 may execute the computer-readable instructions to implement various functions of the electronic device 600. Various applications and various data may also be stored in the storage medium.

[0182] For example, processor 602 can control other components in electronic device 600 to perform desired functions. Processor 602 can be a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The central processing unit (CPU) can be an X86 or ARM architecture, etc.

[0183] For example, in some embodiments, the electronic device 600 may be a mobile phone, a tablet computer, an electronic paper, a television, a monitor, a laptop computer, a digital photo frame, a navigator, a wearable electronic device, a smart home device, etc.

[0184] For example, the electronic device 600 may include a display panel, which may be used to segment an image, etc. For example, the display panel may be a rectangular panel, a circular panel, an elliptical panel, or a polygonal panel, etc. In addition, the display panel may be not only a flat panel but also a curved panel or even a spherical panel.

[0185] For example, the electronic device 600 may have a touch function, that is, the electronic device 600 may be a touch device.

[0186] For example, for a detailed description of the process of the electronic device executing the image processing method, reference may be made to the relevant description in the above-mentioned embodiment of the image processing method, and repeated descriptions will be omitted.

[0187] FIG20 is a schematic diagram of a non-transitory computer-readable storage medium provided by at least one embodiment of the present disclosure.

[0188] For example, as shown in FIG20 , a non-transitory computer-readable storage medium 700 stores computer-executable instructions 701 , which, when executed by a processor, can implement any of the above-mentioned image processing methods.

[0189] For example, the storage medium 700 may be applied to the electronic device 600. For example, the storage medium 700 may include the memory 601 in the electronic device 600.

[0190] For example, for description of the storage medium, reference may be made to the description of the memory in the embodiment of the electronic device, and repeated descriptions will be omitted.

[0191] Figure 21 is a schematic diagram of a hardware environment provided by at least one embodiment of the present disclosure. The electronic device provided by the present disclosure can be applied in an Internet system.

[0192] The computer system provided in the figure can be used to implement the functions of the image processing device and / or electronic device involved in this disclosure. Such computer systems may include personal computers, laptops, tablets, mobile phones, personal digital assistants, smart glasses, smart watches, smart rings, smart helmets, and any smart portable or wearable device. The specific system in this embodiment uses a functional block diagram to illustrate a hardware platform including a user interface. This computer device can be a general-purpose computer device or a computer device with a specific purpose. Both computer devices can be used to implement the image processing device and / or electronic device in this embodiment. The computer system can include any components that implement the information required to implement the image processing described herein. For example, the computer system can be implemented by the computer device through its hardware devices, software programs, firmware, and combinations thereof. For convenience, only one computer device is depicted in the figure, but the relevant computer functions for implementing the information required for image processing described in this embodiment can be implemented in a distributed manner by a group of similar platforms, thereby distributing the processing load of the computer system.

[0193] As shown in FIG21 , the computer system may include a communication port 850 connected to a network for data communication. For example, the computer system can send and receive information and data via the communication port 850. That is, the communication port 850 enables the computer system to communicate with other electronic devices wirelessly or wired to exchange data. The computer system may also include a processor group 820 (i.e., the processor described above) for executing program instructions. The processor group 820 may be composed of at least one processor (e.g., a CPU). The computer system may include an internal communication bus 810. The computer system may include various forms of program storage units and data storage units (i.e., the memory or storage media described above), such as a hard disk 870, a read-only memory (ROM) 830, and a random access memory (RAM) 840, which can be used to store various data files used for computer processing and / or communication, as well as possible program instructions executed by the processor group 820. The computer system may also include an input / output component 860, which is used to implement input / output data flow between the computer system and other components (e.g., a user interface 880, etc.).

[0194] Typically, the following devices can be connected to the input / output component 860: input devices such as a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices such as a display (e.g., an LCD, an OLED display, etc.), a speaker, a vibrator, etc.; storage devices such as a tape, a hard disk, etc.; and a communication interface.

[0195] Although FIG. 21 shows a computer system having various devices, it should be understood that the computer system is not required to have all of the devices shown, and the computer system may alternatively have more or fewer devices.

[0196] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

[0197] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0198] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

[0199] Regarding this disclosure, the following points need to be explained:

[0200] (1) The drawings of the embodiments of the present disclosure only involve structures related to the embodiments of the present disclosure, and other structures can refer to general designs.

[0201] (2) In the absence of conflict, features in the same embodiment and different embodiments of the present disclosure may be combined with each other.

[0202] The foregoing description is merely an exemplary embodiment of the present disclosure and is not intended to limit the scope of protection of the present disclosure. The scope of protection of the present disclosure is determined by the appended claims.

Claims

1. An image processing method, comprising: Obtaining a face image in a first image frame; Performing key point detection on the face image to obtain first coordinates of the key points; The first coordinates of the key points include first coordinates of a first key point and first coordinates of a second key point; Performing face image segmentation on the face image to obtain a face local image; the face local image includes a first local image and a second local image; Performing key point detection on the first local image to obtain second coordinates of the first key point; performing key point detection on the second local image to obtain second coordinates of the second key point; Fusing the first coordinates of the first key point and the second coordinates of the first key point to obtain third coordinates of the first key point; Fusing the first coordinates of the second key point and the second coordinates of the second key point to obtain third coordinates of the second key point; Based on the third coordinates of the first key point, the third coordinates of the second key point and a preset face image, obtaining eye position coordinates.

2. The image processing method according to claim 1, wherein, Performing face image segmentation on the face image to obtain a face local image, including: Segmenting the face image according to the area where the eyes are located and the area where the mouth and nose are located to obtain the first local image and the second local image; The first local image includes an eye area image, and the second local image includes a mouth and nose area image.

3. The image processing method according to claim 2, wherein, Determining the area where the eyes are located based on the first coordinates of the first key point, and determining the area where the mouth and nose are located based on the first coordinates of the second key point.

4. The image processing method according to claim 2, wherein, The first key point includes an eye key point, and the second key point includes a mouth and nose key point.

5. The image processing method according to claim 4, wherein, Fusing the first coordinates of the first key point and the second coordinates of the first key point to obtain third coordinates of the first key point, including: fusing the first coordinates of the eye key point and the second coordinates of the eye key point to obtain third coordinates of the eye key point; Fusing the first coordinates of the second key point and the second coordinates of the second key point to obtain third coordinates of the second key point, including: fusing the first coordinates of the mouth and nose key point and the second coordinates of the mouth and nose key point to obtain third coordinates of the mouth and nose key point.

6. The image processing method according to claim 5, wherein, The eye key points include a left eye center key point and a right eye center key point; The mouth and nose key points include a left mouth corner key point, a right mouth corner key point and a nose tip key point.

7. The image processing method according to any one of claims 1 to 6, wherein, Before performing key point detection on the face image to obtain the first coordinates of the key points, it further includes: Performing feature extraction on the face image through a neural network model to obtain a first feature map; Performing downsampling on the first feature map to obtain a second feature map; Performing key point detection on the face image to obtain the first coordinates of the key points, including: Performing the key point detection on the second feature map to obtain the first coordinates of the key points.

8. The image processing method according to claim 7, wherein, Performing face image segmentation on the face image to obtain a face local image, including: Performing the face image segmentation on the first feature map to obtain the face local image.

9. The image processing method according to claim 7 or 8, wherein, The neural network model is a quantized neural network model.

10. The image processing method according to claim 9 further includes: Performing quantization processing on the initial neural network model to obtain the current neural network model; Determining whether the current neural network model meets the operation restriction conditions; In response to the current neural network model meeting the operation restriction conditions, using the current neural network model as the quantized neural network model; In response to the current neural network model not meeting the operation restriction conditions, obtaining the output errors of each processing layer in the current neural network model, and correcting the processing layer with the largest output error until the current neural network model meets the operation restriction conditions.

11. The image processing method according to claim 10, wherein, The operation restriction conditions include at least one of the restriction conditions on operation speed and model accuracy.

12. The image processing method according to any one of claims 1-11 further includes: Obtaining a second image frame, wherein the first image frame and the second image frame are adjacent in the time domain, and the first image frame is the forward frame of the second image frame; Determining whether the fractional values corresponding to the third coordinates of the first key point and the third coordinates of the second key point obtained based on the first image frame both reach a preset requirement; In response to the fractional values of the third coordinates of the first key point and the fractional values of the third coordinates of the second key point both reaching the preset requirement, obtaining the face image in the second image frame based on the third coordinates of the first key point and the third coordinates of the second key point obtained from the first image frame.

13. The image processing method according to claim 12, wherein, Obtaining the face image in the second image frame based on the third coordinates of the first key point and the third coordinates of the second key point obtained from the first image frame includes: Obtaining an affine transformation matrix based on the third coordinates of the first key point and the third coordinates of the second key point; Obtaining the face image in the second image frame based on the affine transformation matrix.

14. The image processing method according to claim 12 or 13, wherein, The preset requirement includes that the fractional value of the third coordinate of the first key point and the fractional value of the third coordinate of the second key point are both greater than a preset fractional threshold.

15. The image processing method according to any one of claims 1-14 further includes preprocessing the first image frame, and the preprocessing includes: Performing region detection on the first image frame to obtain a face region; Performing face extraction on the face region to obtain the face image.

16. The image processing method according to any one of claims 1 to 15, wherein, Parallelly executing a first thread and a second thread, where the first thread includes an image acquisition thread and the second thread includes an image analysis thread; The image acquisition thread includes acquiring the first image frame and storing it in a buffer space; the image analysis thread includes at least one of the following operations: extracting the first image frame from the buffer space, obtaining the face image in the first image frame, performing key point detection on the face image, performing face image segmentation on the face image, performing key point detection on the first partial image, performing key point detection on the second partial image, fusing the first coordinate of the first key point and the second coordinate of the first key point, fusing the first coordinate of the second key point and the second coordinate of the second key point, and obtaining the eye position coordinates based on the third coordinate of the first key point, the third coordinate of the second key point, and the preset face image.

17. The image processing method according to claim 16, further comprising: Generate a synchronization signal based on pixel rearrangement operation; Wherein, the second thread includes: extracting the first image frame from the buffer space in response to the synchronization signal.

18. The image processing method according to any one of claims 1-17, wherein, Obtaining the eye position coordinates based on the third coordinate of the first key point, the third coordinate of the second key point, and the preset face image includes: Obtaining the rotation angle of the face orientation in the face image relative to the face orientation of the preset face image based on the third coordinate of the first key point, the third coordinate of the second key point, and the coordinates of the face key points of the preset face image; Obtain the pupil distance of the eyes; Based on the rotation angle and the pupil distance, obtain the eye position coordinates.

19. The image processing method according to claim 18, wherein, Obtaining the human eye position coordinates based on the rotation angle and the pupil distance includes: obtaining the component of the human eye position coordinates on the z-axis through the following formula: Among them, X c [2] represents the component of the human eye position coordinate on the z-axis, d represents the pupil distance, [0] represents the component on the x-axis, [1] represents the component on the y-axis, [2] represents the component on the z-axis, f x 、f y 、x0, y0 are preset constants, C y = cosθ y , C z = cosθ z , S y = sinθ y , where, θ y and θ z are respectively the components of the rotation angle on the y-axis and the z-axis, Δx = x r -x l , Among them, the first key points include eye key points, x r represents the third coordinate of the left eye center key point among the eye key points, x l represents the third coordinate of the right eye center key point among the eye key points.

20. An image processing apparatus, comprising: An acquisition module configured to acquire a face image in a first image frame; A first detection module configured to perform key point detection on the face image to obtain the first coordinates of the key points; the first coordinates of the key points include the first coordinates of the first key point and the first coordinates of the second key point; A segmentation module configured to perform face image segmentation on the face image to obtain a face partial image; The face partial image includes a first partial image and a second partial image; A second detection module configured to perform key point detection on the first partial image to obtain the second coordinates of the first key point, and configured to perform key point detection on the second partial image to obtain the second coordinates of the second key point; A fusion module configured to fuse the first coordinates of the first key point and the second coordinates of the first key point to obtain the third coordinates of the first key point, and configured to fuse the first coordinates of the second key point and the second coordinates of the second key point to obtain the third coordinates of the second key point; And A calculation module configured to obtain eye position coordinates based on the third coordinates of the first key point, the third coordinates of the second key point, and a preset face image.

21. An electronic device, comprising: A photographing device configured to photograph a first image frame; An image processing apparatus configured to receive the first image frame and execute the image processing method according to any one of claims 1-19 based on the first image frame.

22. An electronic device, comprising: A memory that stores computer-executable instructions non-transiently; A processor configured to run the computer-executable instructions, wherein, when the computer-executable instructions are run by the processor, an image processing method according to any one of claims 1-19 is implemented.

23. A non-transitory computer-readable storage medium, wherein, The non-transitory computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, an image processing method according to any one of claims 1-19 is implemented.

Citation Information

Patent Citations

  • Human eye attention localization method and system based on depth neural network

    CN109446892A

  • Face pose estimation method and device, storage medium and computer equipment

    CN114581973A

  • Face key point prediction method, APP, terminal device and storage medium

    CN116152896A

  • Method for driving virtual facial expressions by automatically detecting facial expressions of a face image

    US20080037836A1