Facial expression recognition method and device applying embedded device and storage medium
By detecting key points and analyzing geometric features of embedded devices, combined with expression judgment algorithms, the problem of poor facial expression recognition in embedded devices is solved, and real-time and high-precision facial expression recognition is achieved.
Patent Information
- Application Number
- CN202511121881.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-09-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing AI facial expression recognition algorithms perform poorly in embedded devices due to high computational and storage overhead, numerous model parameters, the need for GPU acceleration, difficulty in edge device deployment, poor interpretability, and insufficient performance of the embedded MIPS architecture.
A facial key point detection model is used to achieve expression recognition through key point detection and geometric feature calculation, combined with a preset expression judgment algorithm, reducing model complexity and computing requirements.
Real-time face tracking and expression recognition are achieved on embedded devices, maintaining high accuracy, reducing computing and storage requirements, and adapting to the performance limitations of embedded devices.
Smart Images

Figure CN120635970A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of face recognition, and in particular to a facial expression recognition method, device and storage medium using an embedded device. Background Art
[0002] Currently, there are numerous face recognition and facial expression detection solutions, primarily employing end-to-end CNN classification for emotion analysis. This solution first crops the face region and then uses convolutional neural networks (such as shallow CNN, Inception-V3, MobileNet, and ResNet) for multi-category expression classification. Next, emotion analysis solutions use object detection models (such as Faster R-CNN and the YOLO series) to locate local facial expression regions (such as the corners of the mouth and between the eyebrows) before classifying or regressing the emotion score. This solution achieves high accuracy, with a classic CNN achieving 75.2% on the FER2013 dataset; transfer learning (VGG16, AlexNet) can improve this to 78%. A lightweight SD-CNN achieves 99.7% and 91.3% on CK+ and Oulu-CASIA, respectively. A hybrid LBP-CNN model also achieves 90.2% on CK+. This robustness is achieved through large-scale data augmentation and deep features, making it adaptable to various interferences such as lighting, angle, and occlusion.
[0003] However, this facial expression recognition solution has high computational and storage overhead, often exceeding one million model parameters. Inference requires GPU acceleration or high-performance DSPs, and deployment on edge devices requires high optimization. Interpretability is poor, and feature abstraction makes it difficult to intuitively explain which pixel regions drive emotion judgments, requiring the use of methods such as Gradient-CAM for auxiliary analysis. Due to the limited performance of devices such as embedded MIPS architectures, some AI algorithms may not achieve the expected results or may not function properly.
[0004] Therefore, in order to address the technical problem that the current AI facial expression recognition algorithm does not perform well when deployed in embedded devices, a facial expression recognition solution is needed that is specifically optimized for embedded devices and can meet the needs of various scenarios. Summary of the Invention
[0005] The main purpose of this invention is to solve the technical problem that the current AI facial expression recognition algorithm has poor performance when deployed in embedded devices.
[0006] A first aspect of the present invention provides a method for recognizing facial expressions using an embedded device, the method comprising: Get face image; Performing key point detection processing on the face image according to a preset key point detection model to obtain a face recognition frame and N facial key points corresponding to the face recognition frame; Obtaining center point coordinates based on the coordinate data of the face recognition frame; Calculating the coordinates of the N facial key points based on the center point coordinates to obtain N key point coordinates; Calculate the horizontal relative weight and the vertical relative weight based on the N key point coordinates; Using the horizontal relative weight and the vertical relative weight, numerically analyze the coordinates of the N key points to obtain a relative value of the mouth and a relative value of the eyes; According to a preset expression determination algorithm, expression analysis processing is performed on the mouth relative value and the eye relative value to obtain the facial expression corresponding to the facial image.
[0007] Optionally, in a first implementation of the first aspect of the present invention, the N key point coordinates include: 68 ordered key point coordinates, and the calculating of the horizontal relative weight and the vertical relative weight based on the N key point coordinates includes: Calculate the Euclidean distance between the coordinates of the first key point and the coordinates of the 13th key point to obtain the horizontal relative weight; Calculate the absolute difference between the ordinate of the 7th key point and the ordinate of the 42nd key point to obtain the vertical relative weight.
[0008] Optionally, in a second implementation of the first aspect of the present invention, the step of performing numerical analysis on the coordinates of the N key points using the horizontal relative weight and the vertical relative weight to obtain the relative value of the mouth and the relative value of the eyes includes: Using the horizontal relative weight, the coordinates of the 61st key point and the 65th key point are calculated to obtain the mouth opening and closing value; Using the vertical relative weight, the ordinate of the 59th key point coordinate, the ordinate of the 63rd key point coordinate, and the ordinate of the 65th key point coordinate are calculated to obtain the mouth corner depression value; Calculating the vertical relative weights of the 59th key point coordinate, the 61st key point coordinate, and the 63rd key point coordinate to obtain a smile score; Using the horizontal relative weight, the coordinates of the 59th key point and the 63rd key point are calculated to obtain the relative mouth width; Based on the preset eye ratio algorithm, calculate the coordinates of the 37th key point to the 42nd key point to obtain the aspect ratio of the right eye; Based on the preset eye ratio algorithm, calculate the coordinates of the 43rd key point to the 48th key point to obtain the left eye aspect ratio; Calculating an average value based on the right eye aspect ratio and the left eye aspect ratio to obtain an average value of eye aspect ratio; By using the vertical relative weight, the ordinates of the 25th key point coordinate, the ordinates of the 16th key point coordinate, the ordinates of the 42nd key point coordinate, and the ordinates of the 33rd key point coordinate are calculated to obtain the eyebrow and eye difference.
[0009] Optionally, in a third implementation of the first aspect of the present invention, the step of performing expression analysis on the mouth relative value and the eye relative value according to a preset expression determination algorithm to obtain the facial expression corresponding to the facial image includes: According to a preset numerical judgment table, the mouth opening and closing value, the mouth corner depression value, the smile score, the relative mouth width, the average length and width of the eye, and the eyebrow and eye difference are matched and fitted to obtain the facial expression corresponding to the facial image.
[0010] Optionally, in a fourth implementation of the first aspect of the present invention, the step of obtaining the center point coordinates based on the coordinate data of the face recognition frame includes: When the coordinate data of the face recognition frame is corner point data, the mean value of the corner point data is calculated according to a preset mean formula to obtain the center point coordinates; When the coordinate data of the face recognition frame is the width and height data of the frame center, the width and height data of the frame center are restored according to a preset ratio restoration algorithm to obtain the center point coordinates.
[0011] Optionally, in a fifth implementation of the first aspect of the present invention, after the step of acquiring the facial image and before the step of performing key point detection processing on the facial image according to a preset key point detection model to obtain a facial recognition frame and N facial key points corresponding to the facial recognition frame, the method further includes: performing exposure enhancement processing on the face image according to a preset linear transformation algorithm to obtain an enhanced image; The enhanced image is subjected to noise reduction processing according to a preset bilateral filtering noise reduction algorithm to obtain a preprocessed face image.
[0012] Optionally, in a sixth implementation of the first aspect of the present invention, the step of performing exposure enhancement processing on the facial image according to a preset linear transformation algorithm to obtain an enhanced image includes: g(x,y)=a(f(x,y)-128)+128, Among them, g(x, y) is the enhanced image, a is the contrast parameter, f(x, y) is the face image, x is the horizontal pixel number, and y is the vertical pixel number.
[0013] Optionally, in a seventh implementation of the first aspect of the present invention, the step of performing exposure enhancement processing on the facial image according to a preset linear transformation algorithm to obtain an enhanced image includes: The fastNlMeansDenoisingColored() function in the OpenCV library is used to perform exposure enhancement processing on the face image to obtain an enhanced image.
[0014] A second aspect of the present invention provides a facial expression recognition device for an embedded device, comprising: a memory and at least one processor, wherein instructions are stored in the memory, and the memory and the at least one processor are interconnected via a line; the at least one processor calls the instructions in the memory so that the facial expression recognition device for the embedded device executes the above-mentioned facial expression recognition method for the embedded device.
[0015] A third aspect of the present invention provides a computer-readable storage medium having instructions stored therein, which, when executed on a computer, enables the computer to execute the above-mentioned method for facial expression recognition using an embedded device.
[0016] In an embodiment of the present invention, key point detection is performed using a facial key point detection model, key point geometric features are calculated, and these key point geometric features are used as recognition data. This data is then input into a pre-trained expression determination algorithm module to output the determined facial expression result. This invention achieves real-time face tracking and facial expression recognition with a smaller model, faster speed, and high accuracy, resolving the technical issue of current AI facial expression recognition algorithms operating poorly in embedded devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 A schematic diagram of an embodiment of a method for facial expression recognition using an embedded device in an embodiment of the present invention; Figure 2 103 is a schematic diagram of a specific embodiment of the method for facial expression recognition using an embedded device in an embodiment of the present invention; Figure 3 105 is a schematic diagram of a specific embodiment of the method for facial expression recognition using an embedded device in an embodiment of the present invention; Figure 4 106 is a schematic diagram of a specific embodiment of the method for facial expression recognition using an embedded device in an embodiment of the present invention; Figure 5 Schematic diagram of an embodiment of a facial expression recognition device using an embedded device in an embodiment of the present invention. DETAILED DESCRIPTION
[0018] The embodiments of the present invention provide a facial expression recognition method, a device and a storage medium using an embedded device.
[0019] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0020] In the description of the embodiments disclosed herein, the term "including" and similar terms should be understood as open inclusion, i.e., "including but not limited to." The term "based on" should be understood as "based, at least in part, on." The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment." The terms "first," "second," etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0021] For ease of understanding, the specific process of the embodiment of the present invention is described below. Figure 1 An embodiment of a facial expression recognition method using an embedded device in an embodiment of the present invention includes the following steps: 101. Obtain a facial image; In this embodiment, the current face image is captured by the deployed camera, and the image is transmitted to the device memory or converted into a specified format such as JPEG or NV12, and transmitted to other processing devices for storage.
[0022] Furthermore, after step 101 and before step 102, the following specific implementation methods are also included: 1011. Perform exposure enhancement processing on the facial image according to a preset linear transformation algorithm to obtain an enhanced image; 1012. Perform noise reduction processing on the enhanced image according to a preset bilateral filtering noise reduction algorithm to obtain a preprocessed face image.
[0023] In steps 1011-1012, linear transformation is first used to perform exposure enhancement processing on the face image. The specific processing formula is as follows: g(x, y) = af(x, y) + b, where a is the parameter controlling contrast, b is the parameter controlling brightness, g(x, y) is the enhanced image, f(x, y) is the face image, x is the horizontal pixel number, and y is the vertical pixel number.
[0024] After performing a linear transformation of exposure enhancement on the face image, an enhanced image is obtained. The enhanced image is denoised using a bilateral filtering algorithm, with the following window filter weights: W(x, y) = exp(-(||P x -P y || 2 ) / h 2 ), where for each pixel point P x , find all points P in the window y W(x, y) is used to calculate the weight of the pixel (x, y). Here, Px and Py can be the pixel's color value (e.g., RGB value) or grayscale value. This difference measures the similarity in intensity between the two pixels. h is a constant that controls the sharpness of the weight function. Smaller values of h result in steeper transitions in the weight function, meaning that only very similar pixels are assigned larger weights. Larger values of h result in smoother transitions in the weight function, meaning that more pixels are assigned larger weights.
[0025] Specifically, step 1011 includes the following specific implementation methods: If you want to increase the contrast without changing the "average brightness", you can use the following formula: g(x, y) = a(f(x, y) - 128) + 128, where g(x, y) is the enhanced image, a is the contrast parameter, f(x, y) is the face image, x is the horizontal pixel number, and y is the vertical pixel number.
[0026] Specifically, step 1012 includes the following specific implementation methods: 1021. Use the fastNlMeansDenoisingColored() function in the OpenCV library to perform exposure enhancement processing on the face image to obtain an enhanced image.
[0027] In step 1021, the fastNlMeansDenoisingColored() function is called in the OpenCV library to perform exposure enhancement on the face image to obtain an enhanced image. The calling method is as follows: cv::fastNlMeansDenoisingColored(src, dst, hL, hColor,templateWindowSize, searchWindowSize); src: Input image, must be an 8-bit or floating-point color image (with 3 or 4 channels).
[0028] dst: Output image, of the same size and type as the input image src.
[0029] hL: Luminance filter strength. This parameter controls the intensity of noise reduction for the luminance component. Larger values result in more pronounced noise reduction, but excessively large values may result in blurry images.
[0030] hColor: Chrominance filter strength. This parameter controls the noise reduction strength of the chrominance component. Similar to hL, larger values result in more pronounced noise reduction, but may cause image color distortion.
[0031] templateWindowSize: The size of the template window used to calculate the weights. This window is used to consider the size of the neighborhood when calculating the weight of each pixel. Typically, this value is set to 3, 5, or 7.
[0032] When using the fastNlMeansDenoisingColored() function, you need to adjust the hL, hColor, templateWindowSize, and searchWindowSize parameters according to your specific image and denoising needs.
[0033] 102. Perform key point detection on the face image according to a preset key point detection model to obtain a face recognition frame and N facial key points corresponding to the face recognition frame; In this embodiment, a target detection network or a face key point detection network is used for forward reasoning. Target detection can obtain a facial region, while face key point detection can obtain a facial region and 68 or 72 key points of the face. Taking 68 key points as an example, the data of the 68 key points are as follows in Table 1: Table 1. Description of 68 key points
[0034] 103. Obtaining center point coordinates based on the coordinate data of the face recognition frame; In this embodiment, the coordinate data of the center point of the entire recognition frame is found through the calibrated data of the face recognition frame.
[0035] For details, please refer to Figure 2 , Figure 2 This is a specific embodiment of step 103 of the facial expression recognition method using an embedded device in an embodiment of the present invention, and step 103 includes the following specific implementation methods: 1031. When the coordinate data of the face recognition frame is corner point data, performing mean calculation on the corner point data according to a preset mean formula to obtain the center point coordinates; 1032. When the coordinate data of the face recognition frame is the frame center width and height data, the frame center width and height data are restored according to a preset ratio restoration algorithm to obtain the center point coordinates.
[0036] In steps 1031-1032, the face detection model outputs a bounding box that surrounds the face. Its coordinates can be expressed in two common ways: (1) Corner point data If the coordinate data of the face recognition frame is corner point data, the corner point data format is (x min ,y min , x max ,y max ), where (x min ,y min ) is the pixel coordinate of the upper left corner of the detection box, (x max ,y max ) is the pixel coordinate of the lower right corner of the detection box. The center point coordinate (x center ,y center ) is calculated as follows: x center =(x min +x max ) / 2,y center =(y min +y max ) / 2.
[0037] (2) Frame width and height data When the coordinate data of the face recognition frame is the frame center width and height data, the frame center width and height data is (c x , c y ,w,h) where (c y , c y ) are the normalized pixel coordinates of the frame center, w and h are the width and height of the frame (in pixels or normalized values relative to the image width and height), respectively.
[0038] For the output (c x , c y , w, h) (normalized to [0,1]) model, needs to be restored to pixel (c x px , c y px ): c x px =c x *W img , cy px =c y *H img , w px =w*W img , h px =h*H img , where W img is the width of the image, H img is the length of the image, and the restored pixels (c x px , c y px ).
[0039] Then calculate the corner point data, (x min ,y min ) is the pixel coordinate of the upper left corner of the detection box, (x max ,y max ) are the pixel coordinates of the lower right corner of the detection box: x min =c x px -w px / 2,y min =c y px -h px / 2,x max =c x px +w px / 2,y max =c y px +h px / 2, and then you can execute the above corner point data processing method, which can share some code modules.
[0040] Another way is to use the following method: the center point coordinates (x center ,y center ) is calculated as follows: center =c x *W img ,y center =c y *H img .
[0041] 104. Calculate the coordinates of the N facial key points based on the center point coordinates to obtain N key point coordinates; In this embodiment, after obtaining the facial region and key points, coordinates of the 68 key points are calculated. The independently designed Geometric Calculations module is based on the key point geometric calculation module, the Classify classification network module, and the Detect object detection module. Any module can be selected for expression recognition based on requirements and device hardware performance. This solution primarily uses the Geometric Calculations module for key points.
[0042] 105. Calculate the horizontal relative weight and the vertical relative weight based on the N key point coordinates; In this embodiment, the horizontal and vertical relative weights are calculated for 68 key point coordinates. The horizontal and vertical relative weights are calculated for 72 key point coordinates. The calculation method of the horizontal and vertical relative weights can be modified according to the key point coordinates.
[0043] For details, please refer to Figure 3 , Figure 3 This is a specific embodiment of step 105 of the facial expression recognition method using an embedded device in an embodiment of the present invention. The N key point coordinates include: 68 key point coordinates in order. Step 105 includes the following specific implementation methods: 1051. Calculate the Euclidean distance between the coordinates of the first key point and the coordinates of the 13th key point to obtain the horizontal relative weight; 1052. Calculate the absolute difference between the ordinate of the 7th key point and the ordinate of the 42nd key point to obtain the vertical relative weight.
[0044] In steps 1051-1052, the horizontal relative weight may be calculated as follows: W f =||P0-P 12 ||,H f =|P 6,y -P 41,y |, where W f is the horizontal relative weight, P0 is the coordinate of the first key point, P 12 is the coordinate of the 13th key point, H f is the vertical relative weight, P 6,y is the ordinate of the 7th key point, P 41,y The vertical coordinate of the 42nd key point.
[0045] 106. Using the horizontal relative weight and the vertical relative weight, perform numerical analysis on the coordinates of the N key points to obtain a relative value of the mouth and a relative value of the eyes; In this embodiment, the horizontal relative weight and the vertical relative weight are calculated according to the designed geometric relationship to obtain the mouth relative value and the eye relative value corresponding to the coordinates of the 68 key points.
[0046] For details, please refer to Figure 4 , Figure 4 This is a specific embodiment of step 106 of the facial expression recognition method using an embedded device in an embodiment of the present invention, and step 106 includes the following specific implementation methods: 1061. Calculate the coordinates of the 61st key point and the 65th key point using the horizontal relative weight to obtain a mouth opening and closing value. 1062. Calculate the vertical coordinates of the 59th key point coordinate, the 63rd key point coordinate, and the 65th key point coordinate using the vertical relative weight to obtain a mouth corner depression value. 1063. Calculate the vertical coordinate of the 59th key point coordinate, the vertical coordinate of the 61st key point coordinate, and the vertical coordinate of the 63rd key point coordinate using the vertical relative weight to obtain a smile score; 1064. Calculate the coordinates of the 59th key point and the 63rd key point using the horizontal relative weight to obtain the relative mouth width. 1065. Based on a preset eye ratio algorithm, calculate the coordinates of the 37th key point to the 42nd key point to obtain the aspect ratio of the right eye. 1066. Based on a preset eye ratio algorithm, calculate the coordinates of the 43rd key point to the 48th key point to obtain the left eye aspect ratio; 1067. Calculate an average value based on the right eye aspect ratio and the left eye aspect ratio to obtain an average value of eye aspect ratio. 1068. Utilize the vertical relative weight to calculate the ordinate of the 25th key point coordinate, the ordinate of the 16th key point coordinate, the ordinate of the 42nd key point coordinate, and the ordinate of the 33rd key point coordinate to obtain the eyebrow and eye difference.
[0047] In steps 1061-1068, based on the data of the 68 key point coordinates and the data relationship between the horizontal relative weight and the vertical relative weight, the feature data corresponding to the relative value of the mouth and the relative value of the eyes are calculated. The following is a detailed calculation description: (1) Calculate the mouth opening and closing value N_m corresponding to yawning / opening the mouth: N_m=||P 60 -P 64 || / W f , where W f is the relative weight of the level, P 60 is the coordinate of the 61st key point, P64 The coordinates of the 65th key point.
[0048] (2) Calculate the mouth corner depression value N_m_d corresponding to the angry emotion: N_m_d=(max(P 58,y , P 62,y )-P 64,y ) / H f , H f is the vertical relative weight, P 58,y is the ordinate of the 59th key point, P 62,y is the ordinate of the 63rd key point, P 64,y The vertical coordinate of the 65th key point.
[0049] (3) Calculate the smile score S_s and relative mouth width m_w_n corresponding to the happy emotion: S_s=((P 60,y -P 58,y )+(P 60,y -P 62,y )) / H f =(2P 60,y -P 58,y -P 62,y ) / H f , where H f is the vertical relative weight, P 60,y is the ordinate of the 61st key point, P 60,y is the ordinate of the 59th key point, P 62,y The vertical coordinate of the 63rd key point.
[0050] m_w_n=||P 58 -P 62 || / W f , where W f is the relative weight of the level, P 58 is the coordinate of the 59th key point, P 62 The coordinates of the 63rd key point.
[0051] (4) Calculate the right eye aspect ratio EAR corresponding to the surprise emotion r , left eye aspect ratio EAR l , eyebrow-eye distance d b,e , the actual calculation method is as follows: 4.1 The formula for calculating the aspect ratio (EAR) of a single eye is as follows: EAR = (||e2−e6||+||e3−e5|) / (2||e1-e4||), where e1 and e4 are the points at the left and right corners of the eye and are used to calculate the horizontal width of the eye. e2, e3, e5, and e6 are the points on the upper and lower eyelids and are used to calculate the vertical height of the eye.
[0052] The coordinates of the 37th key point to the 42nd key point are input into the above formula according to the actual physical meaning to obtain the right eye aspect ratio EAR r The coordinates of the 43rd to 48th key points are input into the above formula according to the actual physical meaning to obtain the left eye aspect ratio EAR l Therefore, the actual average length and width of the eye is calculated as EAR = (EAR r +EAR l ) / 2.
[0053] 4.2 Distance between eyebrows and eyes d b,e The calculation formula is as follows: d b,e =(|P 24,y -P 15,y |+|P 41,y -P 32,y |) / 2H f , where H f is the vertical relative weight, P 24,y is the ordinate of the 25th key point, P 15,y is the ordinate of the 16th key point, P 41,y is the ordinate of the 42nd key point, P 32,y The vertical coordinate of the 33rd key point.
[0054] (5) The data corresponding to sleepiness is set as the average length and width of the eye, EAR.
[0055] (6) The data corresponding to the crying emotion are set as the average length and width of the eye EAR and the distance between the eyebrow and the eye d b,e .
[0056] (7) The blink value can be set to △EAR=|EAR r -EAR l |.
[0057] 107. Perform expression analysis on the mouth relative value and the eye relative value according to a preset expression determination algorithm to obtain a facial expression corresponding to the facial image.
[0058] In this embodiment, the data related to the relative value of the mouth and the relative value of the eyes can be combined into a feature vector as feature data and input into the expression judgment algorithm for classification and judgment. The expression judgment algorithm can be set as a classifier, such as MLP, SVM, GCN, etc., and the facial expression corresponding to the face image is obtained through the expression judgment algorithm.
[0059] Specifically, step 107 includes the following specific implementation methods: 1071. According to a preset numerical judgment table, matching and fitting processing is performed on the mouth opening and closing value, the mouth corner depression value, the smile score, the relative mouth width, the average length and width of the eye, and the eyebrow and eye difference to obtain the facial expression corresponding to the facial image.
[0060] In step 1071, the numerical judgment table is fitted with six data, including the mouth opening and closing value, the mouth corner depression value, the smile score, the relative mouth width, the average eye length and width, and the eyebrow difference, corresponding to the emotion discrimination types such as anger, happiness, surprise, sleepiness, and crying. Based on the emotion corresponding data, including the mouth corner depression value, the smile score, the relative mouth width, the average eye length and width, and the eyebrow difference, the numerical query is fitted in combination with the mouth opening and closing value corresponding to yawning / opening the mouth and the blinking value to obtain the facial expression corresponding to the face image. The fitted numerical judgment table can be generated using a classifier or traversed through neural network judgment, without the need for integrated training in the device, thereby improving the speed of classification processing and reducing related performance consumption. Devices with performance limitations such as embedded MIPS architecture can quickly achieve the results of facial expression recognition.
[0061] In an embodiment of the present invention, key point detection is performed using a facial key point detection model, key point geometric features are calculated, and these key point geometric features are used as recognition data. This data is then input into a pre-trained expression determination algorithm module to output the determined facial expression result. This invention achieves real-time face tracking and facial expression recognition with a smaller model, faster speed, and high accuracy, resolving the technical issue of current AI facial expression recognition algorithms operating poorly in embedded devices.
[0062] Figure 5This is a schematic diagram of the structure of a facial expression recognition device for an embedded device, provided in an embodiment of the present invention. This embedded device facial expression recognition device 500 may vary significantly depending on configuration or performance. It may include one or more processors (central processing units, CPUs) 510 (e.g., one or more processors), memory 520, and one or more storage media 530 (e.g., one or more mass storage devices) storing application programs 533 or data 532. The memory 520 and storage medium 530 may be either transient or persistent storage. The program stored in the storage medium 530 may include one or more modules (not shown), each of which may include a series of instructions for operating on the embedded device facial expression recognition device 500. Furthermore, the processor 510 may be configured to communicate with the storage medium 530 to execute the series of instructions stored in the storage medium 530 on the embedded device facial expression recognition device 500.
[0063] The facial expression recognition device 500 based on an embedded device may further include one or more power supplies 540, one or more wired or wireless network interfaces 550, one or more input and output interfaces 560, and / or one or more operating systems 531, such as Windows Server, Mac OS X, Unix, Linux, Free BSD, etc. It will be understood by those skilled in the art that Figure 5 The structure of the facial expression recognition device based on the application embedded device shown does not constitute a limitation on the facial expression recognition device based on the application embedded device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0064] The present invention also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores instructions, and when the instructions are run on a computer, the computer executes the steps of the facial expression recognition method using an embedded device.
[0065] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0066] In addition, although adopting specific order to describe each operation, this should be understood as requiring such operation to be carried out in the specific order shown or in sequential order, or requiring that all illustrated operations should be carried out to obtain desired results. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although comprising some specific implementation details in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of separate embodiment can also be implemented in a single implementation in combination. On the contrary, the various features described in the context of a single implementation also can be implemented in a plurality of implementations individually or in the mode of any suitable subcombination.
[0067] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A facial expression recognition method using an embedded device, characterized in that: Including steps: Get face image; Performing key point detection processing on the face image according to a preset key point detection model to obtain a face recognition frame and N facial key points corresponding to the face recognition frame; Obtaining center point coordinates based on the coordinate data of the face recognition frame; Calculating the coordinates of the N facial key points based on the center point coordinates to obtain N key point coordinates; Calculate the horizontal relative weight and the vertical relative weight based on the N key point coordinates; Using the horizontal relative weight and the vertical relative weight, numerically analyze the coordinates of the N key points to obtain a relative value of the mouth and a relative value of the eyes; According to a preset expression determination algorithm, expression analysis processing is performed on the mouth relative value and the eye relative value to obtain the facial expression corresponding to the facial image.
2. The facial expression recognition method using an embedded device according to claim 1, wherein: The N key point coordinates include: ordered 68 key point coordinates, and the step of calculating the horizontal relative weight and the vertical relative weight based on the N key point coordinates includes: Calculate the Euclidean distance between the coordinates of the first key point and the coordinates of the 13th key point to obtain the horizontal relative weight; Calculate the absolute difference between the ordinate of the 7th key point and the ordinate of the 42nd key point to obtain the vertical relative weight.
3. The facial expression recognition method using an embedded device according to claim 2, wherein: The step of performing numerical analysis on the coordinates of the N key points using the horizontal relative weight and the vertical relative weight to obtain the relative value of the mouth and the relative value of the eyes comprises: Using the horizontal relative weight, the coordinates of the 61st key point and the 65th key point are calculated to obtain the mouth opening and closing value; Using the vertical relative weight, the ordinate of the 59th key point coordinate, the ordinate of the 63rd key point coordinate, and the ordinate of the 65th key point coordinate are calculated to obtain the mouth corner depression value; Calculating the vertical relative weights of the 59th key point coordinate, the 61st key point coordinate, and the 63rd key point coordinate to obtain a smile score; Using the horizontal relative weight, the coordinates of the 59th key point and the 63rd key point are calculated to obtain the relative mouth width; Based on the preset eye ratio algorithm, calculate the coordinates of the 37th key point to the 42nd key point to obtain the aspect ratio of the right eye; Based on the preset eye ratio algorithm, calculate the coordinates of the 43rd key point to the 48th key point to obtain the left eye aspect ratio; Calculating an average according to the right eye aspect ratio and the left eye aspect ratio to obtain an average of the eye aspect ratio; By using the vertical relative weight, the ordinates of the 25th key point coordinate, the ordinates of the 16th key point coordinate, the ordinates of the 42nd key point coordinate, and the ordinates of the 33rd key point coordinate are calculated to obtain the eyebrow and eye difference.
4. The facial expression recognition method using an embedded device according to claim 3, wherein: The step of performing expression analysis on the mouth relative value and the eye relative value according to a preset expression determination algorithm to obtain the facial expression corresponding to the facial image comprises: According to a preset numerical judgment table, the mouth opening and closing value, the mouth corner depression value, the smile score, the relative mouth width, the average length and width of the eye, and the eyebrow and eye difference are matched and fitted to obtain the facial expression corresponding to the facial image.
5. The facial expression recognition method using an embedded device according to claim 1, wherein: The step of obtaining the center point coordinates based on the coordinate data of the face recognition frame includes: When the coordinate data of the face recognition frame is corner point data, the mean value of the corner point data is calculated according to a preset mean formula to obtain the center point coordinates; When the coordinate data of the face recognition frame is the width and height data of the frame center, the width and height data of the frame center are restored according to a preset ratio restoration algorithm to obtain the center point coordinates.
6. The facial expression recognition method using an embedded device according to claim 1, wherein: After the step of acquiring the facial image and before the step of performing key point detection processing on the facial image according to a preset key point detection model to obtain a facial recognition frame and N facial key points corresponding to the facial recognition frame, the method further includes: performing exposure enhancement processing on the face image according to a preset linear transformation algorithm to obtain an enhanced image; The enhanced image is subjected to noise reduction processing according to a preset bilateral filtering noise reduction algorithm to obtain a preprocessed face image.
7. The method for facial expression recognition using an embedded device according to claim 6, wherein: The step of performing exposure enhancement processing on the face image according to a preset linear transformation algorithm to obtain an enhanced image comprises: g(x,y)=a(f(x,y)-128)+128, Among them, g(x, y) is the enhanced image, a is the contrast parameter, f(x, y) is the face image, x is the horizontal pixel number, and y is the vertical pixel number.
8. The facial expression recognition method using an embedded device according to claim 6, wherein: The step of performing exposure enhancement processing on the face image according to a preset linear transformation algorithm to obtain an enhanced image comprises: The fastNlMeansDenoisingColored() function in the OpenCV library is used to perform exposure enhancement processing on the face image to obtain an enhanced image.
9. A facial expression recognition device using an embedded device, characterized in that: The facial expression recognition device using an embedded device includes: a memory and at least one processor, wherein instructions are stored in the memory, and the memory and the at least one processor are interconnected via a line; The at least one processor calls the instructions in the memory to enable the facial expression recognition device of the application embedded device to execute the facial expression recognition method of the application embedded device according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for facial expression recognition using an embedded device as described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Expression recognition method and system
CN111444860A
Facial expression recognition method and system based on video stream
CN113111789A
Method and device for driving expression of virtual character
CN114821734A
Methods and Systems to Modify a Two Dimensional Facial Image to Increase Dimensional Depth and Generate a Facial Image That Appears Three Dimensional
US20180158230A1