Emotion recognition method based on eye expression analysis
By extracting and analyzing the iris, sclera and periophthalmic regions in the eye image, calculating texture, width and position centrifugal values, and combining feature change factors, inputting the emotion recognition neural network model, the problem of low emotion recognition accuracy in the existing technology is solved and higher accuracy is achieved.
Patent Information
- Application Number
- CN202510592066.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-09
AI Technical Summary
The existing methods of emotional recognition through the eyes rely on the distance of eyelid characteristic points, and there is a problem of low emotional recognition accuracy.
By extracting eye images from multiple consecutive moments, the iris region, sclera region and periophthalmic region were segmented, texture centrifugal values, width centrifugal values and position centrifugal values were calculated, and these vectors and their characteristic change factors were input into the emotion recognition neural network model for emotion recognition.
It improves the accuracy of emotion recognition, can reflect the actual movement of the eyes more comprehensively and accurately, and avoids the limitation of relying solely on the distance of eyelid characteristic points.
Smart Images

Figure CN120108028A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to an emotion recognition method based on eye expression analysis. Background Art
[0002] As the windows to the soul, eyes carry a lot of emotional information. The direction of eye movement, the duration of gaze, and the subtle changes in the muscles around the eyes are closely related to the emotional state. Under normal circumstances, people's eyes will naturally and flexibly move during daily communication and environmental observation to obtain visual information around them. The scanning range is relatively wide, and the movement speed is relatively uniform. However, individuals who are in a bad mood or depressed often show obvious slowness and limitation in eye movement. The horizontal and vertical movement amplitude of their eyes is reduced. For example, when reading text or viewing images, it is difficult for the eyes to move to the target position as quickly and smoothly as normal people, and it may take more time and effort to complete the same visual search task.
[0003] Existing methods for emotion recognition through eyes recognize human emotions based on the distances between the feature points of the upper left eyelid, the lower left eyelid, the upper right eyelid, and the lower right eyelid. However, each person's eyes open and close to different degrees, and the distance between the feature points cannot truly reflect the movement of the eyes, resulting in low accuracy in emotion recognition. Summary of the invention
[0004] In view of the above-mentioned deficiencies in the prior art, the present invention provides an emotion recognition method based on eye contact analysis to solve the problem of low emotion recognition accuracy in the prior art.
[0005] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is: an emotion recognition method based on eye expression analysis, comprising the following steps: Extract eye images at multiple consecutive moments, and segment each eye image to obtain the iris area, sclera area, and eye periorbital area; marking a region composed of the iris region and the sclera region where the position contacts as an eye region; Extracting texture values from each sub-region in the periorbital region, obtaining texture centrifugal values, and constructing texture centrifugal vectors; According to the width of the eye area at each moment, the width centrifugal value is obtained and the width centrifugal vector is constructed; According to the distance between the eye area and the iris area, the position centrifugal value is calculated and the position centrifugal vector is constructed; Get the feature change factor for each vector; The texture centrifugal vector, the width centrifugal vector, the position centrifugal vector, and the characteristic change factors corresponding to each vector are input into the emotion recognition neural network model to obtain the emotion type.
[0006] Furthermore, the segmentation process includes: The grayscale of each eye image is processed, and the pixels are clustered according to the grayscale value to obtain multiple clusters; Calculate the similarity between the average grayscale of each cluster and the stored iris grayscale value, and find the cluster corresponding to the maximum similarity as the iris region; Calculate the similarity between the average grayscale of each cluster and the stored sclera grayscale value, and find the cluster corresponding to the maximum similarity as the sclera area; The other clusters adjacent to the sclera region and the iris region are regarded as the periocular region.
[0007] Furthermore, the process of constructing the texture eccentric vector includes: Extract the contour of the peri-eye area to obtain a contour map; Count the number of contour points in the contour map and normalize the number to obtain the contour density of each periocular area; Calculate the standard deviation of each grayscale value in the periocular area, and normalize the standard deviation to obtain the grayscale fluctuation value of each periocular area; The contour density and grayscale fluctuation values belonging to the same periocular area are added together to obtain the texture value; The texture centrifugal value was obtained by subtracting the average texture value from the texture value of each periocular area; The texture centrifugal value of the eye area at each moment is constructed as a texture centrifugal vector.
[0008] Furthermore, the process of constructing the width centrifugal vector includes: Construct a bounding rectangle for the eye area; Extract the width of the bounding rectangle; Subtract the average width from the width of each eye region to obtain the width eccentricity value; The width eccentric value of the eye area at each moment is constructed as a width eccentric vector.
[0009] Furthermore, the process of constructing the position eccentric vector is: Get the eye center position for the eye area at each moment; Obtain the iris center position for the iris area at each moment; Calculate the distance between the center of the iris and the center of the eye to obtain the eccentricity value of the iris position; The position eccentricity value of the iris at each moment is constructed as a position eccentricity vector.
[0010] Furthermore, the formula for the eccentricity of the iris position is: , where γ is the eccentric value of the iris position, x r is the horizontal coordinate of the iris center, y ris the ordinate of the iris center, x e is the horizontal coordinate of the eye center, y e is the ordinate of the eye center, f is the sign function, at x r -x e When f(x r -x e ) is assigned a value of 1, and in x r -x e When f(x r -x e ) is assigned a value of -1.
[0011] Furthermore, the process of obtaining the characteristic variation factor includes: Sum the absolute values of the differences between adjacent elements in each vector to obtain the total feature change value; According to the number of groups of adjacent elements, the total feature change value is averaged to obtain the feature change average value; The feature change average is normalized to obtain the feature change factor.
[0012] Furthermore, the emotion recognition neural network model includes: a texture vector feature extraction module, a width vector feature extraction module, a position vector feature extraction module, a texture feature fusion module, a width feature fusion module, a position feature fusion module and a fully connected layer; The input end of the texture vector feature extraction module is used to input the texture centrifugal vector; the input end of the width vector feature extraction module is used to input the width centrifugal vector; the input end of the position vector feature extraction module is used to input the position centrifugal vector; The first input end of the texture feature fusion module is connected to the output end of the texture vector feature extraction module, and the second input end thereof is used to input a texture feature change factor; The first input end of the width feature fusion module is connected to the output end of the width vector feature extraction module, and the second input end thereof is used to input a width feature change factor; The first input end of the position feature fusion module is connected to the output end of the position vector feature extraction module, and the second input end thereof is used to input the position feature change factor; The input end of the fully connected layer is connected to the output end of the texture feature fusion module, the output end of the width feature fusion module and the output end of the position feature fusion module respectively, and its output end serves as the output end of the emotion recognition neural network model.
[0013] Furthermore, the texture vector feature extraction module is used to extract texture features from the texture centrifugal vector; the width vector feature extraction module is used to extract width features from the width centrifugal vector; the position vector feature extraction module is used to extract position features from the position centrifugal vector; The texture feature fusion module is used to fuse the texture feature and the texture feature change factor to obtain the texture fusion feature; the width feature fusion module is used to fuse the width feature and the width feature change factor to obtain the width fusion feature; the position feature fusion module is used to fuse the position feature and the position feature change factor to obtain the position fusion feature; The fully connected layer is used to classify the emotion type based on the texture fusion features, width fusion features and position fusion features.
[0014] Further, the texture vector feature extraction module, the width vector feature extraction module and the position vector feature extraction module each include an LSTM network, a two-dimensional feature construction layer and a CNN network connected in sequence; The LSTM network is used to extract shallow features from vectors; the two-dimensional data construction layer is used to construct shallow features into two-dimensional features; and the CNN network is used to extract deep features from shallow features.
[0015] The beneficial effects of the present invention are: 1. The present invention extracts eye images at multiple consecutive moments and segments them to obtain iris areas, sclera areas and periocular areas, analyzes different features of multiple areas (such as the texture value of the periocular area, the width of the eye area, and the distance between the eye area and the iris area), and extracts texture centrifugal values, width centrifugal values and position centrifugal values, reflecting the state of texture deviation from the mean, the state of width deviation from the mean, and the state of iris position deviation from the center at each moment, which can more comprehensively and accurately reflect the actual movement of the eyes, avoid the limitations brought by relying solely on the distance of eyelid feature points, and thus improve the accuracy of emotion recognition.
[0016] 2. The present invention obtains a characteristic change factor for each vector, reflecting the change speed of the elements in the vector. The change of emotions is dynamic, which is reflected not only in the static values of eye features, but also in the change speed of these features over time.
[0017] 3. The present invention adopts an emotion recognition neural network model to process texture centrifugal vectors, width centrifugal vectors and position centrifugal vectors, and uses the characteristic change factor corresponding to each vector to further improve the accuracy of the model emotion classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a flow chart of an emotion recognition method based on eye gaze analysis; Figure 2 This is a structural diagram of the emotion recognition neural network model. DETAILED DESCRIPTION
[0019] The specific implementation modes of the present invention are described below so that those skilled in the art can understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation modes. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations utilizing the concept of the present invention are protected.
[0020] like Figure 1 As shown, an emotion recognition method based on eye expression analysis comprises the following steps: Extract eye images at multiple consecutive moments, and segment each eye image to obtain the iris area, sclera area, and eye periorbital area; marking a region composed of the iris region and the sclera region where the position contacts as an eye region; Extracting texture values from each sub-region in the periorbital region, obtaining texture centrifugal values, and constructing texture centrifugal vectors; According to the width of the eye area at each moment, the width centrifugal value is obtained and the width centrifugal vector is constructed; According to the distance between the eye area and the iris area, the position centrifugal value is calculated and the position centrifugal vector is constructed; Get the feature change factor for each vector; The texture centrifugal vector, the width centrifugal vector, the position centrifugal vector, and the characteristic change factors corresponding to each vector are input into the emotion recognition neural network model to obtain the emotion type.
[0021] The iris is the colored, circular part of the eye located in the center of the eye that controls the amount of light that enters the eye. The sclera is often described as the "white" of the eye, which is the white area surrounding the iris and covers most of the surface of the eyeball, protecting the internal tissues of the eyeball.
[0022] In this embodiment, 10-20 eye images may be collected every 30 seconds to facilitate observation of eye movements over a period of time.
[0023] In this embodiment, the segmentation process includes: The grayscale of each eye image is processed, and the pixels are clustered according to the grayscale value to obtain multiple clusters; Calculate the similarity between the average grayscale of each cluster and the stored iris grayscale value, and find the cluster corresponding to the maximum similarity as the iris region; Calculate the similarity between the average grayscale of each cluster and the stored sclera grayscale value, and find the cluster corresponding to the maximum similarity as the sclera area; The other clusters adjacent to the sclera region and the iris region are regarded as the periocular region.
[0024] In this embodiment, the formula for calculating the similarity between the average grayscale of each cluster and the stored iris grayscale value or the stored sclera grayscale value is: , where S is the similarity, G avg is the average grayscale of the cluster, G s To store the iris grayscale value or to store the sclera grayscale value.
[0025] There is an obvious color distinction between the iris and the sclera. Therefore, after clustering the grayscale image, the present invention first finds the iris region and the sclera region according to the similarity of the grayscale values of the clusters, and then uses other clusters adjacent to the sclera region and the iris region as the periocular region.
[0026] In this embodiment, the process of constructing the texture eccentric vector includes: Extract the contour of the peri-eye area to obtain a contour map; Count the number of contour points in the contour map and normalize the number to obtain the contour density of each periocular area; Calculate the standard deviation of each grayscale value in the periocular area, and normalize the standard deviation to obtain the grayscale fluctuation value of each periocular area; The contour density and grayscale fluctuation values belonging to the same periocular area are added together to obtain the texture value; The texture centrifugal value was obtained by subtracting the average texture value from the texture value of each periocular area; The texture centrifugal value of the eye area at each moment is constructed as a texture centrifugal vector.
[0027] The present invention extracts the contour of the eye and counts the number of contour points to obtain the contour density, which can reflect the shape and structural characteristics of the eye. The grayscale value standard deviation is calculated to obtain the grayscale fluctuation value. The contour and grayscale fluctuation value are combined to jointly characterize the texture of the eye. Changes in a person's emotions will drive the activity of the muscles around the eye. For example, when you are happy, more wrinkles will form in the area around the eye, and the contour density and grayscale fluctuation value will increase. However, since the contours of each area around the eye are different and the wrinkles are different, the present invention uses the texture value of each area around the eye to subtract the average texture value to obtain the texture centrifugal value, which reflects the deviation of the texture. In this embodiment, the average texture value can be: the average value of the texture values of the area around the eye at multiple times, or it can be the average value of the texture values of multiple photos of the area around the eye in daily life.
[0028] In this embodiment, the specific process of extracting the contour of the eye area includes: taking each pixel point as the center, comparing the grayscale value of the pixel point at the center with the grayscale value of the pixel point at the center. When the grayscale values within the neighborhood are the same, the pixel point at the center is marked as a non-contour point, the non-contour point is discarded, and the remaining pixel points are contour points to obtain a contour map.
[0029] In this embodiment, the expression for obtaining the contour density of each periorbital area is: , where μ o is the contour density of the eye area, M o is the number of contour points in the contour map, M E is the number of pixels in the periocular area.
[0030] In this embodiment, the expression for obtaining the grayscale fluctuation value of each periocular area is: , where θ is the grayscale fluctuation value of the periocular area, σ is the standard deviation of each grayscale value in the periocular area, G max is the maximum grayscale value in the eye area at each moment, G min is the minimum grayscale value in the periocular area at each moment.
[0031] In this embodiment, the process of constructing the width centrifugal vector includes: Construct a bounding rectangle for the eye area; Extract the width of the bounding rectangle; Subtract the average width from the width of each eye region to obtain the width eccentricity value; The width eccentric value of the eye area at each moment is constructed as a width eccentric vector.
[0032] The present invention constructs a circumscribed rectangle for the eye area, and the width of the circumscribed rectangle reflects the opening and closing state of the eyes. The larger the width of the circumscribed rectangle, the larger the iris area and the sclera area, indicating that the degree of eye opening and closing is greater. For example, when a person is surprised, the eyes will open wider and the width of the circumscribed rectangle will increase; when a person is in a relaxed or tired state, the eyes will squint slightly and the width will decrease.
[0033] Since each person's eyes are of different sizes, the eye width cannot accurately reflect the state of the eyes. The present invention obtains the width eccentricity value by subtracting the average width from the width of each eye area, which reflects the deviation of the width from the average and more accurately reflects the state of the eyes.
[0034] In this embodiment, the average width may be: an average value of the widths of the eye regions at multiple times, or an average value of the widths of the eye regions in multiple daily photos.
[0035] In this embodiment, the process of constructing the position centrifugal vector is: Get the eye center position for the eye area at each moment; Obtain the iris center position for the iris area at each moment; Calculate the distance between the center of the iris and the center of the eye to obtain the eccentricity value of the iris position; The position eccentricity value of the iris at each moment is constructed as a position eccentricity vector.
[0036] In this embodiment, the formula for the eccentricity value of the iris position is: , where γ is the eccentric value of the iris position, x r is the horizontal coordinate of the iris center, y r is the ordinate of the iris center, x e is the horizontal coordinate of the eye center, y e is the ordinate of the eye center, f is the sign function, at x r -x e When f(x r -x e ) is assigned a value of 1, and in x r -x e When f(x r -x e ) is assigned a value of -1.
[0037] The present invention calculates the distance between the eye center and the iris center, and can accurately quantify the position of the iris in the eye area. Adding the sign function f can not only reflect the distance value, but also reflect the direction of the iris relative to the eye center.
[0038] The horizontal coordinate of the eye center is the average of the horizontal coordinates of each pixel point in the eye area, and the vertical coordinate of the eye center is the average of the vertical coordinates of each pixel point in the eye area; another implementation method is: the horizontal coordinate of the eye center is the average of the horizontal coordinates of each pixel point at the outer edge of the eye area, and the vertical coordinate of the eye center is the average of the vertical coordinates of each pixel point at the outer edge of the eye area, that is, the eye center position is the geometric center of the eye area. The horizontal coordinate of the iris area is the average of the horizontal coordinates of each pixel point in the iris area, and the vertical coordinate of the iris area is the average of the vertical coordinates of each pixel point in the iris area; the horizontal coordinate of the iris area is the average of the horizontal coordinates of each pixel point at the outer edge of the iris area, and the vertical coordinate of the iris area is the average of the vertical coordinates of each pixel point at the outer edge of the iris area, that is, the iris center position is the geometric center of the iris area. The method of obtaining the eye center position and the iris center position described in this embodiment is not limited. The outer edge refers to the pixel point at the outermost boundary of the eye area or the iris area.
[0039] In this embodiment, the process of obtaining the characteristic change factor includes: Sum the absolute values of the differences between adjacent elements in each vector to obtain the total feature change value; According to the number of groups of adjacent elements, the total feature change value is averaged to obtain the feature change average value; The feature change average is normalized to obtain the feature change factor.
[0040] The expression for getting the average value of feature change is: , E is the average value of feature change, e t+1is the element in the vector at time t+1, e t is the element in the vector at the tth moment, || is the absolute value operation, T-1 is the number of groups of adjacent elements, and t is the moment number.
[0041] The present invention sums the absolute values of the differences between adjacent elements to obtain the overall feature change situation, and then takes the average of the total feature change values according to the number of groups of adjacent elements to obtain the feature change average value, which reflects the average change speed, and then unifies the evaluation scale through normalization processing.
[0042] In this embodiment, the process of normalizing the feature change average value includes: for the texture centrifugal vector, the feature change average value is divided by the average texture value; for the width centrifugal vector, the feature change average value is divided by the average width; for the position centrifugal vector, the feature change average value is divided by the absolute value of the difference between the maximum value and the minimum value in the position centrifugal vector.
[0043] like Figure 2 As shown, the emotion recognition neural network model includes: a texture vector feature extraction module, a width vector feature extraction module, a position vector feature extraction module, a texture feature fusion module, a width feature fusion module, a position feature fusion module and a fully connected layer; The input end of the texture vector feature extraction module is used to input the texture centrifugal vector; the input end of the width vector feature extraction module is used to input the width centrifugal vector; the input end of the position vector feature extraction module is used to input the position centrifugal vector; The first input end of the texture feature fusion module is connected to the output end of the texture vector feature extraction module, and the second input end thereof is used to input a texture feature change factor; The first input end of the width feature fusion module is connected to the output end of the width vector feature extraction module, and the second input end thereof is used to input a width feature change factor; The first input end of the position feature fusion module is connected to the output end of the position vector feature extraction module, and the second input end thereof is used to input the position feature change factor; The input end of the fully connected layer is connected to the output end of the texture feature fusion module, the output end of the width feature fusion module and the output end of the position feature fusion module respectively, and its output end serves as the output end of the emotion recognition neural network model.
[0044] The texture feature change factor is the feature change factor corresponding to the texture centrifugal vector, the width feature change factor is the feature change factor corresponding to the width centrifugal vector, and the position feature change factor is the feature change factor corresponding to the position centrifugal vector.
[0045] The present invention processes three time series vectors respectively through three feature extraction modules to extract vector features, and then uses three feature fusion modules to fuse feature change factors with vector features to improve the accuracy of emotion recognition.
[0046] In this embodiment, the texture vector feature extraction module is used to extract texture features from the texture centrifugal vector; the width vector feature extraction module is used to extract width features from the width centrifugal vector; the position vector feature extraction module is used to extract position features from the position centrifugal vector; The texture feature fusion module is used to fuse the texture feature and the texture feature change factor to obtain the texture fusion feature; the width feature fusion module is used to fuse the width feature and the width feature change factor to obtain the width fusion feature; the position feature fusion module is used to fuse the position feature and the position feature change factor to obtain the position fusion feature; The fully connected layer is used to classify the emotion type based on the texture fusion features, width fusion features and position fusion features.
[0047] In this embodiment, the expressions of the texture feature fusion module, the width feature fusion module and the position feature fusion module are: , where y is the output of the feature fusion module, g 1 is the input of the first input terminal of the feature fusion module, g 2 is the input of the second input terminal of the feature fusion module, w 1 g 1 The weight, w 2 g 2 The weight of .
[0048] In this embodiment, the texture vector feature extraction module, the width vector feature extraction module and the position vector feature extraction module all include: an LSTM network and a fully connected layer connected in sequence, the LSTM network is used to extract shallow features from the vector, and the fully connected layer is used to perform feature mapping on the shallow features to obtain deep features. More preferably, the texture vector feature extraction module, the width vector feature extraction module and the position vector feature extraction module all include an LSTM network, a two-dimensional feature construction layer and a CNN network connected in sequence; The LSTM network is used to extract shallow features from vectors; the two-dimensional data construction layer is used to construct shallow features into two-dimensional features; the CNN network is used to extract deep features from shallow features. For the texture vector feature extraction module, the deep features are texture features; for the width vector feature extraction module, the deep features are width features; for the position vector feature extraction module, the deep features are position features.
[0049] The two-dimensional feature construction layer is: H=h T h, H are two-dimensional features, and h is the eigenvalues output by the LSTM network.t The vector formed by , T is the transpose operation.
[0050] The present invention can effectively extract the time series information in the vector through the LSTM network, capture the dynamic changes of eye features over time, and then construct the output features of the LSTM network into two-dimensional data, thereby increasing the data volume and facilitating the use of the powerful spatial feature extraction capability of the CNN network to mine the spatial relationship between features, so that the model can comprehensively utilize the time series and spatial information, more comprehensively represent the eye features, and improve the accuracy of emotion recognition.
[0051] In this embodiment, the emotion types include: happiness, sadness, anger, depression, etc.
[0052] The present invention extracts eye images at multiple consecutive moments and segments them to obtain iris areas, sclera areas and periocular areas, analyzes different features of multiple areas (such as the texture value of the periocular area, the width of the eye area, and the distance between the eye area and the iris area), and extracts texture centrifugal values, width centrifugal values and position centrifugal values, reflecting the state of texture deviation from the mean, the state of width deviation from the mean, and the state of iris position deviation from the center at each moment, which can more comprehensively and accurately reflect the actual movement of the eyes, avoid the limitations brought by relying solely on the distance of eyelid feature points, and thus improve the accuracy of emotion recognition.
[0053] The present invention obtains a characteristic change factor for each vector, reflecting the change speed of the elements in the vector. The change of emotions is dynamic, which is reflected not only in the static values of eye features, but also in the change speed of these features over time.
[0054] The present invention adopts an emotion recognition neural network model to process texture centrifugal vectors, width centrifugal vectors and position centrifugal vectors, and adopts a characteristic change factor corresponding to each vector to further improve the accuracy of the model emotion classification.
[0055] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An emotion recognition method based on eye contact analysis, characterized in that: The following steps are involved: Extract eye images at multiple consecutive moments, and segment each eye image to obtain the iris area, sclera area, and eye periorbital area; marking a region composed of the iris region and the sclera region where the position contacts as an eye region; Extracting texture values from each sub-region in the periorbital region, obtaining texture centrifugal values, and constructing texture centrifugal vectors; According to the width of the eye area at each moment, the width centrifugal value is obtained and the width centrifugal vector is constructed; According to the distance between the eye area and the iris area, the position centrifugal value is calculated and the position centrifugal vector is constructed; Get the feature change factor for each vector; The texture centrifugal vector, the width centrifugal vector, the position centrifugal vector, and the characteristic change factors corresponding to each vector are input into the emotion recognition neural network model to obtain the emotion type.
2. The emotion recognition method based on eye expression analysis according to claim 1, characterized in that: The segmentation process includes: The grayscale of each eye image is processed, and the pixels are clustered according to the grayscale value to obtain multiple clusters; Calculate the similarity between the average grayscale of each cluster and the stored iris grayscale value, and find the cluster corresponding to the maximum similarity as the iris region; Calculate the similarity between the average grayscale of each cluster and the stored sclera grayscale value, and find the cluster corresponding to the maximum similarity as the sclera area; The other clusters adjacent to the sclera region and the iris region are regarded as the periocular region.
3. The emotion recognition method based on eye expression analysis according to claim 1, characterized in that: The process of constructing the texture's off-center vector is as follows: Extract the contour of the peri-eye area to obtain a contour map; Count the number of contour points in the contour map and normalize the number to obtain the contour density of each periocular area; Calculate the standard deviation of each grayscale value in the periocular area, and normalize the standard deviation to obtain the grayscale fluctuation value of each periocular area; The contour density and grayscale fluctuation values belonging to the same periocular area are added together to obtain the texture value; The texture centrifugal value was obtained by subtracting the average texture value from the texture value of each periocular area; The texture centrifugal value of the eye area at each moment is constructed as a texture centrifugal vector.
4. The emotion recognition method based on eye expression analysis according to claim 1, characterized in that: The process of constructing the width eccentric vector includes: Construct a bounding rectangle for the eye area; Extract the width of the bounding rectangle; Subtract the average width from the width of each eye region to obtain the width eccentricity value; The width eccentric value of the eye area at each moment is constructed as a width eccentric vector.
5. The emotion recognition method based on eye expression analysis according to claim 1, characterized in that: The process of constructing the position eccentric vector is: Get the eye center position for the eye area at each moment; Obtain the iris center position for the iris area at each moment; Calculate the distance between the center of the iris and the center of the eye to obtain the eccentricity value of the iris position; The position eccentricity value of the iris at each moment is constructed as a position eccentricity vector.
6. The emotion recognition method based on eye expression analysis according to claim 5, characterized in that: The formula for the eccentricity of the iris position is: , where γ is the eccentric value of the iris position, x r is the horizontal coordinate of the iris center, y r is the ordinate of the iris center, x e is the horizontal coordinate of the eye center, y e is the ordinate of the eye center, f is the sign function, at x r -x e When f(x r -x e ) is assigned a value of 1, and in x r -x e When f(x r -x e ) is assigned a value of -1.
7. The emotion recognition method based on eye expression analysis according to claim 1, characterized in that: The process of obtaining the characteristic variation factor includes: Sum the absolute values of the differences between adjacent elements in each vector to obtain the total feature change value; According to the number of groups of adjacent elements, the total feature change value is averaged to obtain the feature change average value; The feature change average is normalized to obtain the feature change factor.
8. The emotion recognition method based on eye expression analysis according to claim 1, characterized in that: The emotion recognition neural network model includes: a texture vector feature extraction module, a width vector feature extraction module, a position vector feature extraction module, a texture feature fusion module, a width feature fusion module, a position feature fusion module and a fully connected layer; The input end of the texture vector feature extraction module is used to input the texture centrifugal vector; the input end of the width vector feature extraction module is used to input the width centrifugal vector; the input end of the position vector feature extraction module is used to input the position centrifugal vector; The first input end of the texture feature fusion module is connected to the output end of the texture vector feature extraction module, and the second input end thereof is used to input a texture feature change factor; The first input end of the width feature fusion module is connected to the output end of the width vector feature extraction module, and the second input end thereof is used to input a width feature change factor; The first input end of the position feature fusion module is connected to the output end of the position vector feature extraction module, and the second input end thereof is used to input the position feature change factor; The input end of the fully connected layer is connected to the output end of the texture feature fusion module, the output end of the width feature fusion module and the output end of the position feature fusion module respectively, and its output end serves as the output end of the emotion recognition neural network model.
9. The emotion recognition method based on eye expression analysis according to claim 8, characterized in that: The texture vector feature extraction module is used to extract texture features from the texture centrifugal vector; the width vector feature extraction module is used to extract width features from the width centrifugal vector; The position vector feature extraction module is used to extract position features from the position centrifugal vector; The texture feature fusion module is used to fuse the texture feature and the texture feature change factor to obtain the texture fusion feature; the width feature fusion module is used to fuse the width feature and the width feature change factor to obtain the width fusion feature; The position feature fusion module is used to fuse the position feature and the position feature change factor to obtain the position fusion feature; The fully connected layer is used to classify the emotion type based on the texture fusion features, width fusion features and position fusion features.
10. The emotion recognition method based on eye expression analysis according to claim 8, characterized in that: The texture vector feature extraction module, the width vector feature extraction module and the position vector feature extraction module all include an LSTM network, a two-dimensional feature construction layer and a CNN network connected in sequence; The LSTM network is used to extract shallow features from vectors; the two-dimensional data construction layer is used to construct shallow features into two-dimensional features; and the CNN network is used to extract deep features from shallow features.
Citation Information
Patent Citations
Identity recognition model training method, testing method, recognition method and device
CN112163456A
Eye image segmentation method based on sclera region supervision
CN113343943A
Image acquisition system for off-axis eye images
US20200364441A1