A method for emotion recognition based on eye movement analysis
By segmenting and feature extraction of eye images, texture, width and position centrifugal vectors are constructed, and combined with neural network models, the problem of low emotion recognition accuracy in the existing technology is solved, and higher precision emotion recognition is achieved.
Patent Information
- Application Number
- CN202510592066.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-05-09
AI Technical Summary
The prior art emotion recognition method based on the eyes has the problem of low emotional recognition accuracy, especially because each person's eyes have different degrees of opening and closing, the distance between feature points cannot truly reflect the movement of the eyes.
By extracting eye images from multiple consecutive moments, segmenting the iris region, sclera region and periophthalmic region, constructing texture centrifugal vector, width centrifugal vector and position centrifugal vector, and using the emotion recognition neural network model for emotion type recognition, combining the feature change factors of texture, width and position to improve the recognition accuracy.
It realizes a more comprehensive and accurate reflection of the actual movement of the eyes, improves the accuracy of emotional recognition, can capture the dynamic characteristics of emotional changes, and improves the accuracy of emotional recognition.
Smart Images

Figure CN120108028B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly relates to an emotion recognition method based on eye movement analysis. Background Art
[0002] As the window of the soul, the eyes carry a large amount of emotional information. The direction of eye movement, the duration of fixation, and the subtle changes in the eye muscles around the eyes are all closely related to the emotional state. Under normal circumstances, during daily communication and environmental observation, people's eyes will rotate naturally and flexibly to obtain surrounding visual information, with a relatively wide scanning range and a relatively uniform rotation speed. However, individuals in a bad mood or in a depressive state often show obvious slowness and limitation in eye movement. The amplitude of their eye movement in the horizontal and vertical directions decreases. For example, when reading text or viewing images, their eyes are difficult to move to the target position as quickly and smoothly as normal people, and may require more time and effort to complete the same visual search task.
[0003] Existing methods for emotion recognition through the eyes identify a person's emotion based on the distances between the feature points of the upper left eyelid, the lower left eyelid, the upper right eyelid, and the lower right eyelid. However, the degree of eye opening and closing varies from person to person, and the distances between the feature points cannot truly reflect the eye movement, resulting in the problem of low emotion recognition accuracy. Summary of the Invention
[0004] In view of the above deficiencies in the prior art, the emotion recognition method based on eye movement analysis provided by the present invention solves the problem of low emotion recognition accuracy in the prior art.
[0005] To achieve the above object of the invention, the technical solution adopted by the present invention is: an emotion recognition method based on eye movement analysis, comprising the following steps:
[0006] Extract eye images at consecutive multiple moments, segment each eye image to obtain the iris region, the sclera region, and the periorbital region;
[0007] Mark the combined region of the iris region and the sclera region in contact as the eye region;
[0008] Extract the texture value of each sub-region in the periorbital region, obtain the texture centrifuge value, and construct a texture centrifuge vector;
[0009] Obtain the width centrifuge value according to the width of the eye region at each moment, and construct a width centrifuge vector;
[0010] Calculate the position centrifuge value according to the distance between the eye region and the iris region, and construct a position centrifuge vector;
[0011] Obtain the feature change factor for each vector;
[0012] Input the texture centrifugal vector, width centrifugal vector, position centrifugal vector, and the feature change factors corresponding to each vector into the emotion recognition neural network model to obtain the emotion type.
[0013] Furthermore, the segmentation process includes:
[0014] Perform grayscale processing on each eye image, and cluster the pixel points according to the grayscale values to obtain multiple clusters;
[0015] Calculate the similarity between the average grayscale of each cluster and the stored iris grayscale value, and find the cluster corresponding to the maximum similarity as the iris region;
[0016] Calculate the similarity between the average grayscale of each cluster and the stored sclera grayscale value, and find the cluster corresponding to the maximum similarity as the sclera region;
[0017] Take the other clusters adjacent to the sclera region and the iris region as the periorbital region.
[0018] Furthermore, the process of constructing the texture centrifugal vector includes:
[0019] Extract the contour of the periorbital region to obtain a contour map;
[0020] Count the number of contour points in the contour map and perform normalization processing on the number to obtain the contour density of each periorbital region;
[0021] Calculate the standard deviation of each grayscale value in the periorbital region and perform normalization processing on the standard deviation to obtain the grayscale fluctuation value of each periorbital region;
[0022] Add the contour density and the grayscale fluctuation value belonging to the same periorbital region to obtain the texture value;
[0023] Subtract the average texture value from the texture value of each periorbital region to obtain the texture centrifugal value;
[0024] Construct the texture centrifugal values of the periorbital regions at each moment into a texture centrifugal vector.
[0025] Furthermore, the process of constructing the width centrifugal vector includes:
[0026] Construct a circumscribed rectangle for the eye region;
[0027] Extract the width of the circumscribed rectangle;
[0028] Subtract the average width from the width of each eye region to obtain the width centrifugal value;
[0029] Construct the width centrifugal values of the eye regions at each moment into a width centrifugal vector.
[0030] Further, the process of constructing the position centrifugal vector is as follows:
[0031] Obtain the eye center position for the eye region at each moment;
[0032] Obtain the iris center position for the iris region at each moment;
[0033] Calculate the distance between the iris center position and the eye center position to obtain the position centrifugal value of the iris;
[0034] Construct the position centrifugal values of the iris at each moment into a position centrifugal vector.
[0035] Further, the formula for the position centrifugal value of the iris is: , where γ is the position centrifugal value of the iris, x r is the abscissa of the iris center, y r is the ordinate of the iris center, x e is the abscissa of the eye center, y e is the ordinate of the eye center, f is the sign function, when x r -x e is greater than 0, assign 1 to f(x r -x e ), when x r -x e is less than 0, assign -1 to f(x r -x e ).
[0036] Further, the process of obtaining the feature change factor includes:
[0037] Sum the absolute values of the differences between adjacent elements in each vector to obtain the total feature change value;
[0038] Take the average of the total feature change value according to the number of groups of adjacent elements to obtain the average feature change value;
[0039] Normalize the average feature change value to obtain the feature change factor.
[0040] Further, the emotion recognition neural network model includes: a texture vector feature extraction module, a width vector feature extraction module, a position vector feature extraction module, a texture feature fusion module, a width feature fusion module, a position feature fusion module, and a fully connected layer;
[0041] The input end of the texture vector feature extraction module is used to input the texture centrifugal vector; the input end of the width vector feature extraction module is used to input the width centrifugal vector; the input end of the position vector feature extraction module is used to input the position centrifugal vector;
[0042] The first input end of the texture feature fusion module is connected to the output end of the texture vector feature extraction module, and its second input end is used to input the texture feature change factor;
[0043] The first input end of the width feature fusion module is connected to the output end of the width vector feature extraction module, and its second input end is used to input the width feature change factor;
[0044] The first input end of the position feature fusion module is connected to the output end of the position vector feature extraction module, and its second input end is used to input the position feature change factor;
[0045] The input end of the fully connected layer is respectively connected to the output ends of the texture feature fusion module, the width feature fusion module, and the position feature fusion module, and its output end serves as the output end of the emotion recognition neural network model.
[0046] Furthermore, the texture vector feature extraction module is used to extract texture features from the texture centrifugal vector; the width vector feature extraction module is used to extract width features from the width centrifugal vector; the position vector feature extraction module is used to extract position features from the position centrifugal vector;
[0047] The texture feature fusion module is used to fuse the texture features and the texture feature change factor to obtain the texture fusion feature; the width feature fusion module is used to fuse the width features and the width feature change factor to obtain the width fusion feature; the position feature fusion module is used to fuse the position features and the position feature change factor to obtain the position fusion feature;
[0048] The fully connected layer is used to classify according to the texture fusion feature, the width fusion feature, and the position fusion feature to obtain the emotion type.
[0049] Furthermore, the texture vector feature extraction module, the width vector feature extraction module, and the position vector feature extraction module all include an LSTM network, a two-dimensional feature construction layer, and a CNN network that are connected in sequence;
[0050] The LSTM network is used to extract shallow features from the vector; the two-dimensional data construction layer is used to construct the shallow features into two-dimensional features; the CNN network is used to extract deep features from the shallow features.
[0051] The beneficial effects of the present invention are:
[0052] 1. The present invention extracts eye images at consecutive multiple moments, segments them to obtain the iris region, sclera region, and periorbital region, analyzes different features of multiple regions (such as the texture value of the periorbital region, the width of the eye region, and the distance between the eye region and the iris region), and extracts the texture eccentricity value, width eccentricity value, and position eccentricity value, which can reflect the state of texture deviation from the mean, width deviation from the mean, and iris position deviation from the center at each moment, and can more comprehensively and accurately reflect the actual movement of the eyes, avoiding the limitations brought by relying only on the distance of eyelid feature points, thereby improving the accuracy of emotion recognition.
[0053] 2. The present invention obtains a feature change factor for each vector to reflect the change speed of the elements in the vector. The change of emotion is dynamic, which is not only reflected in the static numerical values of eye features but also in the change speed of these features over time.
[0054] 3. The present invention uses an emotion recognition neural network model to process the texture eccentricity vector, width eccentricity vector, and position eccentricity vector, and further improves the accuracy of emotion classification of the model by using the feature change factors corresponding to each vector. Description of the Drawings
[0055] Figure 1 is a flowchart of an emotion recognition method based on eye movement analysis;
[0056] Figure 2 is a schematic structural diagram of an emotion recognition neural network model. Detailed Embodiments
[0057] The following describes the detailed embodiments of the present invention to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the detailed embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.
[0058] As Figure 1 shown, an emotion recognition method based on eye movement analysis includes the following steps:
[0059] Extract eye images at consecutive multiple moments, segment each eye image to obtain the iris region, sclera region, and periorbital region;
[0060] Mark the combined region of the iris region and the sclera region in contact with the position as the eye region;
[0061] Extract the texture value of each sub-region in the periorbital region, obtain the texture eccentricity value, and construct a texture eccentricity vector;
[0062] Obtain the width centrifugal value according to the width of the eye region at each moment, and construct a width centrifugal vector;
[0063] Calculate the position centrifugal value according to the distance between the eye region and the iris region, and construct a position centrifugal vector;
[0064] Obtain the feature change factor for each vector;
[0065] Input the texture centrifugal vector, width centrifugal vector, position centrifugal vector, and the feature change factors corresponding to each vector into the emotion recognition neural network model to obtain the emotion type.
[0066] The iris region is located in the central part of the eye and is the colored circular part of the eye that controls the amount of light entering the eye. The sclera region: usually described as the "white of the eye" part, is the white area surrounding the iris, covering most of the surface of the eyeball and playing a role in protecting the internal tissues of the eyeball.
[0067] In this embodiment, 10 - 20 eye images can be collected every 30 seconds to facilitate observing the eye movements over a period of time.
[0068] In this embodiment, the segmentation process includes:
[0069] Perform grayscale processing on each eye image, and cluster the pixel points according to the grayscale values to obtain multiple clusters;
[0070] Calculate the similarity between the average grayscale of each cluster and the stored iris grayscale value respectively, and find the cluster corresponding to the maximum similarity as the iris region;
[0071] Calculate the similarity between the average grayscale of each cluster and the stored sclera grayscale value respectively, and find the cluster corresponding to the maximum similarity as the sclera region;
[0072] Take the other clusters adjacent to the sclera region and the iris region as the periorbital region.
[0073] In this embodiment, the formula for calculating the similarity between the average grayscale of each cluster and the stored iris grayscale value or the stored sclera grayscale value is: , where S is the similarity, G avg is the average grayscale of the cluster, G s is the stored iris grayscale value or the stored sclera grayscale value.
[0074] There are obvious color distinctions between the iris and the sclera. Therefore, after clustering the grayscale image in the present invention, the iris region and the sclera region are first found according to the similarity of the grayscale values of the clusters, and then the other clusters adjacent to the sclera region and the iris region are taken as the periorbital region.
[0075] In this embodiment, the process of constructing the texture centrifugal vector includes:
[0076] Extract the contour of the periorbital region to obtain a contour map;
[0077] Count the number of contour points in the contour map and perform normalization processing on the number to obtain the contour density of each periorbital region;
[0078] Calculate the standard deviation of each gray value in the periorbital region and perform normalization processing on the standard deviation to obtain the gray value fluctuation of each periorbital region;
[0079] Add the contour density and the gray value fluctuation belonging to the same periorbital region to obtain a texture value;
[0080] Subtract the average texture value from the texture value of each periorbital region to obtain a texture eccentricity value;
[0081] Construct the texture eccentricity vectors of the periorbital regions at each moment.
[0082] In the present invention, by extracting the periorbital contour and counting the number of contour points, the contour density can be obtained, which can reflect the shape and structural characteristics of the periorbital region. By calculating the standard deviation of the gray values, the gray value fluctuation can be obtained. Combining the contour and the gray value fluctuation can jointly characterize the texture of the periorbital region. The emotional changes of people will drive the activities of the periorbital muscles. For example, when happy, more wrinkles will be formed in the periorbital region, and both the contour density and the gray value fluctuation will increase. However, due to the different contours and wrinkle conditions of each periorbital region, in the present invention, the average texture value is subtracted from the texture value of each periorbital region to obtain the texture eccentricity value, which reflects the deviation of the texture. In this embodiment, the average texture value can be: the average value of the texture values of the periorbital regions at multiple moments, or the average value of the texture values of multiple daily periorbital regions.
[0083] In this embodiment, the specific process of extracting the contour of the periorbital region includes: taking each pixel point as the center, when the gray value of the pixel point at the center is the same as the gray values within the neighborhood range, the pixel point at the center is marked as a non-contour point, and the non-contour points are discarded. The remaining pixel points are contour points, and a contour map is obtained. When the gray value of the pixel point at the center is the same as the gray values within the neighborhood range, the pixel point at the center is marked as a non-contour point, and the non-contour points are discarded. The remaining pixel points are contour points, and a contour map is obtained.
[0084] In this embodiment, the expression for obtaining the contour density of each periorbital region is: , where μ o is the contour density of the periorbital region, M o is the number of contour points in the contour map, and M E is the number of pixel points in the periorbital region.
[0085] In this embodiment, the expression for obtaining the gray value fluctuation of each periorbital region is: , where θ is the gray value fluctuation of the periorbital region, σ is the standard deviation of each gray value in the periorbital region, and G maxis the maximum gray value in the periorbital region at each moment, G min is the minimum gray value in the periorbital region at each moment.
[0086] In this embodiment, the process of constructing the width centrifugal vector includes:
[0087] Construct a circumscribed rectangle for the eye region;
[0088] Extract the width of the circumscribed rectangle;
[0089] Subtract the average width from the width of each eye region to obtain the width centrifugal value;
[0090] Construct the width centrifugal values of the eye regions at each moment into a width centrifugal vector.
[0091] The present invention constructs a circumscribed rectangle for the eye region, and reflects the opening and closing state of the eyes through the width of the circumscribed rectangle. The larger the width of the circumscribed rectangle, the larger the iris region and the sclera region, indicating a greater degree of eye opening and closing. For example, when a person is surprised, the eyes will open wider, and the width of the circumscribed rectangle will increase; when a person is in a relaxed or tired state, the eyes will squint slightly, and the width will decrease.
[0092] Since the eye sizes of each person are different, the state of the eyes cannot be accurately reflected by the eye width. The present invention subtracts the average width from the width of each eye region to obtain the width centrifugal value, which reflects the deviation of the width from the average and more accurately reflects the state of the eyes.
[0093] In this embodiment, the average width can be: the average value of the widths of the eye regions at multiple moments, or the average value of the widths of the eye regions in multiple daily images.
[0094] In this embodiment, the process of constructing the position centrifugal vector is as follows:
[0095] Obtain the eye center position for the eye region at each moment;
[0096] Obtain the iris center position for the iris region at each moment;
[0097] Calculate the distance between the iris center position and the eye center position to obtain the position centrifugal value of the iris;
[0098] Construct the position centrifugal values of the iris at each moment into a position centrifugal vector.
[0099] In this embodiment, the formula for the position centrifugal value of the iris is: , where γ is the position centrifugal value of the iris, x r is the abscissa of the iris center, y r is the ordinate of the iris center, x eis the abscissa of the eye center, y e is the ordinate of the eye center, f is the sign function, at x r -x e when it is greater than 0, assign f(x r -x e ) the value of 1, at x r -x e when it is less than 0, assign f(x r -x e ) the value of -1.
[0100] The present invention calculates the distance between the eye center and the iris center, and can accurately quantify the position of the iris in the eye region. Adding the sign function f can not only reflect the distance value, but also reflect the direction of the iris relative to the eye center.
[0101] The abscissa of the eye center is the mean value of the abscissas of each pixel point in the eye region, and the ordinate of the eye center is the mean value of the ordinates of each pixel point in the eye region; another implementation is: the abscissa of the eye center is the mean value of the abscissas of each pixel point on the outer edge of the eye region, and the ordinate of the eye center is the mean value of the ordinates of each pixel point on the outer edge of the eye region, that is, the eye center position is the geometric center of the eye region. The abscissa of the iris region is the mean value of the abscissas of each pixel point in the iris region, and the ordinate of the iris region is the mean value of the ordinates of each pixel point in the iris region; the abscissa of the iris region is the mean value of the abscissas of each pixel point on the outer edge of the iris region, and the ordinate of the iris region is the mean value of the ordinates of each pixel point on the outer edge of the iris region, that is, the iris center position is the geometric center of the iris region. It is not limited to the methods of obtaining the eye center position and the iris center position described in this embodiment. The outer edge refers to the pixel points at the outermost boundary of the eye region or the iris region.
[0102] In this embodiment, the process of obtaining the feature change factor includes:
[0103] Sum the absolute values of the differences between adjacent elements in each vector to obtain the total feature change value;
[0104] According to the number of groups of adjacent elements, take the mean value of the total feature change value to obtain the feature change average value;
[0105] Normalize the feature change average value to obtain the feature change factor.
[0106] The expression for obtaining the feature change average value is: , E is the feature change average value, e t+1 is the element at the (t + 1)-th moment in the vector, e t is the element at the t-th moment in the vector, | | is the absolute value operation, T - 1 is the number of groups of adjacent elements, and t is the number of the moment.
[0107] The present invention sums the absolute values of the differences between adjacent elements to obtain the overall characteristic change situation, then takes the average value of the total characteristic change value according to the number of groups of adjacent elements to obtain the average characteristic change value, which reflects the average change speed, and then through normalization processing, the evaluation scale is unified.
[0108] In this embodiment, the process of normalizing the average characteristic change value includes: for the texture centrifugal vector, dividing the average characteristic change value by the average texture value; for the width centrifugal vector, dividing the average characteristic change value by the average width; for the position centrifugal vector, dividing the average characteristic change value by the absolute value of the difference between the maximum value and the minimum value in the position centrifugal vector.
[0109] As Figure 2 shown, the emotion recognition neural network model includes: a texture vector feature extraction module, a width vector feature extraction module, a position vector feature extraction module, a texture feature fusion module, a width feature fusion module, a position feature fusion module, and a fully connected layer;
[0110] The input end of the texture vector feature extraction module is used to input the texture centrifugal vector; the input end of the width vector feature extraction module is used to input the width centrifugal vector; the input end of the position vector feature extraction module is used to input the position centrifugal vector;
[0111] The first input end of the texture feature fusion module is connected to the output end of the texture vector feature extraction module, and its second input end is used to input the texture feature change factor;
[0112] The first input end of the width feature fusion module is connected to the output end of the width vector feature extraction module, and its second input end is used to input the width feature change factor;
[0113] The first input end of the position feature fusion module is connected to the output end of the position vector feature extraction module, and its second input end is used to input the position feature change factor;
[0114] The input end of the fully connected layer is respectively connected to the output ends of the texture feature fusion module, the width feature fusion module, and the position feature fusion module, and its output end serves as the output end of the emotion recognition neural network model.
[0115] The texture feature change factor is the feature change factor corresponding to the texture centrifugal vector, the width feature change factor is the feature change factor corresponding to the width centrifugal vector, and the position feature change factor is the feature change factor corresponding to the position centrifugal vector.
[0116] The present invention processes three time series vectors through three feature extraction modules respectively to extract vector features, and then uses three feature fusion modules to fuse the feature change factors with the vector features, so as to improve the accuracy of emotion recognition.
[0117] In this embodiment, the texture vector feature extraction module is used to extract texture features from the texture centrifugal vector; the width vector feature extraction module is used to extract width features from the width centrifugal vector; the position vector feature extraction module is used to extract position features from the position centrifugal vector.
[0118] The texture feature fusion module is used to fuse the texture features and the texture feature change factors to obtain texture fusion features; the width feature fusion module is used to fuse the width features and the width feature change factors to obtain width fusion features; the position feature fusion module is used to fuse the position features and the position feature change factors to obtain position fusion features.
[0119] The fully connected layer is used to classify according to the texture fusion features, width fusion features and position fusion features to obtain the emotion type.
[0120] In this embodiment, the expressions of the texture feature fusion module, width feature fusion module and position feature fusion module are: , where y is the output of the feature fusion module, g1 is the input of the first input end of the feature fusion module, g2 is the input of the second input end of the feature fusion module, w1 is the weight of g1, and w2 is the weight of g2.
[0121] In this embodiment, the texture vector feature extraction module, width vector feature extraction module and position vector feature extraction module all include, connected in sequence: an LSTM network and a fully connected layer. The LSTM network is used to extract shallow features from the vector, and the fully connected layer is used to perform feature mapping on the shallow features to obtain deep features. More preferably, the texture vector feature extraction module, width vector feature extraction module and position vector feature extraction module all include an LSTM network, a two-dimensional feature construction layer and a CNN network connected in sequence;
[0122] The LSTM network is used to extract shallow features from the vector; the two-dimensional data construction layer is used to construct the shallow features into two-dimensional features; the CNN network is used to extract deep features from the shallow features. For the texture vector feature extraction module, the deep feature is the texture feature; for the width vector feature extraction module, the deep feature is the width feature; for the position vector feature extraction module, the deep feature is the position feature.
[0123] The two-dimensional feature construction layer is: H = h T h, H is the two-dimensional feature, h is the vector composed of each feature value h output by the LSTM network t constituted, and T is the transpose operation.
[0124] The present invention can effectively extract the temporal information in the vector through the LSTM network, capture the dynamic changes of eye features over time, and then form two-dimensional data from the output features of the LSTM network, increasing the amount of data and facilitating the use of the powerful spatial feature extraction ability of the CNN network to explore the spatial relationship between features, enabling the model to comprehensively utilize temporal and spatial information, represent eye features more comprehensively, and improve the accuracy of emotion recognition.
[0125] In this embodiment, the emotion types include: happy, sad, angry, depressed, etc.
[0126] The present invention extracts eye images at consecutive multiple moments, segments them to obtain the iris region, sclera region, and periorbital region, analyzes different features of multiple regions (such as the texture value of the periorbital region, the width of the eye region, and the distance between the eye region and the iris region), and extracts the texture eccentricity value, width eccentricity value, and position eccentricity value, reflecting the state of texture deviation from the mean, width deviation from the mean, and iris position deviation from the center at each moment, being able to more comprehensively and accurately reflect the actual movement of the eyes, avoiding the limitations brought by only relying on the distance between eyelid feature points, and thus improving the accuracy of emotion recognition.
[0127] The present invention obtains a feature change factor for each vector to reflect the change speed of the elements in the vector. The change of emotion is dynamic, which is not only reflected in the static values of eye features but also in the change speed of these features over time.
[0128] The present invention uses an emotion recognition neural network model to process the texture eccentricity vector, width eccentricity vector, and position eccentricity vector, and uses the feature change factors corresponding to each vector to further improve the accuracy of model emotion classification.
[0129] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for emotion recognition based on eye movement analysis, characterized in that, It includes the following steps: Extract eye images at consecutive moments, segment each eye image to obtain the iris region, sclera region, and periorbital region; Mark the composed region of the iris region and sclera region in contact as the eye region; Extract texture values for each sub-region in the periorbital region, obtain the texture centrifugal value, and construct a texture centrifugal vector; According to the width of the eye region at each moment, obtain the width centrifugal value and construct a width centrifugal vector; According to the distance between the eye region and the iris region, calculate the position centrifugal value and construct a position centrifugal vector; Obtain the feature change factor for each vector; Input the texture centrifugal vector, width centrifugal vector, position centrifugal vector, and the feature change factors corresponding to each vector into the emotion recognition neural network model to obtain the emotion type.
2. The emotion recognition method based on eye analysis according to claim 1, wherein The segmentation process includes: Perform grayscale processing on each eye image, and perform pixel clustering according to the grayscale value to obtain multiple clusters; Calculate the similarity between the average grayscale of each cluster and the stored iris grayscale value respectively, and find the cluster corresponding to the maximum similarity as the iris region; Calculate the similarity between the average grayscale of each cluster and the stored sclera grayscale value respectively, and find the cluster corresponding to the maximum similarity as the sclera region; Take the other clusters adjacent to the sclera region and iris region as the periorbital region.
3. The emotion recognition method based on eye analysis according to claim 1, characterized in that The process of constructing the texture centrifugal vector includes: Extract the contour of the periorbital region to obtain a contour map; Count the number of contour points in the contour map and perform normalization processing on the number to obtain the contour density of each periorbital region; Calculate the standard deviation of each grayscale value in the periorbital region and perform normalization processing on the standard deviation to obtain the grayscale fluctuation value of each periorbital region; Add the contour density and grayscale fluctuation value belonging to the same periorbital region to obtain the texture value; Subtract the average texture value from the texture value of each periorbital region to obtain the texture centrifugal value; Construct the texture centrifugal values of the periorbital regions at each moment into a texture centrifugal vector.
4. The emotion recognition method based on eye analysis according to claim 1, characterized in that The process of constructing the width centrifugal vector includes: Construct a circumscribed rectangle for the eye region; Extract the width of the circumscribed rectangle; Subtract the average width from the width of each eye region to obtain the width centrifugal value; Construct the width centrifugal values of the eye regions at each moment into a width centrifugal vector.
5. The emotion recognition method based on eye analysis according to claim 1, characterized in that The process of constructing the position centrifugal vector is: Obtain the eye center position for the eye region at each moment; Obtain the iris center position for the iris region at each moment; Calculate the distance between the iris center position and the eye center position to obtain the position centrifugal value of the iris; Construct the position centrifugal values of the iris at each moment into a position centrifugal vector.
6. The emotion recognition method based on eye analysis according to claim 5, wherein The formula for the eccentricity value of the iris position is: , where γ is the eccentricity value of the iris position, x r is the abscissa of the iris center, y r is the ordinate of the iris center, x e is the abscissa of the eye center, y e is the ordinate of the eye center, and f is the sign function. When x r - x e is greater than 0, f(x r - x e ) is assigned the value 1. When x r - x e is less than 0, f(x r - x e ) is assigned the value -1.
7. The emotion recognition method based on eye analysis according to claim 1, wherein The process of obtaining the feature change factor includes: Sum the absolute values of the differences between adjacent elements in each vector to obtain the total feature change value; Take the average value of the total feature change value according to the number of groups of adjacent elements to obtain the average feature change value; Perform normalization processing on the average feature change value to obtain the feature change factor.
8. The emotion recognition method based on eye movement analysis according to claim 1, wherein The emotion recognition neural network model includes: a texture vector feature extraction module, a width vector feature extraction module, a position vector feature extraction module, a texture feature fusion module, a width feature fusion module, a position feature fusion module, and a fully connected layer; The input end of the texture vector feature extraction module is used to input the texture centrifugal vector; the input end of the width vector feature extraction module is used to input the width centrifugal vector; the input end of the position vector feature extraction module is used to input the position centrifugal vector; The first input end of the texture feature fusion module is connected to the output end of the texture vector feature extraction module, and its second input end is used to input the texture feature change factor; The first input end of the width feature fusion module is connected to the output end of the width vector feature extraction module, and its second input end is used to input the width feature change factor; The first input end of the position feature fusion module is connected to the output end of the position vector feature extraction module, and its second input end is used to input the position feature change factor; The input end of the fully connected layer is respectively connected to the output ends of the texture feature fusion module, the width feature fusion module and the position feature fusion module, and its output end serves as the output end of the emotion recognition neural network model.
9. The emotion recognition method based on eye analysis according to claim 8, characterized in that The texture vector feature extraction module is used to extract texture features from the texture centrifugal vector; the width vector feature extraction module is used to extract width features from the width centrifugal vector; The position vector feature extraction module is used to extract position features from the position centrifugal vector; The texture feature fusion module is used to fuse the texture features and the texture feature change factor to obtain the texture fusion feature; the width feature fusion module is used to fuse the width features and the width feature change factor to obtain the width fusion feature; The position feature fusion module is used to fuse the position features and the position feature change factor to obtain the position fusion feature; The fully connected layer is used to classify according to the texture fusion feature, the width fusion feature and the position fusion feature to obtain the emotion type.
10. The emotion recognition method based on eye analysis according to claim 8, wherein The texture vector feature extraction module, the width vector feature extraction module and the position vector feature extraction module all include an LSTM network, a two-dimensional feature construction layer and a CNN network connected in sequence; The LSTM network is used to extract shallow features from the vector; the two-dimensional data construction layer is used to construct the shallow features into two-dimensional features; the CNN network is used to extract deep features from the shallow features.
Citation Information
Patent Citations
Eye image segmentation method based on sclera region supervision
CN113343943A
Image acquisition system for off-axis eye images
US20200364441A1
Cited By
An end-to-end emotion recognition method and system based on seat cushion pressure sensor array
CN122624071A