Facial expression recognition method, system, device and storage medium

By acquiring images of multiple facial regions and determining the effective expression time periods, combined with optical navigation sensors and classification models, the problem of low accuracy and recognition rate in existing expression recognition technologies has been solved, achieving a more efficient expression recognition effect.

CN115393939BActive Publication Date: 2025-10-24GOERTEK INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211049807.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-30
Publication Date
2025-10-24
Estimated Expiration
2042-08-30

AI Technical Summary

Technical Problem

Existing facial expression recognition methods have low accuracy and recognition rate, and it is difficult to effectively utilize information from multiple facial regions for accurate expression recognition.

Method used

By acquiring images of multiple facial regions, the effective time period for facial expressions is determined. Then, by using an optical navigation sensor to acquire images of facial regions and combining them with a classification model to determine facial expression labels, the accuracy and recognition rate of facial expression recognition are improved.

Benefits of technology

By acquiring images of multiple facial regions and determining the effective expression time periods, the computational load is reduced, improving the accuracy and recognition rate of expression recognition, reducing data processing volume, and increasing the efficiency of expression recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115393939B_ABST
    Figure CN115393939B_ABST
Patent Text Reader

Abstract

The present application relates to expression recognition method, system, equipment and storage medium. Expression recognition method includes: according to the preset sampling frequency, simultaneously carry out image acquisition to multiple face regions, obtain the first image of each face region;Wherein, multiple face regions belong to different face region groups, and each face region group includes at least one face region;The effective expression time period is determined according to the first image collected in continuous multiple sampling periods;Wherein, the effective expression time period includes continuous multiple sub-time periods, and the number of sampling periods corresponding to each sub-time period is same;According to the first image of face region group in each sub-time period, the expression label of face region group in each sub-time period is determined;According to the expression label of multiple face region groups in multiple sub-time periods in the effective expression time period, the expression of the effective expression time period is determined.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present disclosure relate to the field of face recognition, and more particularly, to an expression recognition method, system, device and storage medium. BACKGROUND

[0002] With the continuous development of expression recognition research on facial images, the current mainstream research direction is to input the collected facial images into a convolutional neural network, deeply mine and extract the features of expressions in the images, and train the network to learn the classification of expressions. However, there are still many deficiencies in facial expression recognition using such methods, and the recognition accuracy and recognition rate are not high. SUMMARY

[0003] One object of embodiments of the present disclosure is to provide a new technical solution for an expression recognition method.

[0004] According to a first aspect of the present disclosure, an expression recognition method is provided, the expression recognition method comprising: collecting images of a plurality of facial regions simultaneously according to a preset sampling frequency to obtain first images of each facial region; wherein the plurality of facial regions belong to different facial region groups, and each facial region group includes at least one facial region; determining an effective expression time period according to the first images collected in a plurality of consecutive sampling periods; wherein the effective expression time period includes a plurality of consecutive sub-time periods, and the number of sampling periods corresponding to each sub-time period is the same; determining an expression label of the facial region group in each sub-time period according to the first images of the facial region group in each sub-time period; and determining an expression of the effective expression time period according to the expression labels of the plurality of facial region groups in the plurality of sub-time periods in the effective expression time period.

[0005] Optionally, the first images are collected based on an optical navigation sensor.

[0006] Optionally, the facial region group is selected as a first facial region and a second facial region that are left-right symmetrical; and / or the facial region group is selected as a plurality of facial regions belonging to the same part of the face.

[0007] Optionally, determining the effective expression time period according to the first images collected in the plurality of consecutive sampling periods comprises: calculating a change amount of the first images of the facial region compared to a preset image of the facial region, and taking the change amount as a reference value of the first image; wherein the preset image is an image collected in a relaxed state of the face; and determining an expression start time and an expression end time according to the reference values of the first images collected in the plurality of consecutive sampling periods, and taking a time period from the expression start time to the expression end time as the effective expression time period.

[0008] Optionally, the expression label of the face region group in each sub time period is determined according to the first image of the face region group in each sub time period, comprising: splicing the first image of the face region group in the same sampling period in the sub time period into a second image of the face region group in the sub time period; splicing the second image of the face region group in the sub time period into a third image of the face region group in the sub time period according to the order of the sampling period; inputting the third image of the face region group in the sub time period into the pre-trained classification model, and outputting the expression label of the face region group in the sub time period through the classification model.

[0009] Optionally, the expression of the effective expression time period is determined according to the expression labels of the plurality of face region groups in the plurality of sub time periods in the effective expression time period, comprising: if the expression labels of the plurality of face region groups in the target sub time period are the same, setting the expression determination result of the target sub time period as valid and taking the expression labels of the plurality of face region groups in the target sub time period as the expression label of the target sub time period;

[0010] If the expression labels of the plurality of face region groups in the target sub time period are not the same, the expression determination result of the target sub time period is set as invalid, and the determination of the expression label of the target sub time period is skipped, and the expression labels of the plurality of face region groups in the next target sub time period are determined; wherein the target sub time period is any sub time period in the effective expression time period.

[0011] Optionally, the expression of the effective expression time period is determined according to the expression labels of the plurality of face region groups in the plurality of sub time periods in the effective expression time period, and further comprising:

[0012] If the expression determination results of at least N sub time periods in the continuous M sub time periods in the effective expression time period are valid and the expression labels of the N sub time periods are the same, the expression labels of the N sub time periods are taken as the expression of the effective expression time period;

[0013] If the expression determination results of less than N sub time periods in the continuous M sub time periods in the effective expression time period are valid, or the expression determination results of at least N sub time periods in the continuous M sub time periods are valid, but the expression labels of the N sub time periods are not the same, the expression labels of the continuous M sub time periods are re-determined from the last sub time period with consistent expression labels; wherein M and N are positive integers, and N is less than or equal to M.

[0014] According to a second aspect of the present disclosure, an expression recognition system is also provided, the expression recognition system comprising: a face image sampling module configured to simultaneously collect images of a plurality of face regions at a preset sampling frequency to obtain first images of each face region; wherein the plurality of face regions belong to different face region groups, and each face region group comprises at least one face region;

[0015] an effective expression determination module configured to determine an effective expression time period according to the first images collected in a plurality of consecutive sampling periods; wherein the effective expression time period comprises a plurality of consecutive sub-time periods, and each sub-time period corresponds to the same number of sampling periods;

[0016] a first expression determination module configured to determine an expression label of each face region group in each sub-time period according to the first images of the face region group in each sub-time period;

[0017] a second expression determination module configured to determine an expression of the effective expression time period according to the expression labels of the plurality of face region groups in the plurality of sub-time periods in the effective expression time period.

[0018] According to a third aspect of the present disclosure, a head-mounted electronic device is also provided, which is applied to a face, the face comprising a plurality of face regions, the plurality of face regions belonging to different face region groups, and each face region group comprising at least one face region, the head-mounted electronic device comprising: a shell, the shell being capable of being fixed to the face, and an accommodation space being arranged in the shell; a face-adhesion assembly, the face-adhesion assembly being arranged in the shell, and the face-adhesion assembly being capable of being adhered to the skin of the face; a plurality of optical navigation sensors, the plurality of optical navigation sensors being arranged in the face-adhesion assembly, the plurality of optical navigation sensors being arranged correspondingly to the plurality of face regions, the plurality of optical navigation sensors being closely adhered to the skin of the face regions, and the plurality of optical navigation sensors being configured to collect images of each face region; wherein the images comprise movement direction and distance information of the skin in each face region; and a digital signal processor, the digital signal processor being arranged in the accommodation space, and the digital signal processor being configured to process the images collected by the plurality of optical navigation sensors according to the expression recognition method of any one of the first aspect, and output an expression of the face.

[0019] According to a fourth aspect of the present disclosure, a terminal device is also provided, comprising: a processor and a memory, the memory storing programs or instructions capable of being executed on the processor, and the programs or instructions being executed by the processor to implement the steps of the expression recognition method of any one of the first aspect.

[0020] According to a fifth aspect of the present disclosure, a computer-readable storage medium is also provided, the storage medium storing programs or instructions, and the programs or instructions being executed by a processor to implement the steps of the expression recognition method of any one of the first aspect.

[0021] The expression recognition method provided by the embodiment of the present disclosure comprises: collecting images of a plurality of face regions in a plurality of sampling periods to obtain first images of the plurality of face regions; determining an effective expression time period according to the first images collected in the plurality of sampling periods; determining an expression label of a face region group in each sub time period according to the first images of the face region group in each sub time period; and determining an expression of the effective expression time period according to the expression labels of the plurality of face region groups in the plurality of sub time periods in the effective expression time period. According to the expression recognition method, the expression recognition is performed by collecting images of the plurality of face regions, so that the accuracy and recognition rate of the expression recognition are improved, and the calculation amount is reduced, and the expression recognition of the face can be realized by using less data.

[0022] Other features and advantages of the present application will become apparent from the following detailed description of illustrative embodiments thereof, which description should be taken in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS

[0023] The accompanying drawings incorporated in and forming a part of the specification illustrate embodiments of the present application and, together with the description, serve to explain the principles of the application.

[0024] Figure 1 is a hardware configuration structure diagram of a terminal device that can be used to implement an embodiment;

[0025] Figure 2 is a flowchart of an expression recognition method according to an embodiment;

[0026] Figure 3 is a flowchart of an expression recognition method according to another embodiment;

[0027] Figure 4 is a schematic diagram of a head-mounted electronic device according to another embodiment;

[0028] Figure 5 is a schematic diagram of a face region of an expression recognition method according to another embodiment;

[0029] Figure 6 is a schematic diagram of an expression recognition method according to another embodiment;

[0030] Figure 7 is a schematic diagram of image splicing of an expression recognition method according to another embodiment;

[0031] Figure 8 is a schematic diagram of an expression label of an expression recognition method according to another embodiment;

[0032] Figure 9 is a structural schematic diagram of an expression recognition system according to another embodiment;

[0033] Figure 10 is a schematic diagram of a head-mounted electronic device according to yet another embodiment;

[0034] Figure 11 is a schematic diagram of a terminal device according to yet another embodiment. DETAILED DESCRIPTION

[0035] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. Note that the relative arrangement, numerical expressions, and numerical values of components and steps set forth in these embodiments are not limiting to the scope of the present disclosure unless otherwise specifically stated.

[0036] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way limiting to the scope of the disclosure and its applications or uses.

[0037] Techniques, methods, and devices known to those of ordinary skill in the relevant art can not be discussed in detail herein. However, where appropriate, techniques, methods, and devices should be considered part of the description of the present disclosure.

[0038] In all of the compositions and methods described herein, any of the specific values can be replaced by alternative values of a similar nature. Thus, other examples of the exemplary embodiments can have different values.

[0039] Note that like references and characters designate like items throughout the attached drawings and the detailed description, and thus once an item is defined in one drawing, it is not necessary to discuss it further in subsequent drawings.

[0040] <Implementation Environment and Hardware Configuration>

[0041] Figure 1 A hardware configuration structure diagram of a terminal device 1000 to which an expression recognition method according to an embodiment of the present disclosure can be applied.

[0042] As Figure 1As shown, the terminal device 1000 can include a processor 1100, a memory 1200, an interface device 1300, a display device 1400, an input device 1500, and the like. Among them, the processor 1100 is configured to execute a computer program, which can adopt an instruction set of an architecture such as x86 (The X86 architecture microprocessor executed computer language instruction set), Arm (Acorn RISC Machine advanced reduced instruction set machine), RISC (Reduced Instruction Set Computer), MIPS (MIPS architecture adopts reduced instruction set (RISC) processor architecture), SSE (Streaming SIMD Extensions Internet data stream single instruction sequence extension), and the like. The memory 1200 includes, for example, ROM (read only memory), RAM (random access memory), non-volatile memory such as a hard disk, and the like. The interface device 1300 is a physical interface, for example, a USB interface or a headphone interface, and the like. The display device 1400 can be a display screen, which can be a touch display screen. The input device 1500 can include a keyboard, a mouse, and the like, and can also include a touch device.

[0043] In the embodiment, the memory 1200 of the terminal device 1000 is configured to store a computer program for controlling the processor 1100 to operate to implement the expression recognition method according to any embodiment. The skilled person can design the computer program according to the disclosed scheme in the specification. How the computer program controls the processor 1100 to operate is known in the art, and therefore will not be described in detail here.

[0044] Those skilled in the art will understand that although a plurality of devices of the terminal device 1000 are shown in Figure 1 , the terminal device 1000 of the embodiments of the present disclosure can only involve part of the devices therein, and can also contain other devices, which are not limited herein.

[0045] <Method Embodiment>

[0046] Figure 2 A flowchart of an expression recognition method according to an embodiment is shown, which is applied to a face, the face includes a plurality of face regions, the plurality of face regions belong to different face region groups, each face region group includes at least one face region, and the expressions of the plurality of face region groups can jointly reflect the expression of the face. The expression recognition method can be implemented by, for example, the terminal device 1000 as shown in Figure 1 .

[0047] The facial expression recognition method includes the following steps S1000 to S1320, which are described in detail below:

[0048] Step S1000: Capturing images of multiple facial regions simultaneously at a preset sampling frequency to obtain a first image of each facial region; wherein the multiple facial regions belong to different facial region groups, and each facial region group includes at least one facial region.

[0049] When facial skin changes, facial muscles stretch or tense, causing the skin to move. The changes appear in specific facial areas, such as the forehead, the sides of the eyes, or the cheekbones. The facial expressions in this disclosure are categorized into six categories: relaxation, laughter, crying, anger, left wink, and right wink. Other possible facial expressions are not limited here.

[0050] By simultaneously capturing images of multiple facial regions at a preset sampling frequency, first images of the multiple facial regions within multiple sampling periods can be obtained. It should also be noted that the multiple facial regions can constitute different facial region groups, with each facial region group including at least one facial region. In other words, the first image of each facial region group includes the first images of each facial region within the group. The preset sampling frequency can be set based on actual circumstances and is not limited here.

[0051] According to one embodiment of the present disclosure, a first image is acquired based on acquisition by an optical navigation sensor.

[0052] Figure 4 The position distribution of multiple optical navigation sensors is shown, namely S1 to S7. Figure 5 The figure shows multiple facial areas set on the face, namely F1 to F7. The positions of the multiple optical navigation sensors correspond to the positions of the multiple facial areas. That is, an optical navigation sensor is set above the skin of each facial area, and the optical navigation sensor is placed close to the skin of the corresponding facial area. The optical navigation sensor can capture images of the skin of the corresponding facial area when changes occur, and the optical navigation sensor can also capture preset images of the facial area in a relaxed state. The image captured by the optical navigation sensor contains information such as the direction and distance of skin movement, and the facial expression can be determined by comparing it with the preset image of the face in a relaxed state. The position distribution and number of the multiple optical navigation sensors, as well as the position distribution and number of the multiple facial areas, can be manually set according to actual conditions and are not limited here.

[0053] When the skin of the face changes, the plurality of facial regions are respectively imaged by the plurality of optical navigation sensors in a plurality of sub-time periods. The occurrence and end of an expression usually takes 2 to 3 seconds.

[0054] For example, the facial regions can be imaged once every 2 milliseconds as a sampling period. Thus, the optical navigation sensors can collect images of the skin of the plurality of facial regions in a plurality of sampling periods. That is, the first image includes facial images of the plurality of facial regions collected in all sampling periods. It should be noted that the interval time of each sampling period is the same.

[0055] According to one embodiment of the present disclosure, the facial region group is selected as a first facial region and a second facial region that are left-right symmetrical; and / or, the facial region group is selected as a plurality of facial regions belonging to the same part of the face.

[0056] When the skin of the face changes, the corresponding skin of the facial regions moves in different directions and distances. For example, when the facial image changes, the skin in the forehead region of the face is stretched, but the directions or distances of the corresponding skin stretching are different. According to the differences in the directions and distances of the skin movement, the expression corresponding to the facial region can be determined.

[0057] On this basis, by grouping the plurality of facial regions of the face, the directions and distances of the skin movement when the facial image changes can be more subtly discovered. The facial region group can include a first facial region and a second facial region that are left-right symmetrical according to the facial midline, and can also include a plurality of facial regions belonging to the same part of the face. For example, the positions of the corners of the eyes on both sides will appear symmetrical skin movement when smiling, and the skin in the facial regions on both sides of the corners of the eyes will be synchronously pulled up, so the two left-right symmetrical facial regions can be divided into the same facial region group. The forehead of the face is a relatively large region, and three facial regions are provided at the forehead position. When the facial image changes, the skin in the three facial regions of the forehead will synchronously move, so the three facial regions belonging to the forehead part can be divided into the same facial region. For the case that the skin of other facial regions may synchronously move, details are not described herein.

[0058] In the embodiments of the present disclosure, the plurality of facial regions belong to different facial region groups, and each facial region group includes at least one facial region. By dividing the plurality of facial regions into the same facial region group, the expression corresponding to the facial region group can be comprehensively determined, and the expression misjudgment caused by the abnormal skin movement of a single facial region in the facial region group in a special case can be avoided. The accuracy of expression recognition can be improved by determining the expression through the plurality of facial regions.

[0059] Exemplarily, as shown in Figure 5 The face is divided into six face regions, namely, face regions F1 to F7. Among them, F1, F2 and F3 are located in the forehead region, F4 and F6 are located in the regions on both sides of the eyes, and F5 and F7 are located in the regions of the two cheekbones. Face region F4 and face region F6 are symmetrically distributed on the face; face region F5 and face region F7 are symmetrically distributed on the face; face region F2 and face region F3 are symmetrically distributed on the face, and face region F1, face region F2 and face region F3 are all located in the forehead position, i.e., distributed in the same part of the face. Since the forehead region is relatively wide, the three face regions located in the forehead can be divided into a face region group, which can be combined to more accurately reflect the expression. Therefore, face region F4 and F6 can be divided into face region group 2, face region F1, F2 and F3 can be divided into face region group 1, and face region F5 and F7 can be divided into face region group 3. The division rule of the face region group can be formulated according to the actual situation, which is not limited here.

[0060] It should be noted that, for example, when the face regions F4 and F6 symmetrically distributed on the face change, the directions of the skin movement in the two face regions are symmetrical and the distances are approximately the same; when the face expression changes, the skin in the face regions F1, F2 and F3 located in the same part moves in the same way, so they can also be combined to reflect the face expression.

[0061] In step S1100, the effective expression time period is determined according to the first images collected in the continuous multiple sampling periods; wherein the effective expression time period includes a plurality of continuous sub-time periods, and the number of sampling periods corresponding to each sub-time period is the same.

[0062] In the embodiment of the present disclosure, the effective sampling period that can accurately reflect the start and end of the expression can be extracted from the first image, so as to determine the effective expression time period in which the expression really changes, and the image in the effective time period is subjected to expression recognition and judgment, thereby reducing the data processing amount, improving the efficiency of expression recognition, and making the expression judgment more accurate.

[0063] In one embodiment, the step S1100 of determining the effective expression time period according to the first images collected in the continuous multiple sampling periods can include the following steps S1110-S1120, which are described in detail as follows:

[0064] In step S1110, the change amount of the first image of the face region compared with the preset image of the face region is calculated, and the change amount is taken as the reference value of the first image; wherein the preset image is an image collected in a relaxed state of the face.

[0065] In the embodiments of the present disclosure, after obtaining the first images of the plurality of facial regions collected in the plurality of sampling periods, the first images can also be subtracted from the preset images (i.e., the images of the facial regions collected in the relaxed state) of the facial regions, so as to obtain the pixel change amount between the two, and the change amount can be taken as the reference value of the first images.

[0066] Here, taking a complete facial image collected by the optical navigation sensor from the facial region F6 as an example, first, the first image of the facial region F6 is compared with the initial image in the relaxed state, and the change amount of each pixel point in the first image of the facial region F6 is obtained. As shown in Figure 6 Figure 6 The image change amount of the left eye blinking before and in the relaxed state in the above-mentioned image is obtained by subtracting the first image collected in the different sampling periods from the initial image of the facial region F6. That is, the change amount of the first image and the preset image can be obtained by the above-mentioned subtraction processing, and the change amount can be taken as the reference value of the first image of the corresponding facial region.

[0067] In step S1120, the reference value of the first image collected in the continuous plurality of sampling periods is determined to determine the expression start time and the expression end time, and the time period from the expression start time to the expression end time is taken as the effective expression time period.

[0068] The reference value of the first image in the above-mentioned content is the change amount of the first image of the facial region compared with the preset image of the facial region. In actual application, the change amount of the first image of each facial region compared with the preset image of the facial region can be reflected in the area of the region where the pixel change occurs obtained by subtracting the two images. Specifically, the reference value of the first image can be represented by the proportion of the area of the region where the pixel change occurs in the total area of the image. The determination method of the effective expression time period is described in detail as follows:

[0069] Exemplarily, according to Figure 6 It can be seen that, from the relaxed state to the left eye blinking, the area of the region where the pixel change occurs, i.e., the area of the white region, becomes larger and larger, and the proportion of the area of the white region in the image also becomes larger and larger. When the proportion of the area of the white region in the image grows to a set proportion, it indicates that the expression action starts to change from the relaxed state to the left eye blinking, and thus the proportion of the area of the white region in the image can be taken as the start of the expression.

[0070] ​Similarly, the white area in the face from the left eye blinking to the relaxation will be smaller and smaller, and the proportion of the area of the white area in the image will also be smaller and smaller. When the proportion of the area of the white area in the image is reduced to a set proportion, it can be considered that the expression has ended. Thus, after determining the time period included between the expression start and the expression end, the time period from the expression start time to the expression end time can be taken as the valid expression time period.

[0071] According to the above principle, in the specific implementation process, the valid expression time period can be determined according to the reference values of the first images of the plurality of face regions collected in a plurality of continuous sampling periods. When the reference values of the first images of the plurality of face regions all exceed the preset proportion in the plurality of continuous sampling periods, it is considered that the sampling period corresponding to the first reference value exceeding the preset proportion in the plurality of continuous sampling periods is the time point of the expression start. Similarly, when the reference values of the first images of the plurality of face regions are all lower than the preset proportion in the plurality of continuous sampling periods, it is considered that the sampling period corresponding to the first reference value lower than the preset proportion in the plurality of continuous sampling periods is the time point of the expression end. The preset proportion can be set according to the specific situation, which is not limited here.

[0072] Exemplarily, taking 10% as an example, when the reference values of the first images of the plurality of face regions all exceed 10% in the plurality of continuous sampling periods, it is considered that the sampling period corresponding to the first reference value exceeding 10% in the plurality of continuous sampling periods is the time point of the expression start. Similarly, when the reference values of the first images of the plurality of face regions are all lower than 10% in the plurality of continuous sampling periods, it is considered that the sampling period corresponding to the first reference value lower than 10% in the plurality of continuous sampling periods is the time point of the expression end. After determining the time points of the expression start and the expression end, the valid expression time period of the expression change is determined.

[0073] The embodiment of the present disclosure can more accurately judge the expression of the face region by extracting the valid expression time period from the plurality of sampling periods, reduce the processing amount of redundant data, and improve the speed of expression recognition.

[0074] In step S1200, the expression label of the face region group in each sub time period is determined according to the first image of the face region group in each sub time period.

[0075] According to the effective expression time period determined according to the first images of the continuous multiple sampling periods, the first image of each facial region in the effective expression time period can be obtained, the effective expression time period includes continuous multiple sub-time periods, and the first image of each facial region group in each sub-time period can be obtained according to the first images of the multiple facial regions in the same facial region group. Then, according to the first images of the facial region groups in each sub-time period, the expression labels of the multiple facial region groups in each sub-time period can be respectively judged.

[0076] In the expression judgment process, the modules for expression judgment are divided into first expression judgment modules and second expression judgment modules. Each first expression judgment module corresponds to a facial region group, and judges the expression label of the facial region group in each sub-time period. The second expression judgment module jointly judges the expression labels of the multiple facial region groups in the multiple sub-time periods.

[0077] In one embodiment, the step S1200 of determining the expression label of the facial region group in each sub-time period according to the first images of the facial region group in each sub-time period can include the following steps S1210-S1230, which are described in detail as follows:

[0078] Step S1210: The first images of the facial region group in the same sampling period in the sub-time period are spliced into the second images of the facial region group in the sub-time period.

[0079] According to the above content, the first images are images collected by the multiple facial regions in the multiple sampling periods. The first images of the multiple facial regions belonging to the same facial region group are spliced to obtain the second images collected by the multiple facial region groups in the multiple sampling periods in the sub-time period. It should be noted that the first images spliced here are the first images in the effective expression time period.

[0080] The specific splicing method is as shown in Figure 7 For example, the first images of the facial region F4 and the facial region F6 in the same facial region group are spliced, the first images of the facial region F4 and the facial region F6 in the same sampling period in the sub-time period are spliced first, the image data of the facial region F4 and the facial region F6 are 39*39 pixels, and the two are spliced into 39*78 pixel image data to obtain the second image, that is, the images collected by the multiple facial region groups in the same sampling period in the same sub-time period.

[0081] Step S1220: The second images of the facial region group in the sub-time period are spliced into the third images of the facial region group in the sub-time period according to the order of the sampling periods.

[0082] On the basis of the second image obtained by splicing the first image, the second image of the face region F4 and the face region F6 in the next sampling period in the sub time period is spliced, and the second images in the subsequent multiple sampling periods are spliced in the same way. 18 groups of images are spliced into 234*234 format image data, and 10 pixels in the length and width directions of the pixels of the image data are discarded respectively to obtain 224*224 format image data. On this basis, 36 groups of images are collected in the same way.

[0083] Finally, 54 groups of images collected in multiple continuous acquisition periods in the same sub time period are spliced into 224*224*3 format image data to obtain the third image. It should be noted that the third image can be 224*224*3 format image data including multiple sub time periods. The 224*224*3 format image data here can meet the format requirements of the input expression recognition model. The format requirements that the spliced image should meet can be adjusted according to the specific situation, which is not limited here.

[0084] Step S1230, input the third image of the face region group in the sub time period to the pre-trained classification model, and output the expression label of the face region group in the sub time period through the classification model.

[0085] In an embodiment of the present disclosure, the third images of multiple sub time periods are input into the classification model, which can output the expression labels corresponding to the third images in each sub time period, i.e. the expression labels corresponding to each face region group in each sub time period.

[0086] The specific process of the classification model outputting the expression label is as follows:

[0087] According to the order of the multiple sub time periods, the third images are input into the pre-trained classification model in turn. The trained classification model outputs the expression labels corresponding to each face region group in the multiple sub time periods according to the differences in the direction and distance of skin movement in the third images.

[0088] The classification model can be pre-trained, and the step of obtaining the classification model can include: obtaining a facial expression sample, wherein the facial expression sample has a label corresponding to the facial expression; training the model parameters of the selected model through the facial expression sample; and configuring the selected model according to the trained model parameters to obtain the classification model.

[0089] As Figure 8As shown, the expression labels can be L0, L1, L2, L3, L4, and L5, respectively, and the corresponding expression names are relaxation, smile, cry, anger, left eye blinking, and right eye blinking, respectively. The expression labels here can be recorded by numbers, and different numbers correspond to different expression names. The expression labels can also have other recording methods, which are not limited here.

[0090] In step S1300, the expression of the effective expression time period is determined according to the expression labels of the plurality of facial region groups in the plurality of sub-time periods in the effective expression time period.

[0091] According to the above content, inputting the third image of each sub-time period into the classification model can obtain the corresponding expression labels of the plurality of facial region groups in the plurality of sub-time periods in the effective expression time period. By analogy, the expression labels of the plurality of facial region groups in the plurality of sub-time periods can be calculated, and the expression labels of the plurality of facial region groups in the plurality of sub-time periods are jointly determined to obtain the expression result of the face, and finally the expression in the effective expression time period is obtained.

[0092] In one embodiment, the step S1300 of determining the expression of the effective expression time period according to the expression labels of the plurality of facial region groups in the plurality of sub-time periods in the effective expression time period can include steps S1310-S1320, which are described in detail as follows:

[0093] In step S1310, if the expression labels of the plurality of facial region groups in the target sub-time period are the same, the expression determination result of the target sub-time period is set to be valid, and the expression labels of the plurality of facial region groups in the target sub-time period are taken as the expression labels of the target sub-time period; and

[0094] If the expression labels of the plurality of facial region groups in the target sub-time period are not the same, the expression determination result of the target sub-time period is set to be invalid, and the determination of the expression labels of the target sub-time period is skipped, and the expression labels of the plurality of facial region groups in the next target sub-time period are determined; wherein the target sub-time period is any sub-time period in the effective expression time period.

[0095] For example, if the target sub-time period is the first sub-time period in the effective expression time period, if the expression labels of the plurality of facial region groups in the first sub-time period are the same, the expression determination result of the first sub-time period is set to be valid, and the expression labels of the plurality of facial region groups in the first sub-time period are taken as the expression labels of the first sub-time period. If the expression labels of the plurality of facial region groups in the first sub-time period are different, the expression determination result of the first sub-time period is set to be invalid, and the information of the expression labels of the plurality of facial region groups in the first sub-time period is not displayed.

[0096] In the embodiments of the present disclosure, seven facial regions are taken as an example. The seven facial regions are divided into three facial region groups. The second processor jointly determines the expression labels corresponding to the multiple facial region groups in the target sub-time period. The specific rules of joint determination are as follows:

[0097] First, the expression of each facial region group in the target sub-time period is determined. If the expression labels corresponding to the multiple facial regions in the same group in the target sub-time period are consistent, the expression label of the facial region group in the target sub-time period is output. If the expression labels corresponding to the multiple facial regions in the same group in the target sub-time period are inconsistent, the expression label of the facial region group in the target sub-time period cannot be output, and the information of the expression label of the facial region group in the target sub-time period is not displayed.

[0098] It should be noted that if the expression labels of the multiple facial regions in any one facial region group in the target sub-time period are inconsistent, the determination of the expression label of the target sub-time period is skipped, and the determination of the expression labels of the multiple facial region groups in the next target sub-time period is directly performed.

[0099] Exemplarily, it is taken as an example that the same facial region group includes three facial regions. If the expression labels corresponding to the three facial regions in the same facial region group in the target sub-time period are consistent, all being L1, L1 and L1, the expression label of the facial region group in the target sub-time period is output as L1. If the expression labels corresponding to the three facial regions in the same group in the target sub-time period are inconsistent, being L1, L1 and L2 respectively, the expression label of the facial region group in the target sub-time period cannot be output, the information of the expression label of the facial region group in the target sub-time period is not displayed, and the determination of the expression label of the target sub-time period is skipped, and the determination of the expression labels of the multiple facial region groups in the next target sub-time period is directly performed.

[0100] On this basis, if the expression labels corresponding to the multiple facial region groups in the target sub-time period are the same, i.e., the expression labels corresponding to the multiple facial region groups in the target sub-time period are all L1, it is indicated that the expression labels corresponding to each facial region group in the target time period are the same, and the expression label L1 of the target sub-time period can be output, and the corresponding expression name is smile. If the expression labels corresponding to the multiple facial region groups in the target sub-time period are inconsistent, it is indicated that the expression labels corresponding to the multiple facial regions in the sub-time period are not the same, and no expression label will be output, and the expression determination result is set to be invalid.

[0101] For example, the plurality of facial region groups are exemplified by three facial region groups. If the expression labels corresponding to the three facial region groups in the target sub-time period are all L1, it indicates that the expression labels corresponding to each facial region group in the target time period are the same, so the expression label L1 of the target sub-time period can be output, and the expression determination result is set as valid, and the corresponding expression name is smile. If the expression labels corresponding to the three facial region groups in the target sub-time period are L1, L2 and L2 respectively, it indicates that the expression labels corresponding to the three facial region groups in the target sub-time period are different, so no expression label will be output, and the expression determination result is set as invalid.

[0102] In step S1320, if the expression determination results of at least N sub-time periods in the continuous M sub-time periods in the valid expression time period are valid and the expression labels of the N sub-time periods are the same, the expression label of the N sub-time periods is taken as the expression of the valid expression time period.

[0103] If the expression determination results of less than N sub-time periods in the continuous M sub-time periods in the valid expression time period are valid, or the expression determination results of at least N sub-time periods in the continuous M sub-time periods are valid, but the expression labels of the N sub-time periods are not the same, the expression labels of the continuous M sub-time periods after the last sub-time period in which the expression labels are the same are re-determined. M and N are positive integers, and N is less than or equal to M.

[0104] In an embodiment of the present disclosure, the second expression determination module can also determine the expression labels of the continuous sub-time periods of the plurality of facial region groups. The specific determination rules are as follows:

[0105] When the plurality of facial region groups appear the following situation in the continuous M sub-time periods: the expression labels of the plurality of facial region groups in each of at least N sub-time periods are the same, and the expression labels of the N sub-time periods are also the same, the expression label is taken as the expression of the valid expression time period.

[0106] When the plurality of facial region groups appear the following situation in the continuous M sub-time periods: the expression labels of the plurality of facial region groups in each of less than N sub-time periods are the same, or the expression labels of the plurality of facial region groups in each of at least N sub-time periods are the same, but the expression labels of the N sub-time periods are not all the same, the information of the expression label of the valid expression time period is not displayed, and the expression labels of the continuous M sub-time periods after the last sub-time period in which the expression labels are the same are re-determined.

[0107] Wherein, in the process of determining the expression label of the plurality of sub-time periods, the number of consecutive sub-time periods M, the number of N, and the condition of determining the expression label as the valid expression time period can be set according to the actual situation, which is not limited here.

[0108] Exemplarily, M is set to 10, and N is set to 7. The following is a specific description:

[0109] When the expression labels of the plurality of face region groups in each of at least 7 consecutive sub-time periods of the 10 consecutive sub-time periods are the same, and the expression labels of the corresponding target sub-time periods are also the same, the expression label is determined as the valid expression time period;

[0110] When the number of the plurality of target sub-time periods in which the expression labels are consistent is less than 7, the information of the expression label of the valid expression time period is not displayed, and the expression labels of the next 10 consecutive sub-time periods are re-determined from the last sub-time period in which the expression labels are consistent;

[0111] Or, when the expression labels of the plurality of face region groups in each of at least 7 consecutive sub-time periods of the 10 consecutive sub-time periods are the same, but the expression labels of each sub-time period are not all the same, the information of the expression label of the valid expression time period is not displayed, and the expression labels of the next 10 consecutive sub-time periods are re-determined from the last sub-time period in which the expression labels are consistent.

[0112] The expression recognition method provided by the embodiment of the present disclosure determines the valid expression time period according to the first images collected in the plurality of sampling periods, and then determines the expression label of the face region group in each sub-time period according to the first images of the face region group in each sub-time period. Finally, the expression of the valid expression time period is determined according to the expression labels of the plurality of face region groups in the plurality of sub-time periods in the valid expression time period. According to the expression recognition method of the present disclosure, the expression recognition is performed by collecting the images of the set plurality of face regions, which improves the accuracy and recognition rate of the expression recognition, and reduces the calculation amount, so that the expression recognition of the face can be realized by less data.

[0113] <Example>

[0114] The following refers to Figure 3 The whole process of outputting the face expression label of the plurality of face region groups in the target sub-time period is described:

[0115] Exemplarily, the face is provided with three face region groups, namely, face region group 1, face region group 2 and face region group 3. The face region group 1 includes face region 4 and face region 5. The face region group 2 includes face region 1, face region 2 and face region 3. The face region group 3 includes face region 6 and face region 7. According to the image sampling on the plurality of face regions simultaneously, the first image of each region collected in the target sub-time period is obtained. On this basis, the third image of the plurality of face region groups in the target sub-time period can be obtained through splicing. Then, the third image of each face region group collected in the target sub-time period is input into the trained classification model for recognition, and the corresponding recognition result, namely, the expression label, can be obtained. Through the joint determination of the expression labels corresponding to the three face region groups, the expression label of the face in the target sub-time period can be output.

[0116] <system embodiment>

[0117] In the embodiment of the present application, an expression recognition system 2000 is further provided, Figure 9 The structure diagram of the expression recognition system 2000 is shown, and the expression recognition system 2000 includes a face image sampling module 2100, an effective expression determination module 2200, a first expression determination module 2300 and a second expression determination module 2400. Wherein:

[0118] The face image sampling module 2100 is used for collecting images of a plurality of face regions simultaneously according to a preset sampling frequency, to obtain the first image of each face region; wherein the plurality of face regions belong to different face region groups, and each face region group includes at least one face region;

[0119] The effective expression determination module 2200 is used for determining the effective expression time period according to the first images collected in the continuous plurality of sampling periods; wherein the effective expression time period includes a plurality of continuous sub-time periods, and the number of sampling periods corresponding to each sub-time period is the same;

[0120] The first expression determination module 2300 is used for determining the expression label of the face region group in each sub-time period according to the first image of the face region group in each sub-time period;

[0121] The second expression determination module 2400 is used for determining the expression of the effective expression time period according to the expression labels of the plurality of face region groups in the plurality of sub-time periods in the effective expression time period.

[0122] It should be noted that in the expression determination process, the first expression determination module can be multiple, each first expression determination module can determine the expression label of the corresponding face region group in the plurality of sub-time periods, and then the second expression determination module determines the expression label of the plurality of face region groups in the plurality of sub-time periods.

[0123] It should be noted that, although several means or units of the system for action execution are mentioned in the above detailed description, such a division is not mandatory. Indeed, according to the method of implementing the application, the characteristics and functions of two or more means or units described above can be embodied in one means or unit. Conversely, the characteristics and functions of one means or unit described above can be further divided into means or units embodied by several means or units.

[0124] <Device Embodiment One>

[0125] In this embodiment, a head-mounted electronic device 3000 is also provided, which is applied to a face, the face comprising a plurality of face regions, the plurality of face regions belonging to different face region groups, and each face region group comprising at least one face region. Figure 10 A structural schematic diagram of the head-mounted electronic device 3000 is shown, the head-mounted electronic device comprising: a shell 3100, a face-adhesion assembly 3200, a plurality of optical navigation sensors 3300, and a digital signal processor 3400.

[0126] Among them:

[0127] The shell 3100 can be fixed to the face, and the shell has an accommodating space inside;

[0128] The face-adhesion assembly 3200 is arranged in the shell, and the face-adhesion assembly can be attached to the skin of the face;

[0129] The plurality of optical navigation sensors 3300 are arranged in the face-adhesion assembly, and the plurality of optical navigation sensors are arranged corresponding to the plurality of face regions, the plurality of optical navigation sensors are closely attached to the skin of the face regions, and the plurality of optical navigation sensors are used to collect images of each face region; wherein the images comprise movement direction and distance information of the skin in each face region;

[0130] Among them, as Figure 4As shown, the facial fitting component is provided with a plurality of optical navigation sensors, S1 to S7 respectively represent the position distribution of the plurality of optical navigation sensors on the facial fitting component. After the facial fitting component is fitted to the face, S1, S2, and S3 are located in the forehead area of ​​the face, S4 and S6 are located in the areas on both sides of the corners of the eyes, and S5 and S7 are located in the areas on both sides of the cheekbones of the face. S4 and S6 are symmetrically distributed about the face; S5 and S7 are symmetrically distributed about the face; facial area S2 and facial area S3 are symmetrically distributed on the face, and S1, S2, and S3 are all located in the forehead area, that is, distributed in the same part of the face. An optical navigation sensor is respectively provided at the positions S1 to S7 mentioned above. The number and position distribution of the optical navigation sensors provided on the facial fitting component can be set according to actual conditions and are not limited here.

[0131] Digital signal processor 3400 is disposed in the accommodation space. The digital signal processor 3400 can be configured to simultaneously capture images of multiple facial regions at a preset sampling frequency to obtain a first image of each facial region; wherein the multiple facial regions belong to different facial region groups, each facial region group including at least one facial region; determine an effective expression time period based on the first images captured during multiple consecutive sampling periods; wherein the effective expression time period includes multiple consecutive sub-time periods, each sub-time period corresponding to the same number of sampling periods; determine an expression label for the facial region group in each sub-time period based on the first image of the facial region group in each sub-time period; and determine an expression for the effective expression time period based on the expression labels of the multiple facial region groups in multiple sub-time periods within the effective expression time period.

[0132] <Equipment Example 2>

[0133] In this embodiment, a terminal device 7000 is also provided. Figure 11 As shown, the terminal device 7000 may include a processor 7100 and a memory 7200. The memory 7200 stores programs or instructions that can be run on the processor. When the programs or instructions are executed by the processor, the steps of the expression recognition method of any method embodiment of the present invention are implemented.

[0134] <Medium Example>

[0135] In this embodiment, a computer-readable storage medium is further provided, on which computer instructions are stored. When the computer instructions are executed by a processor, the steps of the expression recognition method according to any method embodiment of the present invention are implemented.

[0136] The present application can be a system, a method, and / or a computer program product. The computer program product can include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present application.

[0137] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or punched tape, a

[0138] The computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.

[0139] Computer readable program instructions for carrying out operations of the present application can be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate array (FPGA), or programmable logic array (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present application.

[0140] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0141] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0142] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0143] The flow diagrams and the block diagrams in the drawings are presented to illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present application. In this regard, each block in the flow diagrams and the block diagrams can represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logic functions. In some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and

[0144] Embodiments of the application have been described above, and the description is intended to be illustrative of the embodiments of the application and not exhaustive. Numerous modifications and adaptations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The scope of the application is defined by the appended claims.

Claims

1. A facial expression recognition method, characterized in that, The method comprises: collecting images of a plurality of face regions simultaneously according to a preset sampling frequency to obtain a first image of each face region; wherein the plurality of face regions belong to different face region groups, and each face region group comprises at least one face region; determining an effective expression time period according to the first images collected in a plurality of consecutive sampling periods; wherein the effective expression time period comprises a plurality of consecutive sub-time periods, and the number of sampling periods corresponding to each sub-time period is the same; determining an expression label of the face region group in each sub-time period according to the first images of the face region group in each sub-time period; determining an expression of the effective expression time period according to the expression labels of the plurality of face region groups in the plurality of sub-time periods in the effective expression time period; determining an expression label of the face region group in each sub-time period according to the first images of the face region group in each sub-time period, comprising: splicing the first images of the face region group in the same sampling period in the sub-time period into a second image of the face region group in the sub-time period; splicing the second images of the face region group in the sub-time period into a third image of the face region group in the sub-time period according to the order of the sampling periods; inputting the third image of the face region group in the sub-time period into a pre-trained classification model to output the expression label of the face region group in the sub-time period by the classification model.

2. The method of claim 1, wherein, The first image is obtained based on an optical navigation sensor.

3. The method of claim 1, wherein, The face region group is selected as a first face region and a second face region that are left-right symmetrical; and / or The face region group is selected as a plurality of face regions belonging to the same part of the face.

4. The method of claim 1, wherein, determining an effective expression time period according to the first images collected in a plurality of consecutive sampling periods, comprising: calculating a change amount of the first image of the face region compared with a preset image of the face region, taking the change amount as a reference value of the first image; wherein the preset image is an image collected in a relaxed state of the face; determining an expression start time and an expression end time according to the reference values of the first images collected in a plurality of consecutive sampling periods, and taking a time period from the expression start time to the expression end time as an effective expression time period.

5. The method of claim 1, wherein, determining an expression of the effective expression time period according to the expression labels of the plurality of face region groups in the plurality of sub-time periods in the effective expression time period, comprising: if the expression labels of the plurality of face region groups in a target sub-time period are the same, setting the expression determination result of the target sub-time period as valid and taking the expression labels of the plurality of face region groups in the target sub-time period as the expression label of the target sub-time period; if the expression labels of the plurality of face region groups in a target sub-time period are not the same, setting the expression determination result of the target sub-time period as invalid and skipping the determination of the expression label of the target sub-time period and determining the expression labels of the plurality of face region groups in the next target sub-time period; wherein the target sub-time period is any sub-time period in the effective expression time period.

6. The method of claim 5, wherein The determining the expression of the effective expression time period according to the expression labels of the plurality of face region groups in the plurality of sub-time periods in the effective expression time period further includes: If the expression determination results of at least N sub-time periods in the continuous M sub-time periods in the effective expression time period are valid and the expression labels of the N sub-time periods are the same, the expression labels of the N sub-time periods are taken as the expression of the effective expression time period; If the expression determination results of less than N sub-time periods in the continuous M sub-time periods in the effective expression time period are valid, or the expression determination results of at least N sub-time periods in the continuous M sub-time periods are valid, but the expression labels of the N sub-time periods are not the same, the expression labels of the continuous M sub-time periods after the last sub-time period with consistent expression labels are re-determined; wherein the M and the N are positive integers, and the N is less than or equal to the M.

7. An expression recognition system characterized by, The expression recognition system includes: A face image sampling module configured to simultaneously collect images of a plurality of face regions according to a preset sampling frequency to obtain first images of each face region; wherein the plurality of face regions belong to different face region groups, and each face region group includes at least one face region; An effective expression determination module configured to determine an effective expression time period according to the first images collected in a plurality of continuous sampling periods; wherein the effective expression time period includes a plurality of continuous sub-time periods, and each sub-time period corresponds to the same number of sampling periods; A first expression determination module configured to determine expression labels of face region groups in each sub-time period according to the first images of the face region groups in each sub-time period; A second expression determination module configured to determine an expression of the effective expression time period according to the expression labels of the plurality of face region groups in the plurality of sub-time periods in the effective expression time period; The first expression determination module is specifically configured to: splice the first images of the face region groups in the same sampling period in the sub-time period into second images of the face region groups in the sub-time period; splice the second images of the face region groups in the sub-time period into third images of the face region groups in the sub-time period according to the order of the sampling periods; input the third images of the face region groups in the sub-time period into a pre-trained classification model, and output the expression labels of the face region groups in the sub-time period through the classification model.

8. A head-mounted electronic device applied to a face, the face comprising a plurality of face regions, the plurality of face regions belonging to different face region groups, each face region group comprising at least one face region, characterized in that, It includes: A shell that can be fixed to the face, the inside of the shell being provided with a containing space; A face-adhering assembly provided in the shell, the face-adhering assembly being capable of adhering to the skin of the face; A plurality of optical navigation sensors are arranged on the face fitting assembly, and the plurality of optical navigation sensors are arranged on the plurality of face regions correspondingly. The plurality of optical navigation sensors are arranged close to the skin of the face regions, and the plurality of optical navigation sensors are used to collect images of each face region. The images include the moving direction and distance information of the skin in each face region. A digital signal processor is arranged in the accommodating space, and the digital signal processor is used to process the images collected by the plurality of optical navigation sensors according to the expression recognition method in any one of claims 1-6, and output the expression of the face.

9. A terminal device, comprising: A processor and a memory are included, the memory stores programs or instructions executable on the processor, and the programs or instructions are executed by the processor to implement the steps of the expression recognition method in any one of claims 1-6.

10. A computer-readable storage medium, characterized in that, The programs or instructions are stored on the readable storage medium, and the programs or instructions are executed by the processor to implement the steps of the expression recognition method in any one of claims 1-6.

Citation Information

Patent Citations

  • Facial expression recognition-based emotion recognition method and device

    CN108564007A

  • Facial expression recognition method and device

    CN109800734A