Facial expression labeling method and device, and medium

By combining control group comparisons and small-world network annotation methods with sensor monitoring, the subjectivity problem of traditional facial expression annotation was solved, achieving higher annotation accuracy and consistency.

CN117197575BActive Publication Date: 2026-05-29QINGDAO CLASS COGNITIVE ARTIFICIAL INTELLIGENCE CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
QINGDAO CLASS COGNITIVE ARTIFICIAL INTELLIGENCE CO LTD
Filing Date
2023-09-19
Publication Date
2026-05-29

Smart Images

  • Figure CN117197575B_ABST
    Figure CN117197575B_ABST
Patent Text Reader

Abstract

The application discloses a facial expression labeling method and device and a medium, and relates to the technical field of data processing. The method comprises the following steps: determining a plurality of expression data, determining a control group from the plurality of expression data, wherein the control group comprises first expression data and second expression data; determining a labeling type of the plurality of expression data, determining an expression difference of the control group according to the labeling type; and performing expression labeling on the control group according to the expression difference. The application compares facial expression data two by two, compares the differences of the two in terms of "positive degree" and "awakening degree", and adopts a simple selection method to determine the expression with a higher positive degree or a higher awakening degree in the paired facial expression. The application greatly reduces the difficulty of the labeling task, thereby greatly reducing the intra-individual variability and inter-individual variability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a method, device and medium for annotating facial expressions. Background Technology

[0002] Facial expression data is typically labeled manually by annotators. This can involve labeling various expression types, such as joy, anger, and sadness, or using a rating system based on arousal and pleasure levels. However, traditional annotation methods are highly subjective, and the labeling standards may vary depending on the annotator's fatigue and psychological state at different times, leading to errors. Summary of the Invention

[0003] To address the aforementioned issues, this application proposes a method for annotating facial expressions, comprising: determining multiple expression data; determining a control group from the multiple expression data, wherein the control group includes first expression data and second expression data; determining the annotation type of the multiple expression data; determining the expression differences of the control group based on the annotation type; and annotating the control group with expressions based on the expression differences.

[0004] In one example, the method further includes: determining the first expression data as a basic expression and labeling the basic expression as a basic result; and labeling the second expression data with expressions based on the basic result and the expression difference.

[0005] In one example, the method further includes: determining other facial expression data based on the plurality of facial expression data, and determining a new control group based on the other facial expression data and the first facial expression data; determining the facial expression differences of the new control group, and labeling the other facial expression data based on the facial expression differences and the baseline results.

[0006] In one example, the method further includes: determining the expression result of the second expression data, and determining the annotation result of the other expression data;

[0007] The multiple expression data are sorted based on the basic results, the expression results of the second expression data, and the annotation results of the other expression data.

[0008] In one example, the method further includes: determining the difference probability of the control group, wherein the difference probability is calculated using the following formula:

[0009]

[0010] Where P(i>j) is the probability that the degree of expression of the first expression data is greater than that of the second expression data, i represents the first expression data, j represents the second expression data, and λ iλ represents the degree of expression of the primary facial expression. j α represents the degree of expression of the second expression. i Let α be the hidden parameter of the first expression. j This is a hidden parameter for the second expression.

[0011] In one example, the method further includes: calculating a labeling result based on the difference probability, and labeling according to the labeling result, wherein the formula for calculating the labeling result is:

[0012]

[0013] Where lnL represents the annotation result, and n ij The data group whose expression level of the first expression data is greater than that of the second expression data.

[0014] In one example, the method further includes: establishing an expression data network based on the plurality of expression data, using the plurality of expression data as nodes in the expression data network; determining a first node and a second node in the expression data network, establishing a control group based on the expression data corresponding to the first node and the expression data corresponding to the second node, thereby labeling the control group with expressions.

[0015] In one example, the method further includes: determining the gaze duration of the annotator gazing at the facial expression data using a sensor, determining a pre-set gaze threshold, comparing the gaze duration with the gaze threshold; if the gaze duration is greater than the gaze threshold, determining the gaze duration as a valid time; determining the annotator's accuracy based on the valid time, and annotating the facial expression data based on the accuracy.

[0016] On the other hand, this application also proposes a facial expression annotation device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the facial expression annotation device to perform: determining a plurality of expression data; determining a control group from the plurality of expression data, wherein the control group includes first expression data and second expression data; determining an annotation type for the plurality of expression data; determining expression differences in the control group based on the annotation type; and annotating the control group with expressions based on the expression differences.

[0017] On the other hand, this application also proposes a non-volatile computer storage medium storing computer-executable instructions, the computer-executable instructions being configured to: determine multiple facial expression data; determine a control group from the multiple facial expression data, wherein the control group includes first facial expression data and second facial expression data; determine the annotation type of the multiple facial expression data; determine the facial expression differences of the control group according to the annotation type; and annotate the control group with facial expressions according to the facial expression differences.

[0018] This application compares facial expression data pairwise, comparing the differences in "activity level" and "arousal level," and uses a simple selection method to determine whether the paired facial expressions have a higher level of activity or a higher level of arousal. This application significantly reduces the difficulty of the annotation task, thereby greatly reducing intra-individual and inter-individual variability. Attached Figure Description

[0019] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0020] Figure 1 This is a flowchart illustrating a facial expression annotation method according to an embodiment of this application;

[0021] Figure 2 This is a schematic diagram of a facial expression annotation device according to an embodiment of this application. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0023] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.

[0024] In traditional data annotation processes, facial expression data annotation has a high probability of introducing intra-individual and inter-individual variability, meaning there is a significant risk of introducing human annotation error. To control this error, in addition to using simple averaging methods to process the annotated facial expression data, traditional methods also employ aggregation models with confusion matrices, using binary or categorical label classification problems. Aggregating continuous labels is reminiscent of analysis of variance models and factor analysis. However, these methods still cannot effectively control the subjective bias introduced by the annotator.

[0025] like Figure 1 As shown, in order to solve the above problems, this application provides a method for annotating facial expressions, the method including:

[0026] S101. Determine multiple facial expression data, and determine a control group from the multiple facial expression data, wherein the control group includes first facial expression data and second facial expression data.

[0027] Define the arousal level (V) or arousal level (A) of each facial expression data point in a dataset as α. v α a The log-odds function of the difference in arousal level between any two facial expressions, i.e., facial expression image i (referred to as the first expression data) and facial expression image j (referred to as the second expression data), is as follows:

[0028]

[0029] Where P(i>j) is the probability that the first expression data has a greater degree of arousal or positivity than the second expression data, and this probability satisfies a log-odds function, where i represents the first expression data, j represents the second expression data, and λ i λ represents the degree of arousal or positivity of the primary facial expression. j α represents the degree of arousal or positivity of the second expression. i Let α be the hidden parameter of the first expression. j α is a hidden parameter for the second expression. i and α j This refers to the true values ​​of facial expression image i and facial expression image j in terms of arousal level or arousal level. These two true values ​​are latent parameters, which are parameters that cannot be directly observed or calculated and need to be estimated by using algorithms.

[0030] S102. Determine the annotation type of the multiple facial expression data, and determine the facial expression differences of the control group based on the annotation type.

[0031] In one embodiment, the above α is obtained by the maximum likelihood method. i and α jThese two latent parameters are estimated. During the estimation process, for all facial expression annotations, an externally sourced Facial ActionUnit (FAUnit) recognition algorithm is used to identify changes in the facial action units within the expression data. An objective, autonomous judgment algorithm is then established based on fundamental emotion theories in the field of emotion psychology. Specifically, the algorithm performs image recognition on the expression data to obtain changes in eye muscle groups such as the orbicularis oculi and supercilium brevis. Contraction of the orbicularis oculi and supercilium brevis is defined as negative; contraction of the depressor supercilium is defined as positive; the greater the deformation amplitude of the relevant muscle group, the higher the degree of negative change. The activity of the mouth muscle groups is also obtained. Contraction of the pronator labii superioris is defined as positive; contraction of the depressor labii superioris is defined as negative; the greater the deformation amplitude of the relevant muscle group, the higher the degree of negative change.

[0032] All facial expression annotations were used to perform pairwise comparisons between two presented facial expression images, comparing the differences between the two images on two dimensions: positivity and arousal. The resulting value is n. ij That is, in terms of positivity or arousal level, the image at index i is greater than the image at index j; n ji This means the image with index j is greater than the image with index i. The label for all facial expression data annotations can be described as:

[0033] Label = {1 12 ,2 12 ,3 21 ,…,b ij}

[0034] For example, if a total of 10 comparisons are made, and the expression level of image i is greater than that of image j in 8 of those comparisons, then n ij =8, and this 8 is the result of the total score.

[0035] For all the facial expression annotation results mentioned above, if a total of n annotations were performed, then N is the sum of all values ​​of n. The total facial expression annotation results N can be further described as follows:

[0036]

[0037] In one embodiment, when estimating all facial expression annotation results N using the maximum likelihood method, the logarithm of the likelihood function L is required to satisfy:

[0038]

[0039] Where ln L represents the annotation result, and n ijThe first facial expression data group represents the data group whose facial expression intensity is greater than that of the second facial expression data group. n refers to the total number of comparisons between all facial expression images i and j during the annotation process, i.e., the total number of annotations as mentioned earlier, n times.

[0040] S103. Mark the facial expressions of the control group according to the differences in facial expressions.

[0041] The first facial expression data is determined as the basic facial expression. The level of positivity and arousal of the first facial expression data is standardized, i.e., α. v =1, α a =1 is used as a baseline, i.e., the basic result. The attributes of the second or other facial expression data compared to the first expression data are adjusted to maximize the logarithm of the likelihood function L of the actual N pairs of comparison results. This yields estimates of the positivity and arousal levels of other facial expressions, used to rank facial expressions in terms of positivity and arousal. This is achieved by pairwise comparisons of multiple control groups, pairing facial expression i and facial expression j images, and counting the number of times i is greater than j or less than j. Then, α is assumed... v =1, α s When 1 is used as a baseline, the absolute values ​​of all facial expression images on V and A can be obtained.

[0042] In one embodiment, in addition to the first and second expression data, other expression data can be selected from multiple expression data sets and grouped with the first expression data to form a new control group. The other expression data can then be labeled using the method described above.

[0043] In one embodiment, for example, three facial expression images are defined as image A, image B, and image C, and the annotator annotates these images.

[0044] The annotation results for the positiveness (V) of facial expressions in the images are shown in the table below:

[0045] Annotation results Greater than Less than A and B 8 4 A and C 3 5 B and C 2 3

[0046] The annotation results of facial expression arousal level (A) in the images are shown in the table below:

[0047]

[0048]

[0049] As shown in the diagram above, comparing facial expression data A and B, the number of times A is greater than B is 8, which is n. AB =8, the number of times A is less than B is 4, which is n BA =4, and then the maximum likelihood method is used for estimation.

[0050] Normalize the actual value of the facial expression positivity (V) of Picture A, i.e., α A-v = 1, and take this value as the reference value, then:

[0051]

[0052] After maximizing the value of lnL, we get α B-v = 0.59, α C-v = 1.32.

[0053] In one embodiment, since the facial expression images need to be paired in pairs, all N facial expression images paired will generate N(N - 1) paired data, with extremely high costs, and it is difficult to label so much data in practice. Therefore, a small-world network (herein called the expression data network) is constructed for pairing. The small-world network regards all facial expression data as a sequence and constructs this sequence into a loop. n expression data are used as the n vertices of this small-world network, each vertex has k edges, and the vertices and edges form a network loop lattice. Each edge is randomly connected with a preset probability p. This structure allows adjusting the graph between regularity (p = 0) and disorder (p = 1), so as to take values between 0 < p < 1. After experiments, when p takes the value of 0.06, a small-world network can be generated fastest, which can minimize the labeling cost of the labeler and meet the requirement of the maximum connectivity of the pairing for the labeling task. After forming the small-world network, the facial expression pairing can connect all facial expression data sets as much as possible with as few pairings as possible, so as to establish a data set that meets the labeling requirements. Select the first node and the second node in the expression data network, and establish a control group according to the expression data corresponding to the first node and the expression data corresponding to the second node, so as to perform expression labeling on the control group.

[0054] In one embodiment, during the process of the labeler performing video labeling, a sensor, such as an eye tracker, is used to track the cognitive information processing process. Define the facial expression region ROA in the image, and define the eye movement fixation region as ROI. Collect the fixation time of the labeler through the eye tracker, determine the preset fixation threshold, compare the fixation time with the fixation threshold. If the fixation time is greater than the fixation threshold, then determine that the fixation time is the effective time, and the eye movement fixation region ROI is the region where the effective time exceeds 3s. If the proportion of the overlapping time T1 between the eye movement fixation region ROI and the facial expression region ROA in the total time T is the degree of earnestness. After the labeler finishes labeling a group of paired images, automatically add a field to the labeled data, which is the earnestness index. The higher the earnestness index, the higher the accuracy of the labeler.

[0055] Such as Figure 2As shown in the illustration, this application also provides a facial expression annotation device, including:

[0056] At least one processor; and,

[0057] A memory that is communicatively connected to at least one processor; wherein,

[0058] The memory stores instructions that can be executed by at least one processor to enable a facial expression annotation device to perform:

[0059] Multiple facial expression data are determined, and a control group is determined from the multiple facial expression data, wherein the control group includes first facial expression data and second facial expression data;

[0060] Determine the annotation type of the multiple facial expression data, and determine the facial expression differences of the control group based on the annotation type;

[0061] The control group was labeled with facial expressions based on the differences in facial expressions.

[0062] This application embodiment also provides a non-volatile computer storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured as follows:

[0063] Multiple facial expression data are determined, and a control group is determined from the multiple facial expression data, wherein the control group includes first facial expression data and second facial expression data;

[0064] Determine the annotation type of the multiple facial expression data, and determine the facial expression differences of the control group based on the annotation type;

[0065] The control group was labeled with facial expressions based on the differences in facial expressions.

[0066] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should understand that by simply performing some logic programming on the method flow using one of these hardware description languages ​​and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.

[0067] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0068] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.

[0069] For ease of description, the above devices are described in terms of function, divided into various units. Of course, in implementing this specification, the functions of each unit can be implemented in one or more software and / or hardware components.

[0070] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device and medium embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the description of the method embodiments.

[0071] The devices and media provided in this application are one-to-one with the methods. Therefore, the devices and media also have similar beneficial technical effects as their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.

[0072] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0073] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0074] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0075] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0076] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0077] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0078] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0079] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0080] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for annotating facial expressions, characterized in that, include: Multiple facial expression data are determined, and a control group is determined from the multiple facial expression data, wherein the control group includes first facial expression data and second facial expression data; Determine the annotation type of the multiple facial expression data, and determine the facial expression differences of the control group based on the annotation type; The control group was labeled with facial expressions based on the differences in facial expressions. The first expression data is determined as the basic expression, and the basic expression is labeled as the basic result; The second expression data is labeled with expressions based on the basic results and the expression differences. Determine the probability of difference in the control group, wherein the probability of difference is calculated using the following formula: in, Let be the probability that the intensity of the first facial expression is greater than that of the second facial expression, in terms of either arousal or positivity. This probability follows a log-odds function. This represents the first expression data. This represents the second expression data. The primary expression reflects the degree of positivity or arousal in the expression. The degree of expression of the second facial expression in terms of its positivity or arousal. The hidden parameter for the first expression. This is a hidden parameter for the second expression; The annotation result is calculated based on the difference probability, and annotation is performed based on the annotation result, wherein the calculation formula for the annotation result is: in, For the annotation results, The data group whose expression level of the first expression data is greater than that of the second expression data.

2. The method according to claim 1, characterized in that, The method further includes: Other facial expression data are determined based on the multiple facial expression data, and a new control group is determined based on the other facial expression data and the first facial expression data; The facial expression differences of the new control group are determined, and the other facial expression data are labeled based on the facial expression differences and the baseline results.

3. The method according to claim 2, characterized in that, The method further includes: Determine the expression result of the second expression data, and determine the annotation result of the other expression data; The multiple expression data are sorted based on the basic results, the expression results of the second expression data, and the annotation results of the other expression data.

4. The method according to claim 1, characterized in that, The method further includes: An expression data network is established based on the multiple expression data, and the multiple expression data are used as nodes in the expression data network; A first node and a second node in the facial expression data network are determined, and a control group is established based on the facial expression data corresponding to the first node and the facial expression data corresponding to the second node, thereby performing facial expression annotation on the control group.

5. The method according to claim 1, characterized in that, The method further includes: The system uses sensors to determine the gaze duration of the annotator on the facial expression data, determines a pre-set gaze threshold, and compares the gaze duration with the gaze threshold. If the fixation time is greater than the fixation threshold, then the fixation time is determined to be a valid time. The accuracy of the annotator is determined based on the effective time, and the facial expression data is annotated based on the accuracy.

6. A facial expression annotation device, characterized in that, include: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the facial expression annotation device to perform the following: Multiple facial expression data are determined, and a control group is determined from the multiple facial expression data, wherein the control group includes first facial expression data and second facial expression data; Determine the annotation type of the multiple facial expression data, and determine the facial expression differences of the control group based on the annotation type; The control group was labeled with facial expressions based on the differences in facial expressions. The first expression data is determined as the basic expression, and the basic expression is labeled as the basic result; The second expression data is labeled with expressions based on the basic results and the expression differences. Determine the probability of difference in the control group, wherein the probability of difference is calculated using the following formula: in, Let be the probability that the intensity of the first facial expression is greater than that of the second facial expression, in terms of either arousal or positivity. This probability follows a log-odds function. This represents the first expression data. This represents the second expression data. The primary expression reflects the degree of positivity or arousal in the expression. The degree of expression of the second facial expression in terms of its positivity or arousal. The hidden parameter for the first expression. This is a hidden parameter for the second expression; The annotation result is calculated based on the difference probability, and annotation is performed based on the annotation result, wherein the calculation formula for the annotation result is: in, For the annotation results, The data group whose expression level of the first expression data is greater than that of the second expression data.

7. A non-volatile computer storage medium storing computer-executable instructions, characterized in that, The computer-executable instructions are set as follows: Multiple facial expression data are determined, and a control group is determined from the multiple facial expression data, wherein the control group includes first facial expression data and second facial expression data; Determine the annotation type of the multiple facial expression data, and determine the facial expression differences of the control group based on the annotation type; The control group was labeled with facial expressions based on the differences in facial expressions. The first expression data is determined as the basic expression, and the basic expression is labeled as the basic result; The second expression data is labeled with expressions based on the basic results and the expression differences. Determine the probability of difference in the control group, wherein the probability of difference is calculated using the following formula: in, Let be the probability that the intensity of the first facial expression is greater than that of the second facial expression, in terms of either arousal or positivity. This probability follows a log-odds function. This represents the first expression data. This represents the second expression data. The primary expression reflects the degree of positivity or arousal in the expression. The degree of expression of the second facial expression in terms of its positivity or arousal. The hidden parameter for the first expression. This is a hidden parameter for the second expression; The annotation result is calculated based on the difference probability, and annotation is performed based on the annotation result, wherein the calculation formula for the annotation result is: in, For the annotation results, The data group whose expression level of the first expression data is greater than that of the second expression data.