An AI Dynamic Emotion Recognition Multi-Dimensional Oriented Training System

Through the AI ​​dynamic emotion recognition multi-dimensional directional training system, the multi-dimensional emotion detection and directional training methods are used to solve the problem of rapid and non-contact detection and intervention in bad emotions in the existing technology, real-time monitoring and regulation of people's emotions and psychological states is achieved.

CN119181119BActive Publication Date: 2025-07-01NINGBO XINYI TECH CO LTD

Patent Information

Application Number
CN202411207642.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-30
Publication Date
2025-07-01
Estimated Expiration
2044-08-30

AI Technical Summary

Technical Problem

The prior art is difficult to quickly and contactlessly detect the dynamic emotions and psychological state of a person, and lacks effective means of real-time monitoring and adverse emotional intervention.

Method used

A multi-dimensional directional training system for dynamic emotions recognition was designed. Through the cooperation of multiple emotions acquisition and training terminals and background servers, contactless face image acquisition and emotional data analysis are realized. The system can detect the emotional state of the person from multiple dimensions and push targeted training tasks based on the detection results.

Benefits of technology

It realizes rapid screening and real-time monitoring of the dynamic psychological and emotional state of people, which can quickly intervene and regulate bad emotions, improve mental health, and avoid sudden malignant events on a large scale.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119181119B_ABST
    Figure CN119181119B_ABST
Patent Text Reader

Abstract

The present invention proposes an AI dynamic emotion recognition multi-dimensional targeted training system, which relates to the fields of computer vision technology and artificial intelligence technology. It includes multiple emotion collection and training terminals at the front end and a background server. The emotion collection and training terminals are used to detect the emotions of the tested personnel. The terminals analyze the face images to obtain the multi-dimensional emotion data of the tested personnel. The multi-dimensional emotion data is uploaded to the background server. The background server analyzes the current emotion state of the tested personnel according to the received multi-dimensional emotion data, and pushes different targeted training tasks to the corresponding emotion collection and training terminals for the tested personnel in different states. After receiving the targeted training tasks, the emotion collection and training terminals, the tested personnel complete the corresponding training content according to the guidance. The present invention can specifically improve the psychological quality of the tested personnel or improve the bad emotion state.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention provides an AI dynamic emotion recognition multi-dimensional orientation training system, which relates to the fields of computer vision technology and artificial intelligence technology. Background Art

[0002] In the current social environment, people have a fast-paced life and work, and are under great pressure. Affected by factors such as the working environment, life stress, and social relations, it is easy for people to have bad emotions, such as irritability, restlessness, depression, negativity, etc. Bad emotions will have an impact on people's physical health and mental state, and may even trigger serious harmful events such as violent conflicts and suicides. If the bad emotional state of people can be quickly recognized and timely intervention and guidance are carried out, the bad psychological state and sudden emotional problems of people can be improved targeted, sudden events can be predicted in advance or even avoided, and the possibility of causing serious consequences can be reduced.

[0003] Traditional emotion assessment methods often require face-to-face diagnosis by a psychologist or filling out a specific scale for evaluation, both of which require a long time and are not applicable to scenarios that require large-scale and rapid screening. And most emotion intervention means require a venue, props or on-site guidance by a psychologist, need to be actively cooperated by personnel, and are not timely. Among the existing new technologies, there are related technical means for psychological assessment by analyzing information such as micro-expressions and micro-actions through video or using voice characteristics. Non-contact detection can be achieved, but the existing technologies generally require the tested person to stay still in front of the camera for about 1-3 minutes, or answer certain specific questions for voiceprint collection. The cooperation of the tested person is still required. A certain collection time is needed and real-time monitoring cannot be achieved, and sudden emotional problems are difficult to detect in time. And the existing technologies lack supporting means for intervention and treatment after detection, and it is difficult to fundamentally adjust the emotions and psychological state of the tested person. Summary of the Invention

[0004] To overcome the problems existing in the related technologies, the present application provides an AI dynamic emotion recognition multi-dimensional orientation training system, which can quickly and non-contact detect the dynamic emotions and psychological states of personnel from multiple dimensions, propose relevant orientation training and guidance content for different states of the tested person, and then improve the bad emotional state and improve the mental health level. Realize rapid emotion screening, dynamic emotion and mental health regulation in large-scale scenarios, and analyze, display and information warning of data.

[0005] An AI dynamic emotion recognition multi-dimensional orientation training system includes: multiple emotion collection and training terminals at the front end and a background server;

[0006] The emotion collection and training terminal is used to detect the emotions of the person to be tested. The terminal analyzes the face image to obtain multi-dimensional emotion data of the person to be tested;

[0007] The multi-dimensional emotion data is uploaded to the background server. The background server analyzes the current emotion state of the person to be tested according to the received multi-dimensional emotion data, and pushes different targeted training tasks to the corresponding emotion collection and training terminals for different state persons to be tested;

[0008] After the emotion collection and training terminal receives the targeted training task, the person to be tested completes the corresponding training content according to the guidance.

[0009] Further, the emotion collection and training terminal includes a camera, a display, an emotion data calculation module, a control and data transmission module;

[0010] The camera is used to collect the face image of the person to be tested;

[0011] The display is used to display the picture collected by the camera and the targeted training content;

[0012] The emotion data calculation module identifies the feature points of the face to obtain the dynamic emotion feature data of the person to be tested within a fixed time;

[0013] The control and data transmission module uploads the dynamic emotion feature data to the background server and receives the targeted training content and control instructions returned by the background server.

[0014] Further, the emotion data calculation module identifies the feature points of the face, obtains the positions of the face feature points in each frame image and the relative changes between frames, represents the change law of the face feature points in the form of a spatial vector to form a face feature change matrix; inputs the face feature change matrix into an emotion recognition feature model for calculation to obtain the emotion feature data of the person to be tested.

[0015] Further, a coordinate system is established with the lower left corner pixel point of the face image as the origin, and the coordinate value of the nth feature point in the face image is: (x n ,y n );

[0016] The coordinate values of the feature points are normalized to obtain the relative position coordinates (x zn ,y zn ) of the nth feature point:

[0017]

[0018]

[0019] Among them, xzn , y zn are the normalized horizontal and vertical coordinate values; mean(x) and mean(y) are the average values of the horizontal and vertical coordinates of all feature points; std(x) and std(y) are the standard deviations of the horizontal and vertical coordinates of all feature points.

[0020] Furthermore, the relative position coordinates of 78 feature points of the current frame image are selected to form a 156-dimensional feature point space vector Ve:

[0021] Ve = {(x z1 , y z1 ), (x z2 , y z2 ), …, (x z78 , y z78 )}

[0022] For the previous frame image, the same processing is performed to obtain the feature point space vector V e ';

[0023] Subtract the feature point space vectors of the front and rear frames to obtain the face feature change matrix V c :

[0024] V c = V e ' - V e = {(x z1 ' - x z1 , y z1 ' - y z1 ), …, (x z78 ' - x z78 , y z78 ' - y z78 )};

[0025] Input the face feature change matrix V c into the emotion recognition feature model for calculation to obtain the emotion feature data of the person being tested.

[0026] Furthermore, the emotion feature recognition model includes a spatial feature model, a change feature model, and a regional feature model:

[0027] The input of the spatial feature model is: the relative position coordinate matrix Ve of the facial feature points extracted from each frame image; the output is: the spatial feature emotion label

[0028] The input of the change feature model is: the change matrix V c of the face features between each frame image and its previous frame image; the output is: the change feature emotion label

[0029] The input of the regional feature model is: the regional images obtained by segmenting each frame of image according to the distribution of facial features, including the left eye image I eyel , the right eye image I eyer , the nose image I nos , the left cheek image I facl , the right cheek image I facr , the mouth image I mou ;

[0030] At the feature layer of the regional feature model, the output features of each segmentation image model network are concatenated to obtain the concatenation result ConFea:

[0031] ConFea = [F(I eyel ), F(I eyer ), F(I nos ), F(I facl ), F(I facr ), F(I mou )]

[0032] where F(·) is the shallow convolutional neural network of each regional image;

[0033] Output the regional feature emotion label

[0034]

[0035] where FC(·) is the fully connected layer network and softmax(·) is the normalized exponential function.

[0036] Furthermore, through weighted cumulative calculation, the emotion label of the current frame of the picture is obtained

[0037]

[0038] where W s , W c , W a are the spatial feature weight value, the change feature weight value and the regional feature weight value respectively.

[0039] Furthermore, the emotion data calculation module collects multiple frames of face images within a period of time, obtains multiple emotion labels corresponding to the multiple frames through the emotion recognition feature model, accumulates the multiple emotion labels and takes the average value to obtain the dynamic emotion feature data of the measured person during this period of time.

[0040] Furthermore, the emotion label is a collection of 12-dimensional emotion feature parameters, represented in the form of a 12-dimensional vector, which respectively represent: aggression, stress, tension, suspiciousness, balance, confidence, energy, self-regulation, depression, neuroticism, inhibition, happiness.

[0041] The technical solution provided by this application may include the following beneficial effects: This application can quickly screen the dynamic psychological and emotional states of personnel in a non-contact manner, quickly intervene and conduct targeted training for bad emotions and psychological problems, and promptly give early warnings for abnormal emotions. It can quickly adjust the psychological level and emotional state of personnel on a large scale, and to a certain extent, avoid sudden vicious events caused by bad emotions. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0043] Figure 1 is the overall architecture diagram of the AI dynamic emotion recognition multi-dimensional targeted training system of the present invention;

[0044] Figure 2 is the front structure schematic diagram of the emotion collection and training terminal of the present invention;

[0045] Figure 3 is the back structure schematic diagram of the emotion collection and training terminal of the present invention;

[0046] Figure 4 is the working process schematic diagram of the AI dynamic emotion recognition multi-dimensional targeted training system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the drawings in the embodiments of this application. Obviously, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of this application without creative efforts fall within the scope of protection of this application.

[0048] In the attached drawings of the specific embodiments of the present invention, in order to better and more clearly describe the working principles of the components in the system and show the connection relationships of the various parts of the device, only the relative positional relationships between the components are clearly distinguished, and it does not constitute a limitation on the signal transmission direction, connection sequence, and the sizes, shapes of the various parts of the structure within the component or structure.

[0049] Figure 1It is the overall architecture diagram of an AI dynamic emotion recognition multi-dimensional orientation training system of the present invention. The AI dynamic emotion recognition multi-dimensional orientation training system includes multiple emotion collection and training terminals at the front end and a background server. A local area network is formed between the multiple emotion collection and training terminals at the front end and the background server through an Ethernet and network switching devices to transmit data and control information.

[0050] The background server is internally equipped with data service and management programs, which are used to analyze and process the emotion feature data uploaded by the emotion collection and training terminals, analyze the current emotion state of the measured person according to the score distribution of different dimensions in the emotion feature data, generate corresponding orientation training content according to the analysis results, push the orientation training content to the emotion collection and training terminals, and display it through a display to guide the measured person to complete the training task.

[0051] As Figure 2 and Figure 3 shown, they are respectively the front and back structure schematic diagrams of the emotion collection and training terminal. The emotion collection and training terminal includes a housing 1, a camera 2, a display 3, an emotion data calculation module 4, and a control and data transmission module 5.

[0052] Among them, the camera 2 is installed at the front end of the housing, facing the direction of the measured person directly, and is used to collect the face image of the measured person.

[0053] The display 3 is installed on the front of the housing and is used to display the images collected by the camera and the orientation training content.

[0054] The emotion data calculation module 4 and the control and data transmission module 5 are installed inside the same host, and can be composed of devices with data access and calculation functions such as tablet computers and computer hosts and the application software they carry.

[0055] The emotion data calculation module 4 identifies the feature points of the face, obtains the positions of the face feature points in each frame of the image and the relative changes between frames, represents the change law of the face feature points in the form of a spatial vector, and forms a face feature change matrix.

[0056] Specifically, for each frame of the collected image, through a face detection algorithm, a rectangular face area in the image is obtained. Let this area be a rectangular image area with X pixel points in the horizontal direction and Y pixel points in the vertical direction. Using the face area as the detection object, the positions of the face feature points are detected.

[0057] First, the positions of the facial contour feature points are detected, a deep convolutional neural network is constructed, the input of the deep convolutional neural network is the face area image, and the output is the positions of the facial contour feature points. A total of 17 facial contour feature points are taken. After obtaining the facial contour feature points, the face is divided into six regions: the left eye, the right eye, the nose, the left cheek, the right cheek, and the mouth according to the positions of the facial features.

[0058] Six different deep neural networks are constructed. The input of the deep neural network is the cropped image of each region (six regions: left eye, right eye, nose, left cheek, right cheek, and mouth), and the output is the position coordinates of the facial contour feature points within the region. For each eye region, 5 feature points on the eyebrows and 6 feature points on the eyes are taken; 5 feature points are taken for each cheek region; 9 feature points are taken for the nose; and 20 feature points are taken for the mouth. There are a total of 61 facial feature points of the facial features.

[0059] For each frame of the facial image, after feature point detection, a total of 78 position coordinate values of facial feature points are obtained, and each feature point is numbered from 1 to 78. Taking the lower left pixel point of the facial image as the origin to establish a coordinate system, the coordinate value of the nth feature point in the facial image is: (x n , y n );

[0060] In order to represent the relative positions between the facial feature points, the Z-socore method is used to normalize the coordinate values of the feature points.

[0061]

[0062] Among them, x zn , y zn are the normalized horizontal and vertical coordinate values, and the relative position coordinates (x zn , y zn ) of the nth feature point are obtained;

[0063] mean(x) and mean(y) are the average values of the horizontal and vertical coordinates of all feature points; std(x) and std(y) are the standard deviations of the horizontal and vertical coordinates of all feature points.

[0064] The relative position coordinates of the 78 feature points of the current frame image are used to form a 156-dimensional feature point space vector Ve:

[0065] Ve = {(x z1 , y z1 ), (x z2 , y z2 ), …, (x z78 , y z78 )}

[0066] For the previous frame image, the same processing is performed to obtain the feature point space vector V e ';

[0067] Subtracting the feature point space vectors of the front and back frames can obtain the facial feature change matrix V c .

[0068] V c = Ve '-V e = {(x z1 '-x z1 , y z1 '-y z1 ),…, (x z78 '-x z78 , y z78 '-y z78 )}

[0069] Input the obtained face feature change matrix V c into the emotion recognition feature model for calculation, and then obtain the emotion feature data of the person. The emotion data calculation module is built-in with an emotion recognition feature model obtained through artificial intelligence algorithms. After being verified by annotation from a large number of face images with different emotion features, it is input into the model for training.

[0070] Specifically, the emotion feature recognition model includes three sub-models: a spatial feature model, a change feature model, and a region feature model. The input and output corresponding to each model are as follows:

[0071] (1) Spatial feature model:

[0072] Input: The relative position coordinate matrix Ve of the facial feature points extracted from each frame of the image;

[0073] Output: Spatial feature emotion label

[0074] (2) Change feature model:

[0075] Input: The change matrix V of the face features between each frame of the image and its previous frame of the image c ;

[0076] Output: Change feature emotion label

[0077] (3) Region feature model:

[0078] Input: The region images obtained by dividing each frame of the image according to the distribution of the facial features, including the left eye image I eyel , the right eye image I eyer , the nose image I nos , the left cheek image I facl , the right cheek image I facr , the mouth image I mou

[0079] At the feature level of the model, splice the model network output features of each segmented image:

[0080] ConFea = [F(I eyel ), F(I eyer ), F(Inos ), F(I facl ), F(I facr ), F(I mou )]

[0081] Among them, F(·) is the shallow convolutional neural network of each regional image;

[0082] Then, the regional feature emotion label is obtained through feature splicing:

[0083]

[0084] Among them, FC(·) is the fully connected layer network, and softmax(·) is the normalization exponential function.

[0085] Output: Regional feature emotion label

[0086] After obtaining the emotion labels output by each sub-model, through weighted accumulation, the emotion label of the current frame of the picture is finally obtained

[0087]

[0088] Among them, W s , W c , W a are the spatial feature weight value, the change feature weight value, and the regional feature weight value respectively.

[0089] W s + W c + W a = 1.

[0090] The above-mentioned emotion labels are a collection of 12-dimensional emotion feature parameters, represented in the form of a 12-dimensional vector. Each dimension is represented by an integer score from 0 to 100 to indicate the intensity of that emotion dimension. Respectively represent: aggression, stress, tension, suspiciousness, balance, confidence, energy, self-regulation, depression, neuroticism, inhibition, happiness.

[0091] The emotion data calculation module collects face images within a period of time, infers and obtains the emotion labels of each frame of the image through the above process, accumulates the emotion labels and then takes the average value to obtain the dynamic emotion feature data of the person being measured during that period of time. Also represented in the form of a 12-dimensional vector.

[0092] The control and data transmission module 5 controls the overall working process of the terminal, uploads the dynamic emotion feature data to the background server, receives the targeted training content and control instructions returned by the background server, and controls the display to display the training content on the current screen to guide the person being measured to perform the targeted emotion training task.

[0093] For example, the emotional characteristic data is multi-dimensional data represented by scores from 0 to 100, denoted by E1 - E 12 The higher the score, the deeper the degree of emotion or psychological state in this dimension. The specific meanings represented by the data in each dimension are as follows:

[0094] E1: Aggressiveness. When E1 > 70, it indicates that the person being tested is in a state of irritability and anger;

[0095] E2: Stress. When E2 > 60, it indicates that the person being tested is at a relatively high level of stress;

[0096] E3: Tension. When E3 > 70, it indicates that the person being tested is nervous and has a sense of anxiety;

[0097] E4: Suspiciousness. When E4 > 70, it indicates that the person being tested lacks a sense of security;

[0098] E5: Balance. When E5 < 40, it indicates that the person being tested is in a state of imbalance between body and mind;

[0099] E6: Confidence. When E6 < 40, it indicates that the person being tested lacks confidence and has a low self - recognition;

[0100] E7: Energy. When E7 < 50, it indicates that the person being tested lacks vitality and is physically exhausted;

[0101] E8: Self - regulation. When E8 < 40, it indicates that the person being tested has low self - control.

[0102] E9: Depression. When E9 > 60, it indicates that the person being tested is in a state of depression and loss;

[0103] E 10 : Neuroticism. When E 10 > 80, it indicates that the person being tested has an obsessive - compulsive disorder tendency;

[0104] E 11 : Inhibition. When E 11 > 70, it indicates that the person being tested has a state of depression and self - enclosure;

[0105] E 12 : Happiness. When E 12 < 40, it indicates that the person being tested has a low happiness level and is not satisfied with their life state.

[0106] The content of the orientation training can be multi - faceted, including encouraging words, physical exercise programs, communication tasks, reading, listening to music, etc. When the psychological and emotional data of the person being tested is in a relatively stable and healthy state, the training content can also be sharing joy, helping others, rewards, etc.

[0107] Table 1 exemplarily lists the content of targeted training corresponding to different emotional feature data states. In actual applications, various different targeted training contents can be added as needed.

[0108]

[0109]

[0110] Figure 4 It is a schematic diagram of the working process of an AI dynamic emotion recognition multi-dimensional targeted training system shown in the embodiments of the present application.

[0111] See Figure 4 The system working process is as follows:

[0112] The person to be tested stands in front of the emotion collection and training terminal and starts emotion detection;

[0113] The terminal collects face images of the person for 5 seconds through the camera. During the collection process, the display prompts the collection time and the position of the face in the picture to prevent the person to be tested from leaving early or the standing position from shifting;

[0114] Through face recognition, the identity information of the person to be tested is obtained;

[0115] The terminal analyzes the face image to obtain the multi-dimensional emotion data of the person to be tested;

[0116] Upload the identity information and multi-dimensional emotion data of the person to be tested to the background server;

[0117] The background server receives the multi-dimensional emotion data of the person, sorts and stores the data, and can also display the existing data in the form of charts, etc. as needed;

[0118] The background server analyzes the current psychological emotion state of the person to be tested according to the received multi-dimensional emotion data, and pushes different targeted training tasks to the emotion collection and training terminal for different state personnel.

[0119] After the terminal receives the targeted training task, it displays the detailed requirements of the task on the display in the form of a text pop-up window;

[0120] The person to be tested completes the corresponding training content according to the text guidance on the display.

[0121] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or a data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)), etc.

[0122] As described above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. An AI dynamic emotion recognition multi-dimensional directional training system, characterized in that: include: Multiple emotion collection and training terminals and backend servers at the front end; the emotion collection and training terminals are used to perform emotion detection on the tested persons, and the terminals analyze facial images to obtain multi-dimensional emotion data of the tested persons; the multi-dimensional emotion data are uploaded to the backend server, and the backend server analyzes the current emotional state of the tested persons based on the received multi-dimensional emotion data, and pushes different directional training tasks to the corresponding emotion collection and training terminals for the tested persons in different states; after the emotion collection and training terminals receive the directional training tasks, the tested persons complete the corresponding training content according to the instructions; The emotion collection and training terminal includes an emotion data calculation module, which identifies feature points for a face, obtains the positions of the face feature points in each frame image and the relative changes between frames, and represents the change rules of the face feature points in the form of spatial vectors to form a face feature change matrix; Inputting the facial feature change matrix into the emotion recognition feature model for calculation to obtain the emotion feature data of the person to be tested; The emotion recognition feature model includes a spatial feature model, a change feature model, and a regional feature model: The input of the spatial feature model is: the relative position coordinate matrix V of the facial feature points extracted from each frame image e ; Output: spatial feature emotion label The input of the change feature model is: the face feature change matrix V between each frame image and its previous frame image c ; Output: Change feature emotion label The input of the regional feature model is: the regional images of each frame image segmented according to the distribution of facial features, including the left eye image I eyel , right eye image I eyer , nose image I nos , left cheek image I facl , right cheek image I facr , mouth image I mou ; In the feature layer of the regional feature model, the network output features of each segmentation image model are spliced ​​to obtain the splicing result ConFea: ConFea=F(I eyel ),F(I eyer ),F(I nos ),F(I facl ),F(I facr ),F(I mou )] Where F(·) is the shallow convolutional neural network of each region image; The output region feature emotion label is Among them, FC(·) is a fully connected layer network, and softmax(·) is a normalized exponential function.

2. According to the AI ​​dynamic emotion recognition multi-dimensional directional training system described in claim 1, it is characterized in that: The emotion collection and training terminal also includes a camera, a display, and a control and data transmission module; the camera is used to collect facial images of the person being tested; the display is used to display the images collected by the camera and the directional training content; the control and data transmission module uploads the dynamic emotion feature data to the background server, and receives the directional training content and control instructions returned by the background server.

3. According to the AI ​​dynamic emotion recognition multi-dimensional directional training system described in claim 1, it is characterized in that: The coordinate system is established with the pixel point at the lower left corner of the face image as the origin. The coordinate value of the nth feature point in the face image is: (x n ,y n ); The coordinate values ​​of the feature points are standardized to obtain the relative position coordinates (x zn ,y zn ): Among them, x zn ,y zn are the normalized horizontal and vertical coordinate values; mean(x) and mean(y) are the average values ​​of the horizontal and vertical coordinates of all feature points; std(x) and std(y) are the standard deviations of the horizontal and vertical coordinates of all feature points.

4. According to the AI ​​dynamic emotion recognition multi-dimensional directional training system described in claim 3, it is characterized in that: The relative position coordinates of the 78 feature points of the current frame image are selected to form a 156-dimensional facial feature point relative position coordinate matrix V e : V e ={(x z1 ,and z1 ),(x z2 ,and z2 ),…,(x z78 ,and z78 )} For the previous frame image, the relative position coordinate matrix V of the facial feature points is obtained by the same processing e '; Subtract the feature point space vectors of the previous and next frames to obtain the face feature change matrix V c : V c =V e '-V e ={(x z1 '-x z1 ,and z1 '-and z1 ),…,(x z78 '-x z78 ,and z78 '-and z78 )}; The face feature change matrix V c The emotion recognition feature model is input for calculation to obtain the emotion feature data of the person being tested.

5. The AI ​​dynamic emotion recognition multi-dimensional directional training system according to claim 1, characterized in that: Through weighted cumulative calculation, the emotion label of the current frame is obtained Among them, W s , W c , W a They are spatial feature weight value, change feature weight value and regional feature weight value respectively.

6. The AI ​​dynamic emotion recognition multi-dimensional directional training system according to claim 5, characterized in that: The emotion data calculation module collects multiple frames of facial images within a period of time, obtains multiple emotion tags corresponding to the multiple frames through the emotion recognition feature model, accumulates the multiple emotion tags and takes the average value to obtain the dynamic emotion feature data of the person being tested during the period of time.

7. The AI ​​dynamic emotion recognition multi-dimensional directional training system according to claim 5, characterized in that: The emotion label is a collection of 12-dimensional emotion feature parameters, expressed in the form of 12-dimensional vectors, representing: aggressiveness, pressure, tension, suspicion, balance, self-confidence, energy, self-regulation, depression, neuroticism, inhibition, and happiness.

Citation Information

Patent Citations

  • Mobile terminal and method and apparatus for pushing video

    CN108989887A

  • Early warning method and device based on face recognition, electronic equipment and medium

    CN113887406A

Cited By

  • Drug addict relapse risk early warning system

    CN120661149A

  • Multi-dimensional early warning system for violent risks of drug addicts

    CN120746262A

  • Public appeal intelligent management and social risk prevention and control system

    CN120822819A