Mental state analysis method and smart mirror
By combining smart mirrors with contactless data collection and EEG data, and utilizing multimodal fusion and generation technologies, the method for analyzing mental states has been solved, achieving high-precision mental state assessment and adapting to the analysis needs of different user groups.
Patent Information
- Application Number
- CN202211571266.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-08
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2042-12-08
AI Technical Summary
Existing methods for analyzing mental state are not precise enough, and there are differences in training data performance among different ethnic groups and cultures, making it difficult to accurately assess a user's mental state.
By using a smart mirror combined with contactless data acquisition and EEG data, and through a multimodal fusion network and a conditional adversarial generative network, generative features that approximate multimodal fusion features are generated and used as input for a mental state assessment model, thereby improving the accuracy of analysis.
It can effectively improve the accuracy of mental state analysis without relying on EEG data, and adapt to the mental state assessment needs of different user groups.
Smart Images

Figure CN115969376B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method for analyzing mental states and a smart mirror. Background Technology
[0002] Mental state is a comprehensive state encompassing a person's feelings, thoughts, and behaviors, playing a vital role in interpersonal communication. Its influence is ubiquitous in people's daily work and life. In medical care, understanding a patient's mental state, especially those with communication difficulties, allows for tailored care and improved attention. In product development, recognizing a user's mental state during product use and understanding the user experience enables improvements to product functionality and the design of products better suited to user needs. In various human-computer interaction systems, the ability to recognize a person's mental state makes human-machine interaction more user-friendly and natural.
[0003] For human-computer interaction processes that employ the recognition of a person's mental state, the accuracy of the results of the interaction depends on the analysis of the user's mental state. Therefore, the accuracy of the specific mental state analysis method used plays a crucial role in this process.
[0004] In existing human-computer interaction products, developers of mental state assessment models often focus on users' blinking, facial expressions, heart rate, pulse, etc., to make mental judgments. The accuracy of mental judgments depends on a large amount of training data.
[0005] However, during the development of the mental state assessment model, the applicant discovered that the existing technological development approach has diminishing returns. Furthermore, significant differences exist in the performance of training data among different ethnic groups and cultures. Therefore, there is an urgent need to supplement the model with more relevant parameters to serve as the basis for mental health assessment.
[0006] Through extensive research, the applicant discovered that electroencephalography (EEG) information plays a crucial role in improving mental state analysis. Because EEG information directly reflects the activity of the central nervous system in a user, it possesses higher objectivity and, unlike blinking or facial expressions, is less easily faked. Therefore, by combining EEG information, the applicant proposed a method for mental state analysis to improve the accuracy of such analysis. Summary of the Invention
[0007] This invention provides a method for analyzing mental state and a smart mirror to address the problem of insufficient accuracy in existing methods for analyzing mental state, thereby effectively improving the accuracy of the overall mental state analysis results when using the method provided by this invention.
[0008] In a first aspect, embodiments of the present invention provide a mental state analysis method applied to a smart mirror kit, the smart mirror kit comprising a smart mirror body and a head-mounted device, including:
[0009] When a user wears the head-mounted device and faces the smart mirror body, the user's physiological data is acquired using the detector carried by the smart mirror body and the head-mounted device. The physiological data includes contactless data acquisition data and electroencephalogram (EEG) data.
[0010] After extracting features from the user's physiological data, input data including non-contact acquisition features and EEG features is obtained. The input data is then fed into the multimodal fusion network in the mental state assessment model to output multimodal fusion features.
[0011] The multimodal fusion features output by the multimodal fusion network are input into the mental state prediction model in the mental state assessment model, and the predicted mental state label is output.
[0012] Secondly, embodiments of the present invention provide another method for analyzing mental states, applied to a smart mirror, including:
[0013] When a user faces the smart mirror, the detector carried by the smart mirror is used to acquire the user's contactless data.
[0014] After extracting features from the user's contactless data collection data, contactless data collection features are obtained.
[0015] The output of the non-contact acquisition features processed by the encoder is spliced together and random noise is added as input to the generator of the conditional adversarial generative network in the mental state assessment model to generate features that approximate multimodal fusion features. The multimodal fusion features are features obtained by inputting input data including non-contact acquisition features and EEG features into the multimodal fusion network.
[0016] The generated features are input into the mental state prediction model in the mental state assessment model, and the predicted mental state label is output.
[0017] Thirdly, embodiments of the present invention provide a smart mirror, comprising:
[0018] The smart mirror consists of the main body, detector, and processor.
[0019] The detector is installed on the smart mirror body and is used to acquire contactless data from users facing the smart mirror body.
[0020] The processor is disposed within the smart mirror body, connected to the detector to obtain contactless acquisition data acquired by the detector, and used to execute the steps of any one of the methods in the first and second aspects described above.
[0021] Fourthly, embodiments of the present invention provide an electronic device, including: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the steps of any one of the methods in the first and second aspects described above.
[0022] Fifthly, embodiments of the present invention provide a storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any one of the methods in the first and second aspects described above.
[0023] The mental state analysis method provided by this invention combines contactless data collected without contact with the user with the user's EEG data as input for training a mental state assessment model. The mental state assessment model fuses data from these two types of data in different modalities to obtain multimodal fusion features, enabling the mental state assessment model to effectively incorporate EEG data to analyze the user's mental state and thus improve the accuracy of mental state analysis.
[0024] Furthermore, based on the aforementioned mental state analysis method, this invention also considers scenarios where there is a lack of or no support for collecting user EEG data. In such scenarios, a generator using a conditional adversarial generative network can be employed, and ordinary non-contact physiological data can be used to generate features that approximate multimodal fusion features (i.e., including EEG information). This ensures the accuracy of the output mental state analysis results even without incorporating EEG data. Attached Figure Description
[0025] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 This is a flowchart of a mental state analysis method according to an embodiment of the present invention;
[0027] Figure 2 This is a flowchart illustrating the training method of the mental state prediction model in a mental state analysis method according to an embodiment of the present invention.
[0028] Figure 3 A flowchart of a mental state analysis method according to another embodiment of the present invention;
[0029] Figure 4for Figure 3 A flowchart of the training method for the conditional adversarial generative network of the implementation method of the mental state analysis method;
[0030] Figure 5 This is a flowchart illustrating the training method of the overall mental state assessment model after incorporating a conditional adversarial generative network in a mental state analysis method according to an embodiment of the present invention.
[0031] Figure 6 for Figure 3 A flowchart illustrating the implementation of the mental state analysis method;
[0032] Figure 7 This is a schematic diagram of the principle of a smart mirror according to an embodiment of the present invention;
[0033] Figure 8 for Figure 7 An exploded view of the implementation method of the smart mirror;
[0034] Figure 9 for Figure 7 A schematic diagram of the structure of a smart mirror according to an implementation method;
[0035] Figure 10 This is a schematic block diagram of a mental state analysis device according to an embodiment of the present invention;
[0036] Figure 11 This is a schematic block diagram of a mental state analysis device according to another embodiment of the present invention;
[0037] Figure 12 This is a schematic diagram of the structure of an embodiment of the electronic device of the present invention. Detailed Implementation
[0038] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0039] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other.
[0040] This invention can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, elements, data structures, etc., that perform a specific task or implement a specific abstract data type. This invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0041] In this invention, terms such as "module," "device," and "system" refer to relevant entities applied to a computer, such as hardware, combinations of hardware and software, software, or software in execution. More specifically, for example, an element can be, but is not limited to, a process running on a processor, a processor, an object, an executable element, an execution thread, a program, and / or a computer. Furthermore, an application program or script running on a server, and the server itself, can also be an element. One or more elements may be in an execution process and / or thread, and elements may be localized on a single computer and / or distributed across two or more computers, and may be run on various computer-readable media. Elements can also communicate via local and / or remote processes based on signals having one or more data packets, for example, signals from data interacting with another element in a local system, a distributed system, and / or interacting with other systems via signals over a network of the Internet.
[0042] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising" or "including" include not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0043] Figure 1 The flowchart of a mental state analysis method according to an embodiment of the present invention is illustrated schematically. The mental state can include positive and negative mental states. Positive mental states can include, for example, a high level of emotional well-being, while negative mental states can include, for example, a high level of depression or anxiety. This method is applied to a smart mirror kit, which includes a smart mirror body and a head-mounted device. (Refer to...) Figure 1 The method includes the following steps:
[0044] Step S11: When the user wears the head-mounted device and faces the smart mirror body, the user's physiological data is obtained by using the detector carried by the smart mirror body and the head-mounted device. The physiological data includes contactless data collection and electroencephalogram (EEG) data.
[0045] Step S12: After extracting features from the user's physiological data, input data including contactless acquisition features and EEG features is obtained. The input data is then input into the multimodal fusion network in the mental state assessment model to output multimodal fusion features.
[0046] Step S13: Input the multimodal fusion features output by the multimodal fusion network into the mental state prediction model in the mental state assessment model, and output the predicted mental state label.
[0047] The physiological data to be acquired in step S11 includes contactless data acquisition and electroencephalogram (EEG) data. Contactless data acquisition refers to physiological data that can be collected without requiring the training subject to wear a device for collecting the corresponding physiological data. This data may include eye movement data, heart rate data, pulse data, and image data, specifically acquired using detectors carried on the smart mirror itself. Taking the required contactless data acquisition, including eye movement data, heart rate data, pulse data, and image data, as an example, eye movement data can be obtained by collecting the training subject's eye movement information using an eye tracker; heart rate data and pulse data can be obtained by monitoring echo information using millimeter-wave radar; and image data can be obtained by capturing images of the training subject using a miniature camera. Therefore, the detectors carried on the smart mirror itself may include an eye tracker, millimeter-wave radar, and a miniature camera to acquire the aforementioned contactless data acquisition. EEG data can be acquired using a head-mounted device; the specific acquisition method will not be elaborated here. It is understood that contactless data collection may also include other data. This embodiment only illustrates a few of the data used in this embodiment. The specific type, quantity, and collection method of the user's physiological data to be acquired will vary depending on the physiological data required by the specific mental state assessment model. This embodiment only uses one method as an example to illustrate the method of the present invention. This embodiment does not limit the specific type, quantity, and collection method of contactless data collection.
[0048] In step S12, after acquiring the user's physiological data, feature extraction processing needs to be performed on the contactless data collected from the user's physiological data to obtain input data for the mental state assessment model used in the mental state analysis method. Specifically, the feature extraction processing of the contactless data can be designed according to the different data types in the contactless data. Taking the contactless data collected earlier, which includes eye movement data, heart rate data, pulse data, and image data, as an example, the specific feature extraction processing may include removing the illumination effect from the pupil diameter in the collected eye movement data and extracting eye movement statistical features to obtain feature data describing the eye-related state, such as the pupil diameter of both eyes, fixation time, saccade time, and blink frequency of the training subject. X EYE The collected heart rate and pulse data are denoised and analyzed, and the velocity and relative position of the training object are calculated through its frequency shift, extracting the heart rate features. X HR and pulse characteristics X P The acquired image data is processed to identify the training object within the image data, and then cropped to ensure that the training object occupies a stable proportion in consecutive images, thereby obtaining image data features. X IMG For feature extraction processing of EEG data, the specific steps can be as follows: first, perform preprocessing steps such as 1-50Hz bandpass filtering, baseline correction, and artifact removal on the EEG data; then, extract the EEG differential entropy features. X EEG Differential entropy is a frequency domain feature designed for electroencephalogram (EEG) signals. It is an extension of Shannon entropy, defined as:
[0049]
[0050] in, f ( x )yes X The probability density function. Assume that the EEG signal follows a Gaussian distribution. N (μ,σ 2 time series X (Where μ is the mean of the EEG signal distribution and σ is the standard deviation), then its corresponding differential entropy is expressed as:
[0051]
[0052] As can be seen from this formula, the differential entropy of a certain frequency band is equivalent to the logarithm of its energy in that frequency band. Compared with frequency band energy characteristics, differential entropy characteristics can balance the significant differences in energy across different frequency domains of EEG, reduce the errors in accuracy caused by energy value differences in subsequent calculations, and improve the discrimination ability of learning algorithms.
[0053] After obtaining the input data, it is then fed into the multimodal fusion network in the mental state assessment model for processing to obtain multimodal fusion features that integrate non-contact acquisition data and EEG data.
[0054] In step S13, after obtaining the multimodal fusion features, it is only necessary to input the multimodal fusion features into the mental state prediction model. The trained mental state prediction model can then analyze the multimodal fusion features to obtain the user's current mental state analysis results.
[0055] The mental state analysis method provided by this invention combines contactless data collected without contact with the user with the user's EEG data as input for training a mental state assessment model. The mental state assessment model fuses data from these two types of data in different modalities to obtain multimodal fusion features, enabling the mental state assessment model to effectively incorporate EEG data to analyze the user's mental state and thus improve the accuracy of mental state analysis.
[0056] The mental state assessment model used in step S12 includes at least a multimodal fusion network and a mental state prediction model. Figure 2 The steps of training a mental state assessment model in a mental state analysis method according to an embodiment of the present invention are illustrated schematically. (Refer to...) Figure 2 As shown, this training method may specifically include the following steps:
[0057] Step S21: Obtain mental training data, which includes the physiological data of the training subject and known actual mental state labels. The physiological data of the training subject includes non-contact data collection and electroencephalogram (EEG) data.
[0058] Step S22: After extracting features from the physiological data of the training subjects, the first input data including non-contact acquisition features and EEG features is obtained. The first input data is input into the multimodal fusion network to output multimodal fusion features.
[0059] Step S23: Input the multimodal fusion features output by the multimodal fusion network into the mental state prediction model, and output the predicted mental state label;
[0060] Step S24: Train the multimodal fusion network by backpropagating the error between the second input data obtained after input reconstruction of the multimodal fusion features and the first input data.
[0061] The mental state prediction model is optimized based on the cross-entropy loss between the predicted mental state label and the actual mental state label.
[0062] In step S21, the acquired mental training data includes the physiological data of the training subjects and known actual mental state labels. The training subjects are the users whose physiological data is collected. The physiological data includes contactless data collection and EEG data. Correspondingly, the known actual mental state labels should be label information that can represent the actual mental state of the corresponding training subject at the time the physiological data is collected. This ensures a one-to-one correspondence between the physiological data of the training subjects and the known actual mental state labels in the mental training data used to train the mental state assessment model, guaranteeing the accuracy of the mental state assessment model training. Contactless data collection refers to physiological data that can be collected without requiring the training subjects to wear devices for collecting the corresponding physiological data. This can include eye movement data, heart rate data, pulse data, and image data. The specific methods for acquiring eye movement data, heart rate data, pulse data, and image data can be referred to the relevant description in step S13 above, and will not be elaborated further here.
[0063] In step S22, after obtaining the mental training data, feature extraction processing needs to be performed on the mental training data to obtain the first input data used to input the mental state assessment model. The specific feature extraction process can be referred to the relevant explanation in step S12, and will not be elaborated here. Taking the contactless data collection mentioned earlier, which includes eye movement data, heart rate data, pulse data, and image data, as an example, the feature data obtained after the specific feature extraction processing is as follows: X EYE Heart rate characteristics X HR Pulse characteristics X P Image data features X IMG and EEG differential entropy characteristics X EEGThe first input data is used as the input data. After obtaining the first input data, it is then fed into the multimodal fusion network in the mental state assessment model for processing to obtain multimodal fusion features. Specifically, the multimodal fusion network is a deep autoencoder, which can automatically assign encoders and process high-dimensional data into low-dimensional data, and the decoder can also restore low-dimensional data into high-dimensional data. It includes an encoder, a single-layer fully connected fusion network, and a decoder corresponding to the encoder. The encoder is used to encode each feature data in the first input data to obtain the first input data of the encoded unified digital signal. Corresponding to the first input data, the encoder includes... E EEG , E EYE , E HR , E P , E IMG A single-layer fully connected fusion network is used to concatenate the encoded first input data to form a multimodal fusion feature as the output. This multimodal fusion feature contains all the feature information from the first input data. The decoder, corresponding to the encoder, is used to decode the formed multimodal fusion feature to reconstruct the corresponding feature information to form the second input data. The decoder includes... D EEG , D EYE , D HR , D P , D IMG It is understood that the second input data corresponds to the first input data and has the same feature data, so that the multimodal fusion network can be trained and optimized in step S24 based on the second input data and the first input data.
[0064] In step S23, the multimodal fusion features output by the multimodal fusion network are input into the mental state prediction model of the mental state assessment model. This allows the mental state prediction model to analyze the multimodal fusion features and obtain predicted mental state labels corresponding to the physiological data of the training subjects in the mental training data. The specific model of the mental state prediction model can be found in relevant descriptions in the prior art, and will not be elaborated upon in this embodiment.
[0065] Step S24 involves training and optimizing the multimodal fusion network and the mental state prediction model in the mental state assessment model to improve the accuracy of the mental state assessment model in mental state analysis. Specifically, when training the multimodal fusion network, the backpropagation of the error between the second input data obtained after input reconstruction using multimodal fusion features and the first input data is used for training. When training the mental state prediction model, it is optimized based on the cross-entropy loss between the predicted mental state label and the actual mental state label.
[0066] By optimizing the multimodal fusion network and the mental state prediction model, a mental state assessment model can be obtained that uses the physiological data of training subjects containing EEG data as input, thereby effectively improving the accuracy of mental state analysis using this mental state assessment model.
[0067] Considering that monitoring EEG data often requires wearing head-mounted devices under current technology, which may be unsuitable in some scenarios due to cost, preference, or other reasons, this invention also provides another method for analyzing mental state. This method can analyze a user's mental state using the mental state prediction model of the aforementioned trained mental state assessment model without requiring the wearing of a head-mounted device or the collection of the user's EEG data. This mental state analysis method also has higher accuracy compared to a mental state assessment model trained without using EEG data at all. Figure 3 The flowchart of another embodiment of the mental state analysis method of the present invention is illustrated schematically. This method is applied to a smart mirror, with reference to... Figure 3 As shown, the method includes the following steps:
[0068] Step S31: When the user faces the smart mirror, the detector carried by the smart mirror is used to acquire the user's contactless data.
[0069] Step S32: After extracting features from the user's contactless data collection data, contactless data collection features are obtained;
[0070] Step S33: After the output of the contactless acquisition features processed by the encoder is spliced together, random noise is added and input into the generator of the conditional adversarial generative network in the mental state assessment model to generate features that approximate multimodal fusion features. The multimodal fusion features are features obtained by inputting input data including contactless acquisition features and EEG features into the multimodal fusion network.
[0071] Step S34: Input the generated features into the mental state prediction model in the mental state assessment model, and output the predicted mental state label.
[0072] Step S31 is similar to step S11, requiring the acquisition of the user's physiological data. However, since this method does not require the use of the user's EEG data, only contactless data collection is needed. Taking the mental state assessment model trained earlier as an example, the contactless physiological data to be collected can include eye movement data, heart rate data, pulse data, and image data. This data can be acquired using detectors carried by the smart mirror. Specifically, these detectors can include an eye tracker for acquiring eye movement data, a millimeter-wave radar for acquiring heart rate and pulse data, and a miniature camera for acquiring image data. The specific acquisition methods can be found in the description of step S11 and will not be repeated here. Similarly, step S32 is similar to the feature extraction process in step S12, obtaining contactless acquisition features corresponding to the contactless acquisition data. The specific extraction methods can be found in the description of step S12 and will not be repeated here.
[0073] In step S33, the obtained non-contact physiological features are encoded by the encoder of the multimodal fusion network, and the output is then concatenated. Random noise is then added as input to the generator of the conditional adversarial generative network. This generator, based on the output of the non-contact acquisition features processed by the encoder and then concatenated with random noise, outputs features that closely resemble the generated features of the multimodal fusion network. It can be understood that this generator, after training, produces features based on the non-contact acquisition features processed by the encoder and then concatenated with random noise. When these generated features are input into the mental state prediction model within the mental state assessment model, the resulting predicted mental state labels show a high degree of similarity to those obtained using multimodal fusion features containing EEG characteristics.
[0074] In step S34, since the corresponding generated features have been obtained through the conditional adversarial generative network, it is only necessary to replace the multimodal fusion features with the generated features and input them into the mental state prediction model to obtain a highly accurate mental state analysis result without adding EEG data.
[0075] In this implementation, since EEG data is not required as input to the mental state assessment model during mental state analysis, product manufacturers can design two different products (one with a data acquisition device for collecting user EEG data and one without) for two different mental state analysis methods. Users can also choose different mental state analysis methods based on their usage habits and their willingness to wear a data acquisition device. For mental state analysis methods that incorporate EEG data as input to the mental state assessment model, the accuracy of the analysis results will be... The accuracy of mental state analysis results is higher than that of the former, while for mental state analysis methods that do not require EEG data as input to the mental state assessment model, the accuracy of the mental state analysis results is somewhat lacking compared to the former. However, due to the current limitations of hardware technology, there is no comfortable wearable EEG acquisition device suitable for long-term wear. When using the corresponding products, users can collect relevant data in a completely contactless manner without wearing a device to collect the user's EEG data. This allows for contactless analysis of the user's mental state, and the accuracy of the mental state analysis results is higher compared to general mental state analysis methods that do not include EEG data.
[0076] The conditional adversarial generative network used in step S33 includes a generator and a discriminator. Figure 4 The steps of the training method for the conditional adversarial generative network in a mental state analysis method according to an embodiment of the present invention are illustrated schematically, with reference to... Figure 4 As shown, this training method may specifically include the following steps:
[0077] Step S41: The output of the contactless acquisition features processed by the encoder is spliced together and random noise is added as input. The multimodal fusion features are used as the target output to train the generator.
[0078] Step S42: Input the multimodal fusion features and the generated features actually output by the generator into the discriminator to train the discriminator;
[0079] Step S43: Perform adversarial training using the generator and the discriminator.
[0080] Step S41 is the training step for the generator in the conditional adversarial generative network. In step S41, the contactless acquisition features after feature extraction are first encoded and concatenated by the encoder of the multimodal fusion network. The specific feature extraction steps can be referred to the relevant descriptions above, and will not be repeated here. The processed data contains all the feature information of the contactless acquisition data. Compared with the multimodal fusion features used as input to the mental state prediction model, this data lacks EEG data features. Therefore, random noise is added as input to this data, and the multimodal fusion features containing EEG data features are used as the target output to train the generator of the conditional adversarial generative network. This is to ensure that the generated features output by the generator of the conditional adversarial generative network, which is based on the output of the contactless acquisition features processed by the encoder and then concatenated with random noise as input, can be as close as possible to the multimodal fusion features containing EEG data features.
[0081] Step S42 is the training step for the discriminator in the conditional adversarial generative network. In step S42, the discriminator is trained by using the multimodal fusion feature containing EEG data features and the generated feature obtained by splicing the output of the generator based on non-contact acquisition features processed by the encoder in step S41 and adding random noise as input. This trains the discriminator to improve its judgment accuracy, so that the discriminator can more accurately judge the difference between the generated feature output by the generator of the conditional adversarial generative network and the multimodal fusion feature containing EEG data features.
[0082] Step S43 involves adversarial training using the generator and discriminator to further optimize the conditional adversarial generative network, thereby improving the realism of the generated features and the accuracy of the discriminator. Specifically, this adversarial training involves iteratively training the generator and discriminator separately, repeating steps S41 and S42. First, the generator is trained by concatenating the output of the non-contact acquisition features processed by the encoder and adding random noise as input. This enables the generator to generate features that closely resemble multimodal fusion features containing EEG data, making the generated features as close as possible to the multimodal fusion features to "deceive" the discriminator. Then, the generated features output by the generator trained in step S41 and the multimodal fusion features are used as input to train the discriminator, improving its accuracy and enabling it to more accurately distinguish between the generated features output by the generator and the multimodal fusion features containing EEG data. By repeating the above steps multiple times, the generated features obtained by stitching together the output of the encoder based on contactless feature acquisition and adding random noise as input can approach the multimodal fusion features containing EEG data features as closely as possible. The optimization function of the generator obtained after final training is as follows:
[0083]
[0084] The final optimized function of the discriminator obtained through training is:
[0085]
[0086] in, F concat It is a spliced feature obtained by processing non-contact acquisition features with an encoder and then splicing them together.
[0087] G θ It is a generator with parameter θ;
[0088] D The parameter is The discriminator;
[0089] r G ~ G θ (.| F concat )middle, r G These are the generated features output by the generator. G θ (.| F concat ) is the inputF concat The distribution of generated features output by the generator. r G ~ G θ (.| F concat ) is consistent with the input. F concat The distribution of generated features output by the generator r G ;
[0090] D ( r G , F concat ) is a condition F concat At that time, the discriminator output r G The probability value of the true multimodal fusion feature;
[0091] ( r , F concat )~ P d ( r )middle, r It is a multimodal fusion feature. P d ( r ) is the distribution of multimodal fusion features, ( r , F concat )~ P d ( r () is a multimodal fusion feature that conforms to the distribution of multimodal fusion characteristics. r splicing features F concat ;
[0092] D ( r , F concat ) is a condition F concat At that time, the discriminator output r This represents the probability value of the true multimodal fusion feature.
[0093] For the generator, the closer the result is to 1, the closer the generated features are to the multimodal fusion features; therefore, a result closer to 1 is better. For the discriminator, the closer the result is to 0, the less similar the generated features it judges are to the multimodal fusion features; therefore, a result closer to 0 is better. By replacing the multimodal fusion features with this generated feature in the mental state prediction model of the mental state assessment model, the accuracy of mental state analysis results can be significantly improved without collecting the user's EEG data.
[0094] Figure 5 schematically shown Figure 2 The training method for the overall mental state assessment model after incorporating a conditional adversarial generative network is described in reference to... Figure 5 The training method first acquires the EEG signals, eye movement events, echo signals, and image data of the training subjects, and then obtains the first input data through feature extraction. X EEG (EEG data) X EYE (Eye-tracking data) X HR (Heart rate data) X P (Pulse data) and X IMG (Image data), and then the first input data is input into the encoder through a multimodal fusion network. E EEG (EEG data encoder) E EYE (Eye-tracking data encoder) E HR (Heart rate data encoder) E P (Pulse data encoder) and E IMG The image data is encoded by an image encoder, and the encoded first input data is concatenated through a single-layer fully connected fusion network to obtain multimodal fusion features. These multimodal fusion features are then deconcatenated through the single-layer fully connected fusion network and passed through a decoder. D EEG (Brainwave Data Decoder) D EYE (Eye-tracking data decoder) D HR (Heart rate data decoder) D P (Pulse data decoder) and D IMGThe image data decoder reconstructs the second input data. Using backpropagation of the error between the second and first input data, a multimodal fusion network is trained. The multimodal fusion features are then input into a mental state prediction model, which outputs predicted mental state labels. The model is optimized using cross-entropy loss between the predicted and actual mental state labels. After training the multimodal fusion features and the mental state prediction model, a generator that constructs a conditional adversarial generative network is used. G and discriminator D ,Will X EYE , X HR , X P and X IMG The pre-trained multimodal fusion features are encoded and concatenated, and random noise is added to the processed data to serve as a generator. G The input is used to train the generator G The output is a generated feature that approximates the multimodal fusion feature, and the generated feature and the multimodal fusion feature are used as a discriminator. D The input is used to train the discriminator. D The accuracy of the judgment is improved by utilizing the generator. G and discriminator D Adversarial training is conducted to complete the training of the conditional adversarial generative network.
[0095] Figure 6 schematically shown Figure 3 The workflow for a mental state analysis method that does not require collecting EEG data is as follows: Figure 6 This method first collects user data without contact using contactless data acquisition devices. For example, it obtains eye-tracking data by using an eye tracker embedded in the product to collect eye-tracking events; it obtains heart rate and pulse data by monitoring echo signals using a millimeter-wave radar detector within the product; and it collects image data from a miniature camera within the product. Then, it extracts features from the contactless data to obtain… X EYE , X HR , X P and X IMG encoder through multimodal fusion network E EYE , E HR , E P and E IMGThe encoded data is then concatenated, random noise is added, and the data is input into the generator of a conditional generative adversarial network to obtain generated features. Finally, the generated features are input into a mental state prediction model to obtain predicted mental state labels, thus obtaining the mental state analysis results.
[0096] Figure 7 A smart mirror according to an embodiment of the present invention is illustrated schematically, with reference to... Figure 7 As shown, the smart mirror includes:
[0097] The smart mirror body 11, detector 12, and processor 13;
[0098] The detector 12 is installed on the smart mirror body 11 and is used to acquire contactless data from users facing the smart mirror body 11.
[0099] The processor 13 is disposed inside the smart mirror body 11, connected to the detector 12 to obtain contactless data collected by the detector 12, and is used to execute the steps of the mental state analysis method described in the above embodiments.
[0100] Among them, reference Figure 8 and Figure 9 As shown, in some embodiments, the smart mirror body 11 includes a mirror 111 and a display 112, and the detector 12 includes an eye tracker 114, a millimeter-wave radar 113 and a miniature camera 115. The contactless data acquisition includes eye movement data, heart rate data, pulse data and image data.
[0101] The eye tracker 114, millimeter-wave radar 113, and miniature camera 115 are all embedded in the lower part of the smart mirror body 11.
[0102] It should be noted that the implementation process and principle of the smart mirror in this embodiment of the invention can be found in the corresponding descriptions of the above-mentioned mental state analysis method embodiments, such as the descriptions of feature extraction of physiological data of training subjects, processing of input data, and methods for obtaining predicted mental state labels in the method embodiment section. Therefore, they will not be repeated here.
[0103] Figure 10 A mental state analysis device according to an embodiment of the present invention is illustrated schematically, with reference to... Figure 10 As shown, the device includes:
[0104] The physiological data acquisition module 21 is used to acquire the user's physiological data by using the detector carried on the smart mirror body and the head-mounted device when the user wears the head-mounted device and faces the smart mirror body. The physiological data includes contactless acquisition data and electroencephalogram (EEG) data.
[0105] The feature fusion module 22 is used to extract features from the user's physiological data to obtain input data including non-contact acquisition features and EEG features, input the input data into the multimodal fusion network in the mental state assessment model, and output multimodal fusion features.
[0106] The mental state prediction module 23 is used to input the multimodal fusion features output by the multimodal fusion network into the mental state prediction model in the mental state assessment model, and output the predicted mental state label.
[0107] It should be noted that the implementation process and principle of the mental state analysis device in this embodiment of the invention can be specifically referred to in the corresponding descriptions of the above method embodiments, such as the descriptions of physiological data acquisition, multimodal fusion feature acquisition, and output of predicted mental state labels in the method embodiments, and therefore will not be repeated here. Exemplarily, the mental state analysis device in this embodiment of the invention can be any intelligent device with a processor, including but not limited to smart mirrors, mental health guidance devices, intelligent interactive robots, computers, smartphones, personal computers, cloud servers, etc.
[0108] Figure 11 The diagram schematically illustrates a mental state analysis device according to another embodiment of the present invention. This mental state analysis device does not require the collection of the user's electroencephalogram (EEG) data as input to the model when performing mental state analysis. (Refer to...) Figure 11 As shown, the device includes:
[0109] The contactless physiological data acquisition module 31 is used to acquire contactless data of the user by using the detector carried by the smart mirror when the user is facing the smart mirror.
[0110] Feature extraction module 32 is used to extract features from the user's contactless collection data to obtain contactless collection features;
[0111] The feature acquisition module 33 is used to splice the output of the non-contact acquisition features after processing by the encoder, add random noise, and input it into the generator of the conditional adversarial generative network in the mental state assessment model to generate features that approximate multimodal fusion features. The multimodal fusion features are features obtained by inputting input data including non-contact acquisition features and EEG features into the multimodal fusion network.
[0112] The mental state prediction module 34 inputs the generated features into the mental state prediction model in the mental state assessment model and outputs the predicted mental state label.
[0113] It should be noted that the implementation process and principle of the mental state analysis device in this embodiment of the invention can be specifically referred to in the corresponding descriptions of the above method embodiments, such as the descriptions of contactless data acquisition, feature extraction methods, feature generation, and output of predicted mental state labels in the method embodiments, and therefore will not be repeated here. Exemplarily, the mental state analysis device in this embodiment of the invention can be any intelligent device with a processor, including but not limited to smart mirrors, mental health guidance devices, intelligent interactive robots, computers, smartphones, personal computers, cloud servers, etc.
[0114] In some embodiments, the present invention provides a non-volatile computer-readable storage medium storing one or more programs including execution instructions, which can be read and executed by electronic devices (including but not limited to computers, servers, or network devices, etc.) to perform the sensor-based strong interaction method based on artificial intelligence mental state analysis according to any of the above embodiments of the present invention.
[0115] In some embodiments, the present invention also provides a computer program product, the computer program product including a computer program stored on a non-volatile computer-readable storage medium, the computer program including program instructions, which, when executed by a computer, cause the computer to perform the sensor-based strong interaction method based on artificial intelligence mental state analysis of any of the above embodiments.
[0116] In some embodiments, the present invention also provides an electronic device comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the sensor-based strong interaction method based on artificial intelligence mental state analysis of any of the above embodiments.
[0117] In some embodiments, the present invention also provides a storage medium storing a computer program, characterized in that, when the program is executed by a processor, it implements the sensor-based strong interaction method based on artificial intelligence mental state analysis of any of the above embodiments.
[0118] Figure 12 This is a schematic diagram of the hardware structure of an electronic device for performing a mental state analysis method according to another embodiment of this application, as shown below. Figure 12 As shown, the device includes:
[0119] One or more processors 410 and memory 420, Figure 12 Take a processor 410 as an example.
[0120] The device for performing mental state analysis methods may also include an input device 430 and an output device 440.
[0121] The processor 410, memory 420, input device 430, and output device 440 can be connected via a bus or other means. Figure 12 Taking the example of a connection between China and Israel via a bus.
[0122] The memory 420, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the mental state analysis method in the embodiments of this application. The processor 410 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 420, thereby implementing the mental state analysis method of the above-described method embodiments.
[0123] The memory 420 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the mental state analysis method, etc. Furthermore, the memory 420 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 420 may optionally include memory remotely located relative to the processor 410, and these remote memories can be connected to the electronic device via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0124] Input device 430 can receive input digital or character information and generate signals related to user settings and function control of the image processing device. Output device 440 may include display devices such as a display screen.
[0125] The one or more modules are stored in the memory 420, and when executed by the one or more processors 410, they perform the mental state analysis method in any of the above method embodiments.
[0126] The above-described product can perform the methods provided in the embodiments of this application, and has the corresponding functional modules and beneficial effects for performing the methods. Technical details not described in detail in this embodiment can be found in the methods provided in the embodiments of this application.
[0127] The electronic devices in this application embodiments exist in various forms, including but not limited to:
[0128] (1) Smart mirror: This type of device is an all-in-one machine that combines a mirror and multiple sensors. While realizing a variety of convenient functions, it can perform multi-faceted real-time AI detection, analysis and feedback on users. Combined with mental state analysis, it can provide users with services such as mental health diagnosis and improve the user experience.
[0129] (2) Mental health triage equipment: This type of equipment has display and data processing functions. By connecting with the data acquisition equipment, it can collect the patient's physiological parameters, analyze the patient's current mental state, and assist doctors in determining the patient's condition.
[0130] (3) Intelligent interactive robots: These devices can generally be set up in public areas such as shopping malls and office buildings. They can interact with users and provide more personalized consulting services by combining mental state analysis.
[0131] (4) Mobile communication devices: These devices are characterized by their mobile communication capabilities and primarily aim to provide voice and data communication. These terminals include: smartphones (e.g., iPhones), multimedia phones, feature phones, and low-end phones, etc.
[0132] (5) Ultra-mobile personal computer devices: These devices fall under the category of personal computers, possessing computing and processing capabilities, and generally also have mobile internet access features. These terminals include PDAs, MIDs, and UMPCs, such as the iPad.
[0133] (6) Server: A device that provides computing services. The components of a server include a processor, hard disk, memory, system bus, etc. Servers are similar to general computer architectures, but because they need to provide highly reliable services, they have higher requirements in terms of processing power, stability, reliability, security, scalability, and manageability.
[0134] (7) Other electronic devices with data processing and interactive functions.
[0135] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0136] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0137] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A smart mirror, characterized in that, The application relates to a smart mirror body, a detector and a processor. The detector is arranged on the smart mirror body and is used for acquiring non-contact acquisition data of a user facing the smart mirror body. The processor is arranged in the smart mirror body and is connected with the detector to acquire the non-contact acquisition data acquired by the detector and is used for executing a mental state analysis method. The detector comprises an eye tracker, a millimeter wave radar and a miniature camera, and the non-contact acquisition data comprises eye movement data, heart rate data, pulse data and image data. The eye tracker, the millimeter wave radar and the miniature camera are all embedded in the lower part of the smart mirror body. The mental state analysis method comprises the following steps: When the user faces the smart mirror, the non-contact acquisition data of the user is acquired by using the detector carried by the smart mirror; After the non-contact acquisition data of the user is extracted, non-contact acquisition features are obtained; The outputs of the non-contact acquisition features after being processed by encoders are spliced and added with random noise to generate a generator of a conditional adversarial generative network in a mental state evaluation model to generate generation features close to multi-modal fusion features, wherein the multi-modal fusion features are features obtained by inputting input data comprising non-contact acquisition features and electroencephalogram features into a multi-modal fusion network; The generation features are input into a mental state prediction model in the mental state evaluation model, and a predicted mental state label is output. The electroencephalogram features are differential entropy features of electroencephalograms.
2. The smart mirror of claim 1, wherein, The mental state evaluation model at least comprises a multi-modal fusion network and a mental state prediction model, and the mental state evaluation model is trained by the following steps:
3. The smart mirror of claim 1, wherein, Acquire mental training data, wherein the mental training data comprises physiological data of a training object and known actual mental state labels, and the physiological data of the training object comprises non-contact acquisition data and electroencephalogram data; After the physiological data of the training object is extracted, first input data comprising non-contact acquisition features and electroencephalogram features are obtained, and the first input data is input into a multi-modal fusion network to output multi-modal fusion features; The multi-modal fusion features output by the multi-modal fusion network are input into the mental state prediction model to output a predicted mental state label; Wherein, the error between the second input data obtained by inputting the multi-modal fusion features and the first input data is back propagated to train the multi-modal fusion network, The mental state prediction model is optimized based on the cross-entropy loss between the predicted mental state label and the actual mental state label. The multi-modal fusion network comprises an encoder, a single-layer full connection fusion network and a decoder corresponding to the encoder, 4. The intelligent mirror of claim 3, wherein, The first input data is input into the single-layer full connection fusion network after being processed by the encoder to output the multi-modal fusion features; The second input data is obtained by inputting the multi-modal fusion features into the decoder for input reconstruction. The mental state evaluation model further comprises a conditional adversarial generative network, the conditional adversarial generative network comprises a generator and a discriminator, and the training method comprises the following steps:
5. The intelligent mirror of claim 4, wherein, The output of the contactless collection feature processed by the encoder is spliced and added with random noise as input, and the multi-modal fusion feature is taken as target output to train the generator; The multi-modal fusion feature and the generated feature actually output by the generator are input into the discriminator to train the discriminator; The generator and the discriminator are used for adversarial training. 6.The smart mirror of claim 5, wherein An optimization function of the generator is An optimization function of the discriminator is wherein, F concat is the spliced feature obtained by splicing the output of the contactless acquisition feature after the encoder processing. G θ is a generator with parameter θ; is a discriminator with parameters ; r G ~ G θ (.| F concat ) in which, r G is a generated feature output by the generator, G θ (.| F concat ) is a distribution of generated features output by the generator input with F concat , r G ~ G θ (.| F concat ) is a generated feature that conforms to the distribution of generated features output by the generator input with F concat , r G ; is conditional F concat the discriminator output r G is a probability value of the real multimodal fusion feature; ( r , F concat )~ P d ( r )middle, r It is a multimodal fusion feature. P d ( r ) is the distribution of multimodal fusion features, ( r , F concat )~ P d ( r () is a multimodal fusion feature that conforms to the distribution of multimodal fusion characteristics. r splicing features F concat ; is conditional F concat the discriminator output r is the probability value of the real multimodal fusion feature.
Citation Information
Patent Citations
Mental state analysis method, apparatus, equipment, computer medium and multifunctional chair
CN109124655A