Emotion intervention system, method and device based on multi-modal analysis and medium

Through a multimodal analysis emotional intervention system, combined with emotion analysis and behavior prediction, audio and lighting intervention is implemented, which solves the problem that smart home cannot regulate user emotions and improves user experience.

CN120340539APending Publication Date: 2025-07-18HANSONG NANJING TECH LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510735650.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing smart home system cannot intelligently regulate the user's emotions, resulting in poor user experience.

Method used

Using a multimodal analysis emotion intervention system, through the emotion analysis module, behavior prediction module and intervention module, based on user voice data and historical behavior data, emotional characteristics and predict behavior are determined, and audio intervention and/or lighting intervention are implemented to regulate user emotions.

Benefits of technology

It realizes early warning and positive emotional guidance for user emotions, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340539A_ABST
    Figure CN120340539A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an emotion intervention system, method and device based on multi-modal analysis and a medium. The system comprises an emotion analysis module, a behavior prediction module and an intervention module. The emotion analysis module is configured to determine emotion features based on the user voice data; the behavior prediction module is configured to determine a predicted behavior based on the historical behavior data; the intervention module is configured to determine whether to intervene and an intervention strategy based on the emotion features and the predicted behaviors; in response, determining a first audio parameter corresponding to the audio intervention, and driving the audio equipment to play sound based on the first audio parameter; and / or determining a first illumination parameter corresponding to the illumination intervention and driving the illumination device to illuminate based on the first illumination parameter.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of emotion intervention technologies, and particularly to an emotion intervention system, method, device and medium based on multimodal analysis. Background Art

[0002] Existing smart homes mainly combine software and environmental sounds to achieve intelligent control. However, this control method cannot perform intelligent control according to the user's emotions, resulting in a poor user experience of smart homes. In the context of the integrated development of smart homes and affective computing technologies, how to intervene in the user's emotions through multimodal means is an urgent problem to be solved.

[0003] Therefore, it is desired to provide an emotion intervention system, method, device and medium based on multimodal analysis, which can achieve early warning of the user's emotions through multimodal means, provide positive emotion guidance to the user through smart homes, and improve the user's experience. Summary of the Invention

[0004] One or more embodiments of this specification provide an emotion intervention system based on multimodal analysis. The system includes: an emotion analysis module, a behavior prediction module, and an intervention module; the emotion analysis module is configured to determine emotion features based on user voice data; the behavior prediction module is configured to determine a predicted behavior based on historical behavior data; the intervention module is configured to: based on the emotion features and the predicted behavior, determine whether to intervene and an intervention strategy, the intervention strategy including audio intervention and / or lighting intervention; in response, determine a first audio parameter corresponding to the audio intervention, and drive an audio device to play a sound based on the first audio parameter; and / or, determine a first lighting parameter corresponding to the lighting intervention and drive the lighting device to perform lighting based on the first lighting parameter.

[0005] One embodiment of this specification provides an emotion intervention method based on multimodal analysis. The method includes: determining emotion features based on user voice data; determining a predicted behavior based on historical behavior data; based on the emotion features and the predicted behavior, determining whether to intervene and an intervention strategy, the intervention strategy including audio intervention and / or lighting intervention; in response, determining a first audio parameter corresponding to the audio intervention, and driving an audio device to play a sound based on the first audio parameter; and / or, determining a first lighting parameter corresponding to the lighting intervention and driving the lighting device to perform lighting based on the first lighting parameter.

[0006] One embodiment of this specification provides an emotion intervention device based on multimodal analysis. The device includes at least one processor and at least one memory. The at least one memory is used to store computer instructions. The at least one processor is used to execute at least some of the computer instructions to implement the above-mentioned emotion intervention method based on multimodal analysis.

[0007] One embodiment of this specification provides a computer-readable storage medium. The storage medium stores computer instructions. When a computer reads the computer instructions in the storage medium, the computer executes the above-mentioned emotion intervention method based on multimodal analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] This specification will be further described by way of exemplary embodiments, which will be described in detail through the drawings. These embodiments are not restrictive. In these embodiments, the same numbers represent the same structures, where: Figure 1 is an exemplary module diagram of an emotion intervention system shown in some embodiments of this specification; Figure 2 is an exemplary flowchart of an emotion intervention method based on multimodal analysis shown in some embodiments of this specification; Figure 3 is an exemplary schematic diagram of an emotion analysis model shown in some embodiments of this specification; Figure 4 is an exemplary flowchart of determining an intervention strategy shown in some embodiments of this specification. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0009] To more clearly illustrate the technical solutions of the embodiments of this specification, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some examples or embodiments of this specification. For those of ordinary skill in the art, without creative efforts, this specification can also be applied to other similar scenarios based on these drawings. Unless obvious from the language context or otherwise stated, the same reference numerals in the drawings represent the same structure or operation.

[0010] It should be understood that the "system", "device", "unit" and / or "module" used herein is a method for distinguishing different components, elements, parts, portions or assemblies at different levels. However, if other words can achieve the same purpose, the said words can be replaced by other expressions.

[0011] Flowcharts are used in this specification to illustrate the operations performed by the system according to the embodiments of this specification. It should be understood that the preceding or subsequent operations do not necessarily have to be executed precisely in sequence. On the contrary, the steps can be processed in reverse order or simultaneously. Also, other operations can be added to these processes, or one or several steps can be removed from these processes.

[0012] Figure 1 is an exemplary module diagram of an emotion intervention system shown according to some embodiments of this specification.

[0013] Some embodiments of this specification provide an emotion intervention system based on multimodal analysis (hereinafter simply referred to as the emotion intervention system). As Figure 1 shown, the emotion intervention system 100 may include an emotion analysis module 110, a behavior prediction module 120, an intervention module 130, etc.

[0014] In some embodiments, the emotion analysis module 110 is configured to determine emotion features based on user voice data.

[0015] In some embodiments, the emotion analysis module 110 is further configured to determine emotion features based on user voice data through an emotion analysis model.

[0016] In some embodiments, the behavior prediction module 120 is configured to determine a predicted behavior based on historical behavior data.

[0017] In some embodiments, the intervention module 130 is configured to determine whether to intervene and an intervention strategy based on emotion features and predicted behavior; in response thereto, determine a first audio parameter corresponding to audio intervention, and drive an audio device to play sound based on the first audio parameter; and / or, determine a first lighting parameter corresponding to lighting intervention and drive a lighting device to perform lighting based on the first lighting parameter.

[0018] In some embodiments, the intervention module 130 is further configured to determine whether the user's emotion is a positive emotion based on emotion features and predicted behavior.

[0019] In some embodiments, the intervention module 130 is further configured to determine an intervention strategy based on emotion features and second environmental data through an anchor database.

[0020] In some embodiments, the anchor module 140 is configured to generate an anchor combination in response to the user's positive emotion, where the anchor combination includes at least two of a second audio parameter, a second lighting parameter, and a fragrance parameter, and perform at least two of the following operations: drive an audio device to play a sound based on the second audio parameter, drive a lighting device to perform lighting based on the second lighting parameter, drive a fragrance device to emit a fragrance based on the fragrance parameter; store the positive emotion, the first environmental data, and the anchor combination in the anchor database.

[0021] In some embodiments, the customization module 150 is configured to receive the emotion intervention conditions and / or intervention preferences input by the user, and update the intervention strategy based on the emotion intervention conditions and / or intervention preferences.

[0022] In some embodiments, the emotion intervention system 100 may further include a processor, a storage device, a user terminal, etc.

[0023] In some embodiments, the processor may process information and / or data related to the emotion intervention system 100 to perform one or more functions described in this application. In some embodiments, the processor may include one or more processing engines (for example, a single-chip processing engine or a multi-chip processing engine). By way of example only, the processor may include a central processing unit (CPU), an application specific integrated circuit (ASIC), an application specific instruction processor (ASIP), a microprocessor, etc. or any combination of the above. In some embodiments, the processor may obtain the pre-stored data and / or information related to the emotion intervention system 100 from the storage device.

[0024] In some embodiments, one or more of the emotion analysis module 110, the behavior prediction module 120, the intervention module 130, the anchor module 140, and the customization module 150 may be integrated into the processor.

[0025] The user terminal refers to one or more terminal devices or software used by the user. The user may include the manager, operator, user, etc. of the emotion intervention system 100. For example, the user may input emotion intervention conditions and intervention preferences through the user terminal.

[0026] For more descriptions of the above content, see Figures 2 - 4 and its related descriptions.

[0027] It should be noted that the above descriptions of the emotion intervention system 100 and its modules are only for convenience of description, and do not limit this specification within the scope of the examples given. It can be understood that for those skilled in the art, after understanding the principle of the system, they may, without departing from this principle, make any combination of the various modules, or form a subsystem and connect it with other modules. In some embodiments, Figure 1The sentiment analysis module 110, behavior prediction module 120, intervention module 130, anchor module 140, and customization module 150 disclosed in

[0028] Figure 2 is an exemplary flowchart of a multi-modal analysis-based emotion intervention method according to some embodiments of the present specification. As Figure 2 shown, process 200 includes the following steps. In some embodiments, process 200 can be executed by a processor.

[0029] Step 210, determining an emotional feature based on user voice data.

[0030] User voice data refers to the collected voice data related to the user. In some embodiments, the processor can collect the user's voice data in real time based on hardware devices such as microphones.

[0031] An emotional feature refers to feature data characterizing the user's emotions. For example, an emotional feature can characterize the user's emotions such as happiness, anger, sadness, calmness, etc. In some embodiments, the emotional feature can be represented by a category. For example, positive emotions, negative emotions, etc. In some embodiments, the emotional feature can be represented by the percentage of negative emotions. For example, 0% negative emotion indicates the happiest mood, 50% negative emotion indicates neither happy nor negative (i.e., calm), 70% negative emotion indicates sadness, and 99% negative emotion indicates the most negative (i.e., anger).

[0032] In some embodiments, the processor can analyze the user voice data and extract the user's acoustic features. The user's acoustic features can include at least one of the fundamental frequency range, speech rate, energy variance, acoustic scene, etc. of the user's voice. Among them, the energy variance refers to the variance of the signal energy of the user's voice; the acoustic scene refers to the scene features of the user's voice, such as rising tone, short stress, tremor, etc.

[0033] In some embodiments, the processor may determine the emotional characteristics based on the user voice data in various ways. For example, the processor may determine the emotional characteristics by querying a first preset table based on the acoustic characteristics of the user obtained by analyzing the user voice data. Among them, the first preset table may be determined based on artificial preset. The first preset table may include the correspondence between the emotional characteristics and different acoustic characteristics of the user. For example, when the emotional characteristic is happy, the corresponding fundamental frequency range is 180 - 400 Hz, the speech rate is 5.2 - 6.8 syllables per second, the energy variance is greater than 0.35, and the acoustic scene is laughter and / or rising tone. Another example is that when the emotional characteristic is sad, the corresponding fundamental frequency range is 80 - 180 Hz, the speech rate is 3.0 - 4.2 syllables per second, the energy variance is less than 0.2, and the acoustic scene is long pause and / or tremor, etc.

[0034] In some embodiments, the processor may also determine the emotional characteristics based on the user voice data through an emotion analysis model. For more related content, see Figure 3 its description.

[0035] Step 220: Determine the predicted behavior based on the historical behavior data.

[0036] The historical behavior data refers to the relevant data of the user's historical behavior within a preset historical period from the current moment. For example, historical behavior logs, etc. The historical behavior log refers to the interaction record between the user and the device. For example, the operation records of the user on devices such as audio devices and lighting devices, or the device operation methods preset by the user. Among them, the device operation methods preset by the user may include the intervention strategies preset by the user. In some embodiments, the historical behavior data may be obtained based on the historical data recorded by the device. For more descriptions about audio devices and lighting devices, see the relevant descriptions in step 230 below.

[0037] The predicted behavior is the predicted user behavior and the degree of out-of-control of the user behavior within a period of time after the current moment. For example, the user frequently turns on and off the lighting device at the first moment, and the degree of out-of-control of this user behavior is 90%. Among them, the first moment (a period of time after the current moment) refers to a moment relatively close to the current moment in the future, such as 1 second, 2 seconds after the current moment, etc.

[0038] In some embodiments, the processor may determine the predicted behavior based on the historical behavior data in various ways. For example, the processor may determine the predicted behavior by a behavior prediction model based on the historical behavior data.

[0039] A behavior prediction model is a model used to determine predicted behaviors. In some embodiments, the behavior prediction model can be a machine learning model, for example, any one or combination of a Deep Neural Networks (DNN) model, a Convolutional Neural Networks (CNN) model, or other custom model structures, etc.

[0040] In some embodiments, the processor can train a behavior prediction model based on multiple first training samples with first labels. For example, the processor can input the first training samples into an initial behavior prediction model, construct a loss function based on the first labels and the output of the initial behavior prediction model, iteratively update the parameters of the initial behavior prediction model based on the loss function, and end the iteration when the iteration end condition is met to obtain a trained behavior prediction model. Among them, the methods of iterative update include but are not limited to the gradient descent method, and the iteration end condition can be that the loss function converges or the number of iterations reaches a threshold, etc.

[0041] The first training samples and the first labels can be determined based on historical data. The first training samples include actual historical behavior data obtained at a first time. Among them, the actual historical behavior data refers to the user behaviors that actually occurred in the historical data.

[0042] The first label is the actual historical behavior data and the corresponding historical out-of-control degree at a second time corresponding to the first training sample. Among them, the first time is before the second time. In some embodiments, the first label can be determined based on manual annotation. For example, the historical out-of-control degree corresponding to the actual historical behavior data can be determined based on a second preset table. The second preset table can be preset based on experience and includes the correspondence between the actual historical behavior data and the historical out-of-control degree.

[0043] Step 230, determine whether to intervene and the intervention strategy based on the emotional characteristics and the predicted behavior.

[0044] Intervention refers to intervening in the user's behavior. For example, controlling an audio device to play specific music, etc. When the user's mood is relatively negative, intervention is generally required to reduce the user's negative mood.

[0045] The intervention strategy refers to the method and strategy of intervention. In some embodiments, the intervention strategy can include audio intervention and / or lighting intervention, etc. The processor can determine intervention strategies such as audio intervention and / or lighting intervention to achieve multimodal analysis of the emotional characteristics. Among them, multimodal analysis refers to a way of achieving understanding or data analysis through multiple different forms or perceptual channels. It should be noted that the intervention system only executes the corresponding intervention strategy when intervention is needed.

[0046] Audio intervention refers to the intervention of a user's behavior by controlling an audio device. For example, playing audio at a specific frequency (sounds that are easily overlooked by the human ear, such as high-frequency audio of 15 - 20 kHz, low-frequency rhythm pulses of 20 - 40 Hz, etc.), cheerful songs, cross talks, jokes, songs with a specific timbre, etc. In some embodiments, by performing audio intervention on a user, it can help the user adjust their mood. For example, by playing audio at a specific frequency, it can help a user with an angry emotional characteristic calm down, changing their emotional characteristic to calm. Another example is that by playing a song with a specific timbre (such as white noise recorded in a natural environment like a forest), imitating the state of the user when in a natural environment, thus helping the user restore calmness.

[0047] An audio device refers to a device that plays audio. In some embodiments, the audio device can be a speaker, an MP3 player, etc.

[0048] Lighting intervention refers to the intervention of a user's behavior by controlling a lighting device. For example, controlling the light color, brightness, lighting pattern, lighting area, etc. of the lighting device. In some embodiments, by performing lighting intervention on a user, it can help the user adjust their mood. For example, by adjusting the light color to a warm color, it can help a user with a sad emotional characteristic feel a warm atmosphere, thus calming down their mood.

[0049] A lighting device refers to a device that provides lighting. In some embodiments, the lighting device can be at least one of a chandelier, a central control lighting device, etc.

[0050] In some embodiments, the processor can determine whether to intervene and the intervention strategy in various ways based on the emotional characteristic and the predicted behavior. For example, the processor can construct a feature vector based on the emotional characteristic and the predicted behavior, and based on the feature vector, by querying a vector database, take the whether-to-intervene and the reference intervention strategy corresponding to the reference vector with the highest similarity to the feature vector as the whether-to-intervene and the intervention strategy corresponding to the feature vector. Among them, the calculation of the similarity can be the calculation of the Euclidean distance, the cosine similarity, etc.

[0051] The vector database includes multiple reference vectors constructed based on the historical emotional characteristics and historical predicted behaviors in the historical data, as well as the whether-to-intervene and the reference intervention strategy corresponding to each reference vector. The processor can determine the actual intervention situation and the actual intervention strategy corresponding to the reference vector in the historical data as the whether-to-intervene and the reference intervention strategy. In some embodiments, the vector database can be modified and supplemented manually, etc.

[0052] Step 240, in response to this, determine the first audio parameter corresponding to the audio intervention, and drive the audio device to play sound based on the first audio parameter; and / or, determine the first lighting parameter corresponding to the lighting intervention and drive the lighting device to perform lighting based on the first lighting parameter.

[0053] The first audio parameter refers to the operating parameter for the audio device to perform an intervention. For example, it includes the audio frequency to be played, a specified song, etc.

[0054] The first lighting parameter refers to the operating parameter for the lighting device to perform an intervention. For example, it includes the color of the light to be illuminated, etc. In some embodiments, the processor can determine the first audio parameter corresponding to the audio intervention and the first lighting parameter corresponding to the lighting intervention through a PID (proportion integration differentiation) control algorithm.

[0055] In some embodiments, in response to the need for an intervention, the processor can perform the intervention in various ways based on the intervention strategy. For example, if the intervention strategy only includes an audio intervention, the processor can determine the first audio parameter corresponding to the audio intervention and drive the audio device to play sound based on the first audio parameter. Another example is that if the intervention strategy only includes a lighting intervention, the processor can determine the first lighting parameter corresponding to the lighting intervention and drive the lighting device to perform lighting based on the first lighting parameter. Another example is that if the intervention strategy includes both an audio intervention and a lighting intervention, the processor can perform the above two operations simultaneously.

[0056] In some embodiments, the intervention strategy further includes an aroma intervention. In response to the need for an intervention, the processor can determine the first aroma parameter corresponding to the aroma intervention and drive the aroma device to emit an aroma smell based on the first aroma parameter.

[0057] An aroma intervention refers to an action of performing an intervention by emitting an aroma smell through an aroma device.

[0058] The first aroma parameter refers to data related to the aroma. For example, the first aroma parameter includes the smell of the aroma and combinations of different smells, the concentration of the aroma, etc. In some embodiments, the processor can determine the first aroma parameter corresponding to the aroma intervention through a PID control algorithm.

[0059] An aroma device is a device that satisfies the aroma intervention by releasing an aroma smell. In some embodiments, the aroma device can be an aromatherapy machine, etc.

[0060] In some embodiments of this specification, by introducing the first aroma parameter to perform an aroma intervention on the environment where the user is located, it helps the user relieve anxiety, thereby generating positive emotions, significantly improving the user experience, and enhancing the comfort of the user's current environment.

[0061] In some embodiments of this specification, by analyzing the user's voice data and historical behavior data, the emotional characteristics and predicted behaviors of the user can be determined, and early warning of the user's emotions can be achieved through multi-modal means; by adopting diversified audio interventions (such as playing cheerful music, playing cross-talk to divert attention, etc.), and at the same time cooperating with lighting interventions in a specific atmosphere for positive emotional guidance, the user experience can be improved.

[0062] In some embodiments, the processor can receive the emotion intervention conditions and / or intervention preferences input by the user, and update the intervention strategy based on the emotion intervention conditions and / or intervention preferences.

[0063] The emotion intervention condition refers to the relevant conditions set by the user for judging whether intervention needs to be triggered and how to intervene. For example, the emotion intervention condition is that intervention is needed when angry and no intervention is needed when sad, etc.

[0064] The intervention preference refers to the personalized choice of the user in aspects such as intervention means. For example, the intervention preference can be to use at least one of only audio intervention, lighting intervention, and fragrance intervention, etc.

[0065] In some embodiments, the processor can update the emotion intervention conditions and / or intervention preferences based on the historical feedback data in the historical behavior data.

[0066] The historical feedback data refers to the user logs obtained after the intervention is issued at a historical time. In some embodiments, the processor can actively collect and record the user logs after the intervention is implemented, and use the user logs as the historical feedback data.

[0067] In some embodiments, the processor can query the third preset table based on the historical feedback data, and update the emotion intervention conditions and / or intervention preferences. The third preset table includes multiple historical feedback data and the corresponding adjustment strategies of emotion intervention conditions and / or intervention preferences, etc. The third preset table can be preset in advance according to experience or historical data.

[0068] An adjustment strategy refers to a plan for adjusting the current emotion intervention conditions and intervention preferences. In some embodiments, the adjustment strategy includes an adjustment amplitude and whether to turn off one or more intervention strategies, etc. For example, after a certain intervention, if within the first preset time length, the user turns off the lighting device and / or the audio device, etc., the intervention preference corresponding to the historical feedback data in the third preset table is updated to the intervention strategy of turning off the lighting device and / or the audio device, etc. (i.e., no intervention); if the lighting device and / or the audio device, etc. are turned off within a time length greater than the first preset time length and less than the second preset time length, the emotion intervention condition corresponding to the historical feedback data in the third preset table is updated to lower the parameter corresponding to the intervention strategy. For example, lower the first audio parameter corresponding to the audio intervention. The first preset time length and the second preset time length can be preset based on prior experience. The second preset time length is greater than the first preset time length.

[0069] In some embodiments of this specification, updating the emotion intervention conditions and intervention preferences based on the historical feedback data in the historical behavior data can improve personalized services, accurately adapt to the different needs of different users, reduce the frequency of ineffective interventions, further improve the accuracy of the subsequent determined intervention strategies, and further enhance the user experience.

[0070] In some embodiments, the processor can update the intervention strategy corresponding to the user based on the user input or the updated emotion intervention conditions and intervention preferences; at this time, the vector databases of different users are different. For example, when the emotion intervention condition is that no intervention is required when the user is sad, the processor can modify the need for intervention when the user is sad to no intervention. For another example, if the user's intervention preference is to only use audio intervention, the processor can only use audio intervention for this user. For another example, if the user's emotion intervention condition is to lower the parameter corresponding to the intervention strategy, the processor can lower the parameter corresponding to the intervention strategy for this user.

[0071] In some embodiments of this specification, by allowing the user to actively input emotion intervention conditions and / or intervention preferences, the user autonomy can be improved, the user control can be enhanced, and the user satisfaction can be improved; updating the intervention strategy based on the emotion intervention conditions and / or intervention preferences can enhance the intervention pertinence for different users and further enhance the user experience.

[0072] Figure 3 It is an exemplary schematic diagram of an emotion analysis model shown in some embodiments of this specification.

[0073] In some embodiments, as Figure 3 shown, the processor can determine the emotion feature 350 based on the user voice data 310 through the emotion analysis model 320. The emotion feature can be represented by the percentage of negative emotions. For more descriptions of the user voice data and the emotion feature, seeFigure 2 and its related content.

[0074] In some embodiments, before inputting into the sentiment analysis model, the processor also needs to collect and preprocess the user voice data 310. For example, perform noise reduction, filtering, artifact removal, alignment of different modality data, etc. on the user voice data.

[0075] The sentiment analysis model 320 is a model for determining sentiment features. In some embodiments, the sentiment analysis model can be a machine learning model. For example, the sentiment analysis model can be any one or a combination of a Convolutional Neural Network (CNN) model or other custom model structures, etc.

[0076] In some embodiments, the processor can train the sentiment analysis model based on a large number of second training samples with second labels. The training process of the sentiment analysis model is similar to that of the behavior prediction model and will not be elaborated here.

[0077] The second training samples include sample user voice data. The second training samples can be obtained from historical data / network data. The second labels include the sentiment features corresponding to the sample voice data. In some embodiments, the processor can obtain the second labels through automatic labeling.

[0078] In some embodiments, the sentiment analysis model 320 includes a feature extraction layer 321 and a sentiment analysis layer 322.

[0079] The feature extraction layer 321 is a layer structure for extracting acoustic features and semantic features, etc. of the user voice data. In some embodiments, the feature extraction layer is a machine learning model. For example, the feature extraction layer can include any one or a combination of a Convolutional Neural Network (CNN) model or other custom model structures, etc.

[0080] In some embodiments, the input of the feature extraction layer 321 includes the user voice data 310; the output includes the acoustic features 330 and the semantic features 340.

[0081] The acoustic features refer to the acoustic physical quantities related to the user voice data. For example, the acoustic features can include speech rate, intonation, and volume, etc.

[0082] The semantic features refer to the features that reflect the user's intention or sentiment tendency, etc. in the content of the user voice data. In some embodiments, the processor can characterize the semantic features through keywords and emotion words, etc. in the text of the user voice data.

[0083] In some embodiments, the processor may train the feature extraction layer based on a large number of third training samples with third tags. The training process of the feature extraction layer is similar to that of the behavior prediction model and will not be elaborated here.

[0084] The third training samples include sample user voice data. The third training samples may be obtained based on historical data / network data. The third tags include the acoustic features and semantic features corresponding to the third training samples. In some embodiments, the processor may obtain the third tags through automatic tagging.

[0085] The sentiment analysis layer 322 is used to perform sentiment analysis on the acoustic features and semantic features and obtain sentiment features. In some embodiments, the sentiment analysis layer is a machine learning model. For example, the sentiment analysis layer may include any one or combination of a deep neural networks (DNN) model or other custom model structures, etc.

[0086] In some embodiments, the processor may train the sentiment analysis layer based on a large number of fourth training samples with fourth tags. The training process of the sentiment analysis layer is similar to that of the behavior prediction model and will not be elaborated here.

[0087] The fourth training samples include sample acoustic features and sample semantic features. The fourth training samples may be obtained based on historical data / network data. The fourth tags include sample sentiment features. In some embodiments, the processor may perform normalization processing on the sentiment features of the sample acoustic features and the sentiment features of the sample semantic features and then perform weighted summation (e.g., Min-Max normalization, Z-score normalization, etc.) to obtain the corresponding sample sentiment features. The weighting values may be preset based on experience. In some embodiments, the weight of the sentiment features of the sample acoustic features is higher; the weight of the sentiment features of the sample semantic features is lower.

[0088] In some embodiments, the output of the feature extraction layer may be used as the input of the sentiment analysis layer. The feature extraction layer and the sentiment analysis layer may also be obtained through joint training.

[0089] In some embodiments, the sentiment features output by the sentiment analysis model may also include the scene type. For more descriptions about the scene type, see Figure 4 and its related content.

[0090] In some embodiments, the processor may determine the sentiment features based on the user voice data, image data, and physiological signals through the sentiment analysis model and / or user active tagging.

[0091] In some embodiments, the input of the feature extraction layer further includes image data and physiological signals, and the output further includes image features and physiological features.

[0092] In some embodiments, when the input of the feature extraction layer further includes image data and physiological signals, the third training sample further includes sample image data and sample physiological signals.

[0093] In some embodiments, the input of the sentiment analysis layer further includes image features and physiological features. When the input of the sentiment analysis layer further includes image features and physiological features, the fourth training sample further includes sample image features and sample physiological features.

[0094] Image data refers to the data of images captured by an image acquisition device. For example, the image data includes the user's facial expressions, body language, and the scene where the user is located. Among them, the scene where the user is located can be that the user is watching a movie or having a conversation with others. The image acquisition device can be a camera or the like.

[0095] Physiological signals refer to the physiological information of the user obtained through wearable devices or the like. For example, the physiological signals can include the user's heart rate, blood oxygen saturation rate, galvanic skin response, etc. The wearable device can be a smart bracelet, a mobile phone, an earphone, etc. The physiological signals can also be obtained by the user's input.

[0096] User active marking refers to the emotional features actively input by the user. For example, the user active marking can be that the user is in a happy mood at present.

[0097] In some embodiments, the processor can determine the emotional features through the sentiment analysis model and the user active marking. For example, the processor can determine the user active marking as the emotional feature corresponding to the data (including the user's voice data, image data, and physiological signals); for those that are not user active markings, the processor can determine the emotional features through the sentiment analysis model.

[0098] In some embodiments of the present specification, based on the user's voice data, image data, and physiological signals, the emotional features are determined through two methods: the sentiment analysis model and the user active marking, which significantly improves the accuracy of emotion recognition and further enhances the user experience; by combining multi-modal data, the emotional features can be determined more accurately. For example, the image data can identify the user's facial expressions, body language, and the scene where the user is located, and the physiological signals can obtain the heart rate, etc.

[0099] In some embodiments of the present specification, by fusing the acoustic features and semantic features, the emotion recognition accuracy can be improved and the risk of misjudgment can be reduced; through the hierarchical model architecture and the efficient training strategy, a high-precision sentiment analysis model is constructed to enhance the practicality of the system.

[0100] Figure 4 is an exemplary flowchart for determining the intervention strategy according to some embodiments of the present specification. As Figure 4 shown, process 400 includes the following steps. In some embodiments, process 400 can be executed by a processor.

[0101] Step 410: Determine whether the user's emotion is a positive emotion based on the emotional feature and the predicted behavior. For more descriptions of the emotional feature and the predicted behavior, see Figure 2 and its description.

[0102] A positive emotion refers to a positive and affirmative emotion. For example, positive emotions such as happiness. In some embodiments, the user's emotions include positive emotions and negative emotions, etc. For more descriptions of negative emotions, see Figure 3 and its description.

[0103] In some embodiments, the processor can determine whether the user's emotion is a positive emotion based on the emotional feature and the predicted behavior in various ways. For example, the processor can determine whether the emotional feature (percentage of negative emotions) output by the emotion analysis model is less than the emotion threshold and whether the degree of out-of-control of the predicted behavior is less than the out-of-control threshold, so as to determine whether the user's emotion is a positive emotion. Among them, the emotion threshold and the out-of-control threshold can be preset in advance. For example, both are 10%, etc. For more descriptions of the emotion analysis model, see Figure 3 and its related description. For the degree of out-of-control of the predicted behavior, see Figure 2 and its related description.

[0104] In some embodiments, the processor can also determine whether there is a positive emotion based on the user's active marking. For example, if the user actively marks that there is a positive emotion currently, the processor determines that there is a positive emotion.

[0105] Step 420: In response to the user's emotion being a positive emotion, generate an anchor point combination and perform at least two of the following operations: drive the audio device to play sound based on the second audio parameter, drive the lighting device to perform lighting based on the second lighting parameter, and drive the fragrance device to emit fragrance based on the fragrance parameter.

[0106] An anchor point combination refers to a combination of at least two parameters. For example, the anchor point combination can include at least two of the second audio parameter, the second lighting parameter, and the fragrance parameter. Among them, the anchor point can be at least one of the second audio parameter, the second lighting parameter, and the fragrance parameter.

[0107] The definitions of the second audio parameter, the second lighting parameter, and the fragrance parameter are similar to those of the first audio parameter, the first lighting parameter, and the first fragrance parameter. The difference is that the second audio parameter, the second lighting parameter, and the fragrance parameter are actively released when the emotion is positive in order to implant the corresponding parameters of the positive emotion anchor point; the first audio parameter, the first lighting parameter, and the first fragrance parameter are the corresponding parameters when the emotion needs to be intervened. In some embodiments, in order to distinguish from the above-mentioned first fragrance parameter, the fragrance parameter in step 420 can also be referred to as the second fragrance parameter.

[0108] In some embodiments, the second audio parameter, the second lighting parameter, and the fragrance parameter can be generated based on preset conditions. For example, parameters that meet the preset conditions are randomly selected. Among them, the preset conditions can include that the intensity of the anchor point is less than the intensity threshold and / or the change amplitude of the anchor point is less than the amplitude threshold. The intensity of the anchor point can be the volume of the second audio parameter, the brightness of the second lighting parameter, the fragrance concentration of the fragrance parameter, etc. The change amplitude of the anchor point can be the change amplitude of data such as volume, brightness, and fragrance concentration at adjacent time points. The intensity threshold and the amplitude threshold can be preset and determined based on experience. By setting the intensity threshold and the amplitude threshold, it is ensured that the second audio parameter, the second lighting parameter, and the fragrance parameter will not have an obvious impact on the user's current mood, so that positive emotion anchor points can be implanted inadvertently.

[0109] In some embodiments, the processor can generate anchor point combinations in various ways. For example, the processor can generate anchor point combinations that do not exist or have a low existence quantity in the current anchor point database based on the above-mentioned generated second audio parameter, second lighting parameter, and fragrance parameter, so as to enhance the binding of the anchor point combination to the corresponding positive emotion. Among them, the anchor point combination includes at least two of the second audio parameter, the second lighting parameter, and the fragrance parameter. The above-mentioned anchor point combination is a combination of at least two of the sound, lighting, and fragrance actively released by the system when the user is in a positive emotion.

[0110] In some embodiments, the processor can perform at least two of the following operations based on the parameter types included in the anchor point combination. For example, if the anchor point combination includes the second audio parameter and the second lighting parameter, the processor drives the audio device to play sound based on the second audio parameter and drives the lighting device to perform lighting based on the second lighting parameter.

[0111] Step 430, store the positive emotion, the first environmental data, and the anchor point combination into the anchor point database.

[0112] The first environmental data refers to the environmental data when the current user is in a positive emotion. In some embodiments, the first environmental data can include at least one of the first temperature, the first humidity, the first brightness, the first color temperature, etc. In some embodiments, the first environmental data can be obtained by devices such as temperature sensors, humidity sensors, and light sensors configured indoors.

[0113] The anchor database refers to a personalized database that stores information related to anchors. In some embodiments, the anchor database may include multiple first vectors. The first vector may include the correspondence between positive emotions, first environmental data, and anchor combinations. For example, the first vector may be in the following form: {8% negative positive emotion, (first temperature, first brightness, first color temperature,...), (second audio parameter, second lighting parameter, fragrance parameter)}. In some embodiments, different users are affected by anchor combinations to different degrees, so the anchor databases corresponding to different users are also different.

[0114] Step 440, based on the emotional characteristics and the second environmental data, determine an intervention strategy through the anchor database.

[0115] The second environmental data refers to the environmental data when the current user is in a situation that requires intervention. For example, the environmental data when the user is in a negative emotion. In some embodiments, the second environmental data may include at least one of a second temperature, a second humidity, a second brightness, and a second color temperature. The method for obtaining the second environmental data is similar to that of the first environmental data.

[0116] In some embodiments, the processor may determine an intervention strategy in multiple ways based on the emotional characteristics, the second environmental data, and the anchor database. For example, when it is determined based on the emotional characteristics that the user's emotion is sadness or anger, the processor may randomly select a first vector from the anchor database and use the anchor combination corresponding to the first vector as the intervention strategy.

[0117] For another example, the processor may determine an intervention strategy based on a preset rule. The preset rule may be determined based on manual presetting. Exemplarily, the preset rule is: match the second environmental data with the anchor database, select the anchor combination corresponding to the first environmental data in the first vector with the highest similarity to the second environmental data, and use the corresponding second audio parameter, second lighting parameter, and fragrance parameter of the anchor combination as the intervention strategy. Among them, the similarity may be calculated based on the degree of numerical proximity between the first environmental data and the second environmental data. Since the degree to which emotions are affected is different during the day and at night, by selecting an anchor combination close to the current environment, the intervention strategy can be better matched with the environment, thereby efficiently awakening the association of positive emotions.

[0118] In some embodiments, the anchor database further includes a sound anchor library, a light and shadow anchor library, and a fragrance anchor library.

[0119] The sound anchor library refers to an anchor database that stores only anchors related to audio data. In some embodiments, the sound anchor library may include a second vector. The second vector may be the correspondence between positive emotions, first environmental data, and second audio parameters.

[0120] The light and shadow anchor library refers to an anchor database that stores only the anchors related to lighting data. In some embodiments, the light and shadow anchor library may include a third vector. The third vector may be the correspondence between positive emotions, first environmental data, and second lighting parameters.

[0121] The fragrance anchor library refers to an anchor database that stores only the anchors related to fragrance data. In some embodiments, the fragrance anchor library may include a fourth vector. The fourth vector may be the correspondence between positive emotions, first environmental data, and fragrance parameters.

[0122] In some embodiments, the processor may determine an intervention strategy in multiple ways based on the emotional characteristics and second environmental data, and at least one of the sound anchor library, light and shadow anchor library, and fragrance anchor library. For example, when it is determined based on the emotional characteristics that the user's emotion is sad or angry, if the processor only selects the sound anchor library, the processor may match the second environmental data with the sound anchor library and select the second audio parameter in the second vector with the highest similarity between the first environmental data and the second environmental data as the intervention strategy.

[0123] In some embodiments of this specification, the anchor database further includes a sound anchor library, a light and shadow anchor library, and a fragrance anchor library. By selecting specific categories of anchor combinations, the randomness of determining the anchor combinations is introduced, which can prevent users from perceiving the rules of the anchors. In addition, the system can learn based on the subsequent intervention results to determine which categories of anchor combinations are more effective for specific users, so as to optimize the subsequent user experience.

[0124] In some embodiments, the first audio parameter may include a guiding audio. The guiding audio refers to the audio used to guide the user. For example, the voice that guides the user to have a subconscious association with the execution of the intervention strategy. In some embodiments, the guiding audio may be determined based on a fourth preset table. The fourth preset table may include the correspondence between the guiding audio and the intervention strategy. In some embodiments, the fourth preset table may be determined based on artificial preset.

[0125] In some embodiments, the guiding audio is associated with the first lighting parameter and / or the first fragrance parameter corresponding to the intervention strategy. For example, the guiding audio may include audio suggesting the fragrance parameter and / or audio suggesting the lighting parameter. In some embodiments, the content of the guiding audio includes indirect questions, ambiguous descriptive statements, etc. An indirect question refers to a way of asking questions that indirectly mentions the lighting parameter / fragrance parameter. For example, "It seems there is a special smell in the air. Can you smell it?" By asking an indirect question, it implies that the user senses the fragrance to generate associations. An ambiguous descriptive statement refers to a statement that ambiguously describes the lighting parameter / fragrance parameter. For example, "That fleeting light and shadow just now was a bit like that day..." After playing this section of audio, pause. By describing an ambiguous feeling, it implies that the user makes associations based on the lighting parameter. In some embodiments, the need for guiding audio means that on the basis that the intervention strategy includes audio intervention, it also includes at least one of lighting intervention and fragrance intervention. Otherwise, the processor does not need to determine the guiding audio and guide.

[0126] In some embodiments of this specification, by using the first audio parameter, the audio content tries not to directly reveal the existence of the anchor point (when implanting the anchor point) or the purpose of the intervention (during the intervention), but rather guides the user's attention to subconscious cues and promotes associations through indirect questions, describing ambiguous feelings, etc., so as to successfully complete the anchor point implantation or intervention under the condition that it is difficult for the user to notice.

[0127] In some embodiments, in response to the generation time of the anchor point combination being less than the time threshold, the intensity of the positive emotion signal of the anchor point combination matching the intensity of the emotional feature, and the second environmental data when generating the anchor point combination being closest to the first environmental data, the processor may determine the intervention strategy based on the anchor point combination.

[0128] The generation time of the anchor point combination refers to the time when the anchor point combination is stored in the anchor point database.

[0129] In some embodiments, the time threshold can be determined based on manual presetting. The generation time of the anchor point combination being less than the time threshold means that the anchor point combination was generated recently, and the effect of the intervention strategy generated using this anchor point combination will be better.

[0130] The positive emotion signal refers to the data representing the intensity of the positive emotion signal. For example, the smaller the percentage of the negative emotion corresponding to the positive emotion, the higher the intensity of the positive emotion signal. In some embodiments, the higher the intensity of the positive emotion signal, the stronger the positive emotion.

[0131] The intensity of the emotional feature refers to the data representing the intensity of the negative emotion signal. For example, the larger the percentage of the negative emotion corresponding to the emotional feature, the higher the intensity of the emotional feature. In some embodiments, the higher the intensity of the emotional feature, the stronger the negative emotion.

[0132] In some embodiments, the processor may determine whether the intensity of the positive emotion signal of the anchor combination matches the intensity of the emotional feature. For example, the processor may compare the intensity of the positive emotion signal of the anchor combination (which may be referred to as the first intensity) with the difference between 1 and the intensity of the emotional feature (the difference may also be referred to as the second intensity). The processor may determine whether the first intensity and the second intensity are the same or similar. Similar means that the numerical difference between the first intensity and the second intensity is less than the difference threshold. The difference threshold may be preset in advance, such as 1. Exemplarily, if the intensity of the positive emotion signal (i.e., the first intensity) is 5%, the intensity of the emotional feature is 95%, and the difference between 1 and the intensity of the emotional feature (i.e., the second intensity) is 5%, then the first intensity and the second intensity are the same. If the first intensity and the second intensity are the same or similar, the intensity of the positive emotion signal of the anchor combination matches the intensity of the emotional feature. If they match, it means that the anchor combination can correspond to a very strong positive emotion, so the intervention strategy generated using this anchor combination can be used to intervene in strong negative emotions.

[0133] In some embodiments, the processor may determine the proximity of the second environmental data to the first environmental data by calculating the similarity between the second environmental data and the first environmental data. The method for calculating the similarity has been mentioned above and will not be elaborated here.

[0134] In some embodiments, in response to the generation time of the anchor combination being less than the time threshold, the intensity of the positive emotion signal of the anchor combination matching the intensity of the emotional feature, and the second environmental data at the time of generating the anchor combination being closest to the first environmental data, the intervention effect of the intervention strategy generated using this anchor combination is better. Therefore, the processor may determine this anchor combination as the intervention strategy.

[0135] In some embodiments of this specification, by screening the generation time of the anchor combination, the intensity of the positive emotion signal, and the similarity between the second environmental data and the first environmental data at the time of generating the anchor combination, the intervention effect of the corresponding generated intervention strategy can be improved, further enhancing the user experience.

[0136] In some embodiments, in response to the image data including the scene where it is located, the emotional feature further includes the scene type; in response to the historical scene type at the time of generating the anchor combination matching the scene type, the processor may determine the intervention strategy based on the anchor combination. For the description of the image data, please refer to Figure 2 and its description.

[0137] The scene type refers to the type of the scene where the user is located. For example, home, exercise, office, etc. In some embodiments, the emotional features output by the emotion analysis model may include the scene type.

[0138] The historical scenario type refers to the scenario type in historical data. In some embodiments, the historical scenario type can be determined based on historical data. For example, the process can determine the scenario type in which the user is located when generating the anchor point combination in the historical data as the historical scenario type.

[0139] In some embodiments, the processor can determine whether the historical scenario type when generating the anchor point combination matches the scenario type. In response, the processor can determine the anchor point combination as an intervention strategy. If the historical scenario type is a sports type and the scenario type is a home type, since the scenario types are too different, it may cause the generated intervention strategy to be ineffective. Therefore, the selected anchor point combination needs to meet the scenario type matching.

[0140] In some embodiments of this specification, by taking the scenario type into consideration, the similarity between the anchor point combination and the user's needs can be improved, and the effect of the corresponding generated intervention strategy can be ensured.

[0141] In some embodiments of this specification, when the user is in a positive mood or engaged in a pleasant activity, the system pre-releases specific and uncommon anchor point combinations, pre-binds these anchor point combinations with the positive mood, and helps the user establish a connection between the anchor point combination and the positive mood. When it is subsequently detected that the user is in a negative mood, the system reactivates these infrequently realized anchor point combinations to help the user evoke positive emotions through subconscious association.

[0142] It should be noted that the above descriptions of processes 200 and 400 are only for illustration and explanation, and do not limit the scope of application of this specification. For those skilled in the art, various modifications and changes can be made to processes 200 and 400 under the guidance of this specification. However, these modifications and changes are still within the scope of this specification.

[0143] One or more embodiments of this specification provide an emotion intervention device based on multimodal analysis, including at least one processor and at least one memory. The at least one memory is used to store computer instructions, and the at least one processor is used to execute at least part of the computer instructions to implement the above emotion intervention method based on multimodal analysis.

[0144] One or more embodiments of this specification provide a computer-readable storage medium that stores computer instructions. When the computer reads the computer instructions in the storage medium, the computer executes the above emotion intervention method based on multimodal analysis.

[0145] The basic concepts have been described above. Obviously, for those skilled in the art, the above detailed disclosure is only an example and does not constitute a limitation to this specification. Although not explicitly stated here, those skilled in the art may make various modifications, improvements, and corrections to this specification. Such modifications, improvements, and corrections are proposed in this specification, so such modifications, improvements, and corrections still fall within the spirit and scope of the exemplary embodiments of this specification.

[0146] In addition, unless clearly stated in the claims, the order of the processing elements and sequences, the use of numbers, letters, or other names described in this specification is not used to limit the order of the processes and methods of this specification. Although some currently considered useful embodiments of the invention are discussed through various examples in the above disclosure, it should be understood that such details only serve the purpose of illustration. The appended claims are not limited to the disclosed embodiments. On the contrary, the claims are intended to cover all modifications and equivalent combinations that conform to the essence and scope of the embodiments of this specification. For example, although the system components described above can be implemented by hardware devices, they can also be implemented only through software solutions, such as installing the described system on existing servers or mobile devices.

[0147] Similarly, it should be noted that, in order to simplify the expression of the disclosure in this specification and thus help the understanding of one or more embodiments of the invention, in the previous description of the embodiments of this specification, sometimes multiple features are grouped into one embodiment, drawing, or description thereof. However, this disclosure method does not mean that the features required by the subject matter of this specification are more than those mentioned in the claims. In fact, the features of the embodiments are fewer than all the features of the individual embodiments disclosed above.

[0148] Finally, it should be understood that the embodiments described in this specification are only used to illustrate the principles of the embodiments of this specification. Other variations may also fall within the scope of this specification. Therefore, by way of example and not limitation, alternative configurations of the embodiments of this specification can be considered to be in accordance with the teachings of this specification. Accordingly, the embodiments of this specification are not limited to the embodiments explicitly introduced and described in this specification.

Claims

1. An emotion intervention system based on multimodal analysis, characterized in that, The system includes an emotion analysis module, a behavior prediction module, and an intervention module; The emotion analysis module is configured to determine emotion features based on user voice data; The behavior prediction module is configured to determine a predicted behavior based on historical behavior data; The intervention module is configured to: Based on the emotion features and the predicted behavior, determine whether to intervene and an intervention strategy, the intervention strategy including audio intervention and / or lighting intervention; In response, determine first audio parameters corresponding to the audio intervention, and drive an audio device to play sound based on the first audio parameters; And / or, determine first lighting parameters corresponding to the lighting intervention and drive the lighting device to perform lighting based on the first lighting parameters.

2. The system according to claim 1, wherein The emotion analysis module is further configured to: Based on the user voice data, determine the emotion features through an emotion analysis model, the emotion analysis model being a machine learning model; The emotion analysis model includes a feature extraction layer and an emotion analysis layer; the input of the feature extraction layer includes the user voice data, and the output includes acoustic features and semantic features; The input of the emotion analysis layer includes the acoustic features and the semantic features, and the output includes the emotion features.

3. The system according to claim 2, wherein The intervention module is further configured to determine whether the user's emotion is a positive emotion based on the emotion features and the predicted behavior; The system further includes an anchor point module, the anchor point module being configured to: In response to the user's emotion being the positive emotion, generate an anchor point combination, the anchor point combination including at least two of second audio parameters, second lighting parameters, and fragrance parameters, and perform at least two of the following operations: Drive the audio device to play sound based on the second audio parameters, drive the lighting device Perform lighting based on the second lighting parameters, drive a fragrance device to emit fragrance based on the fragrance parameters; Store the positive emotion, first environmental data, and the anchor point combination in an anchor point database; The intervention module is further configured to: Determine the intervention strategy through the anchor point database based on the emotion features and second environmental data.

4. The system according to claim 1, wherein The system further includes a customization module, the customization module being configured to receive emotion intervention conditions and / or intervention preferences input by a user, and update the intervention strategy based on the emotion intervention conditions and / or the intervention preferences.

5. A method for emotion intervention based on multimodal analysis, characterized in that, The method includes: Determine emotion features based on user voice data; Determine a predicted behavior based on historical behavior data; Based on the emotion features and the predicted behavior, determine whether to intervene and an intervention strategy, the intervention strategy including audio intervention and / or lighting intervention; In response, determine first audio parameters corresponding to the audio intervention, and drive an audio device to play sound based on the first audio parameters; and / or, determine first lighting parameters corresponding to the lighting intervention and drive the lighting device to perform lighting based on the first lighting parameters.

6. The method according to claim 5, characterized in that, The determining emotion features based on user voice data includes: Based on the user voice data, determine the emotion features through an emotion analysis model, the emotion analysis model being a machine learning model; The emotion analysis model includes a feature extraction layer and an emotion analysis layer; the input of the feature extraction layer includes the user voice data, and the output includes acoustic features and semantic features; the input of the emotion analysis layer includes the acoustic features and the semantic features, and the output includes the emotion features.

7. The method according to claim 6, wherein The method further includes: Determining whether the user's emotion is a positive emotion based on the emotion features and the predicted behavior; In response to the user's emotion being the positive emotion, generating an anchor point combination, the anchor point combination including at least two of a second audio parameter, a second lighting parameter, and a fragrance parameter, and performing at least two of the following operations: Driving the audio device to play sound based on the second audio parameter, driving the lighting device to perform lighting based on the second lighting parameter, and driving the fragrance device to emit fragrance based on the fragrance parameter; Storing the positive emotion, the first environmental data, and the anchor point combination in an anchor point database; Determining the intervention strategy through the anchor point database based on the emotion features and the second environmental data.

8. The method according to claim 5, wherein The method further includes: Receiving emotion intervention conditions and / or intervention preferences input by the user, and updating the intervention strategy based on the emotion intervention conditions and / or the intervention preferences.

9. An emotion intervention device based on multimodal analysis, characterized in that, The device includes at least one processor and at least one memory; The at least one memory is used for storing computer instructions; The at least one processor is used for executing at least part of the computer instructions to implement the method according to any one of claims 5 to 8.

10. A computer-readable storage medium, characterized in that, The storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 5 to 8 is implemented.

Citation Information

Cited By

  • Emotion intervention system based on behavior characteristics

    CN122091102A