Emotion intervention method based on audio and video dual-mode emotion recognition
Through the method based on audio and video dual-mode emotion recognition, the problem of low accuracy of emotional intervention in the prior art is solved, the accuracy and personalization of emotional intervention are achieved, and driving safety is improved.
Patent Information
- Application Number
- CN202510266901.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-24
AI Technical Summary
In the prior art, emotional intervention has low accuracy and poor effect, and lacks clear start and stop rules and triggering conditions, resulting in inaccurate timing and methods of emotional intervention, affecting the driver's emotional state and driving safety.
Using a dual-mode emotion recognition method based on audio and video, we use the initial emotional fusion result and real-time emotional fusion result to determine whether the start and stop conditions of the emotional intervention module are triggered, and tasks are performed according to the preset emotional intervention plan to ensure the accuracy and personalization of emotional intervention.
It improves the accuracy and effectiveness of emotional intervention, provides more accurate and personalized emotional intervention services, and enhances driving safety.
Smart Images

Figure CN120189601A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of emotion recognition, and particularly to an emotion intervention method based on audio-visual bimodal emotion recognition. Background Art
[0002] In the existing technologies and patent literatures, we have been able to see some technical solutions for regulating and intervening in the emotions of drivers based on emotion recognition results. These solutions usually use devices such as in-vehicle atmosphere lights and audio systems to provide corresponding stimuli according to the recognized emotional states of drivers in order to achieve the purpose of regulating emotions. However, although these solutions have realized the function of emotion intervention to a certain extent, there are still some significant problems in their practical applications.
[0003] First of all, these technical solutions often lack specific start-stop rules for emotion intervention tasks. When to start emotion intervention, when to stop it, and what the specific conditions for triggering the intervention are, these issues have not been clearly answered in the existing technologies. This results in that in practical applications, the timing and manner of emotion intervention may not be accurate, and may even have a counterproductive effect, further affecting the emotional state of the driver and driving safety.
[0004] Secondly, the emotion recognition methods in the existing technologies may not be precise and reliable enough. Emotion is a complex and variable psychological state, and accurately recognizing the emotional state of a driver requires highly precise and reliable algorithms and technologies. However, the emotion recognition methods in the existing technologies often have errors and limitations, and may not be able to accurately reflect the true emotional state of the driver.
[0005] Therefore, the present invention aims to propose a new driver emotion intervention method based on emotion recognition, which can clarify the start-stop rules and triggering conditions of emotion intervention to ensure the accuracy and effectiveness of emotion intervention. At the same time, the present invention will also adopt more advanced and precise emotion recognition technologies to improve the accuracy and reliability of emotion recognition, and provide more accurate and personalized emotion intervention services for drivers. Through the implementation of the present invention, we are expected to solve the problems existing in the existing technologies, improve driving safety, and contribute new strength to the development of intelligent driving technologies. Summary of the Invention
[0006] The purpose of the present invention is to provide an emotion intervention method based on audio-visual bimodal emotion recognition to solve the problems of low accuracy and poor effect of emotion intervention in the existing technologies.
[0007] To achieve the above object of the present invention, an embodiment of the present invention provides an emotion intervention method based on audio-visual bimodal emotion recognition, which includes the following steps:
[0008] Turn on the emotion intervention module and obtain the initial emotion fusion result;
[0009] Determine whether it is currently during the execution of the emotion intervention task;
[0010] If not, determine whether the initial emotion fusion result triggers the start condition of the emotion intervention module, where the start condition includes: the initial emotion fusion result appears continuously in the time series and the number of appearances exceeds a predetermined threshold; or, in a predetermined time interval traced back from the current time on the time series, the proportion of the number of times the initial emotion fusion result appears in the total number of times all emotion fusion results appear is greater than the first proportion threshold;
[0011] If so, execute the emotion intervention task according to the preset emotion intervention plan corresponding to the initial emotion fusion result and set this emotion intervention task as the current emotion intervention task, and continue to execute the current emotion intervention task.
[0012] As a further improvement of an embodiment of the present invention, where the "determine whether it is currently during the execution of the emotion intervention task" further includes:
[0013] If it is, continue to execute the current emotion intervention task.
[0014] As a further improvement of an embodiment of the present invention, where the process of the "continue to execute the current emotion intervention task" further includes:
[0015] Determine whether the real-time emotion fusion result triggers the start condition;
[0016] If so, determine whether the priority of the real-time emotion fusion result is greater than the priority of the initial emotion fusion result;
[0017] If it is greater, execute a new emotion intervention task according to the emotion intervention plan corresponding to the real-time emotion fusion result;
[0018] If it is not greater, continue to execute the current emotion intervention task.
[0019] As a further improvement of an embodiment of the present invention, where the process of the "continue to execute the current emotion intervention task" further includes:
[0020] Obtain the real-time emotion fusion result and determine whether the real-time emotion fusion result triggers the stop condition of the emotion intervention module, where the stop condition includes:
[0021] When the initial emotion fusion result triggering the start condition is a positive emotion or a neutral emotion, the stop condition is: the duration of executing the emotion intervention task is a predetermined duration; or, when the initial emotion fusion result triggering the start condition is a negative emotion, the stop condition is: the priority of the real-time emotion fusion result is lower than the priority of the initial emotion fusion result, and the sum of the proportions of the occurrences of Neutral and Happy emotions in the sliding time window is greater than the second proportion threshold;
[0022] If so, stop executing the current emotion intervention task;
[0023] If not, continue to execute the current emotion intervention task.
[0024] As a further improvement of an embodiment of the present invention, among them, the emotion intervention plan corresponding to the initial emotion fusion result includes multiple plans from the X1 intervention plan to the Xn intervention plan, where n is a natural number;
[0025] The process of "continuing to execute the current emotion intervention task" includes:
[0026] Sequentially execute multiple intervention tasks from the X1 intervention plan to the Xn intervention plan according to the predetermined duration until all the foregoing intervention tasks are traversed and executed.
[0027] As a further improvement of an embodiment of the present invention, among them, the process of "executing a new emotion intervention task according to the emotion intervention plan corresponding to the real-time emotion fusion result" includes:
[0028] During the execution of the new emotion intervention task, judge whether the real-time emotion fusion result triggers the stop condition of the emotion intervention module;
[0029] If triggered, stop executing the new emotion intervention task;
[0030] If not triggered, continue to execute the new emotion intervention task.
[0031] As a further improvement of an embodiment of the present invention, among them, the emotion intervention plan corresponding to the real-time emotion fusion result includes multiple plans from the Y1 intervention plan to the Yn intervention plan, where n is a natural number;
[0032] The process of "if not triggered, continue to execute the new emotion intervention task" includes:
[0033] Sequentially execute the emotion intervention tasks of multiple plans from the Y1 intervention plan to the Yn intervention plan according to the predetermined duration until all the foregoing intervention tasks are traversed and executed.
[0034] As a further improvement of an embodiment of the present invention, the emotion intervention module includes a music material library that matches the emotion fusion result triggering the start condition of the emotion intervention module. Each of the X1th to Xnth intervention schemes or the Y1th to Ynth intervention schemes includes playing a music clip included in the music material library that matches the respective emotion fusion result, and the predetermined duration is the playing duration of the music clip.
[0035] As a further improvement of an embodiment of the present invention, the emotion intervention scheme includes an independent operation scheme or a combined operation scheme of more than one of the in-vehicle fragrance, audio, and ambient light systems.
[0036] As a further improvement of an embodiment of the present invention, when the emotion intervention scheme includes an in-vehicle fragrance operation scheme, the gear of the in-vehicle fragrance is set according to a predetermined gear setting rule, where the predetermined gear setting rule is:
[0037] Based on the emotion fusion result that triggers the start of the emotion module, the confidence level of the emotion fusion result, or the A value in the PAD at the time of triggering, set the gear of the in-vehicle fragrance operation.
[0038] As a further improvement of an embodiment of the present invention, "based on the emotion fusion result that triggers the start of the emotion module, the confidence level of the emotion fusion result, or the absolute value of the A value in the PAD at the time of triggering, set the gear of the in-vehicle fragrance operation" specifically includes:
[0039] When the emotion fusion result is Happy or Neutral, the in-vehicle fragrance is turned off;
[0040] When the emotion fusion result is Fear, set the gear according to the emotion confidence threshold of Fear;
[0041] When the emotion fusion result is Angry or Sad, set the gear according to the emotion confidence level of Angry or Sad, or according to the absolute value of the A value in the PAD when Angry or Sad is triggered.
[0042] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0043] Based on the time series of emotion confidence levels, emotion fusion results, and triggering rules, trigger the start of the emotion intervention task and stop the emotion intervention task, thereby providing more accurate and personalized emotion intervention services for drivers. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 It is a flowchart of a method for emotion intervention based on audio-visual bimodal emotion recognition provided by an embodiment of the present invention;
[0045] Figure 2 Flowchart of an emotion intervention method based on audio - video bimodal emotion recognition provided by a preferred embodiment of the present invention;
[0046] Figure 3 Flowchart of the process of performing a new emotion intervention task in another embodiment of the present invention. Detailed implementation manners
[0047] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the drawings and in conjunction with the embodiments.
[0048] It should be pointed out that unless otherwise specified, all technical and scientific terms used in this application have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs.
[0049] In the present invention, unless otherwise stated, the orientation words such as "upper, lower, top, bottom" are usually in reference to the direction shown in the drawings, or in reference to the vertical, perpendicular or gravitational direction of the component itself; similarly, for ease of understanding and description, "inner, outer" refer to the inner and outer of the contour of each component itself, but the above orientation words do not limit the present invention.
[0050] In the present invention, the emotion recognition results of the audio - video bimodal are fused by using the enumeration weight method as disclosed in Patent CN119028377A. The calculation method for the fusion of the emotion recognition results of the audio - video bimodal is as follows:
[0051]
[0052] where n is a positive integer, L is the step size, and t nL represents the Lth, 2Lth, 3Lth, ……, n*Lth seconds, represents the confidence level of the emotion recognition of the video modality at the n*Lth second, represents the confidence level of the emotion recognition of the audio modality at the n*Lth second, represents the confidence level of the emotion recognition of the audio - video bimodal fusion at the n*Lth second.
[0053] In this embodiment, the fusion weight coefficient w1 of the video modality is 0.7, and the fusion weight coefficient w2 of the audio modality is 0.3.
[0054] After the above processing, what is finally obtained is a time series of emotion confidence levels, and the emotion category corresponding to the maximum value is used as the final emotion result after the audio - video bimodal fusion.
[0055] Enhance the comprehensiveness and accuracy of emotion recognition by fusing the emotion results of aligned audio-visual modalities. The video modality can provide rich visual information, such as facial expressions, body postures, etc.; the audio modality focuses on sound features, such as intonation, speech rate, timbre, etc. Fusing the classification results of these two modalities can make full use of their respective advantages, avoid information loss or misjudgment that may exist in independent modalities, and thus capture and understand emotions more comprehensively and accurately.
[0056] In an embodiment of the present invention, optionally, L can be set to 2s and substituted into the above formula (1) to obtain the following calculation formula (2):
[0057]
[0058] Where n is a positive integer, t 2n represents the 2nd, 4th, 6th,..., 2n seconds, indicating that both the audio and video are sliced at 2-second intervals; represents the confidence of emotion recognition of the video modality at the 2nth second, represents the confidence of emotion recognition of the audio modality at the 2nth second, represents the confidence of emotion recognition of the audio-visual dual-modal fusion at the 2nth second.
[0059] After the above processing, what is finally obtained is a time series of emotion confidences. The emotion category corresponding to the maximum value is used as the final emotion after audio-visual dual-modal fusion. The following table is an example of the emotion recognition results of audio-visual dual-modal fusion:
[0060]
[0061]
[0062] For the convenience of explanation, in this article, the emotion fusion results in the above table are denoted as En, and the PAD value sequences are denoted as Pn, An, Dn. Here, n is a positive integer, representing the nth result of dual-modal emotion fusion, rather than using 2-second intervals as the subscript. Pn, An, Dn represent the nth results of Pleasure, Arousal, Dominance, and the value range of PAD is from -3 to 3. For the convenience of explanation, the subscripts of En, Pn, An, Dn are not represented by positive even numbers, that is, not using positive even numbers with 2-second intervals as the subscript. Therefore, n represents the nth result of dual-modal emotion fusion.
[0063] In the present invention, the categories of emotion fusion results usually include the following: Angry, Happy, Neutral, Sad, Surprised.
[0064] In the present invention, Neutural is further defined as a neutral emotion, Happy is defined as a positive emotion, and the fusion result of the three emotions of Angry, Sad, and Fear is defined as a negative emotion.
[0065] Based on the above statements, in order to solve the problems of poor accuracy and low effectiveness of emotion intervention in the prior art, the present invention provides an emotion intervention method based on audio-visual bimodal emotion recognition.
[0066] The following Figures 1-3 and specific embodiments are used to further describe the present invention in detail.
[0067] Embodiment
[0068] As Figure 1 shown, an emotion intervention method based on audio-visual bimodal emotion recognition provided by the present invention includes the following steps:
[0069] S11. Turn on the emotion intervention module and obtain the initial emotion fusion result;
[0070] S13. Determine whether it is currently during the execution of the emotion intervention task;
[0071] S15. If not, determine whether the initial emotion fusion result triggers the start condition of the emotion intervention module;
[0072] S17. If so, execute the emotion intervention task according to the preset emotion intervention plan corresponding to the initial emotion fusion result and set the emotion intervention task as the current emotion intervention task, and continue to execute the current emotion intervention task.
[0073] It should be noted that in the present invention, the start conditions for triggering the emotion intervention module mainly include the following two methods:
[0074] The first triggering method: In the time series, the initial emotion fusion result continuously appears and the number of appearances exceeds a predetermined threshold, that is, the start condition is triggered.
[0075] Specifically, the number of consecutive appearances of the initial emotion fusion result can be recorded as SameEmotion, and this number SameEmotion is a positive integer. In this embodiment, optionally, the number SameEmotion can be set to 5 times.
[0076] A more specific description is as follows:
[0077] 1) Take the current time point as the time reference point and record the initial emotion fusion result as E n , so as to count the sequence segment {E n-SameEmotion , E n-SameEmotion+1, ……, the number of occurrences of the emotional fusion results of each category in En. For example, if the initial emotional fusion result has the largest number of occurrences, record this number of occurrences as k1.
[0078] 2) If k1 = SameEmotion, it means that the category of the initial emotional fusion result in the sequence segment {E n-SameEmotion 、E n-SameEmotion+1 、……、E n} is a certain independent emotion. At this time, the initial emotional fusion result triggers the start condition of the emotion intervention module, and then, the content of step S17 is executed.
[0079] The second triggering method: In the predetermined time interval traced forward from the current time on the time series, if the proportion of the number of occurrences of the initial emotional fusion result in the number of occurrences of all emotional fusion results is greater than the predetermined proportion threshold, the start condition is triggered.
[0080] To distinguish the predetermined proportion threshold in the following text, here, the predetermined proportion threshold here is denoted as the first proportion threshold.
[0081] More specifically, the predetermined time interval traced forward from the current time can be vividly defined as a time window on the time series, and this time window moves as the confirmation time moves on the time series, so it is further defined as a sliding time window.
[0082] During the real-time confirmation process, if the proportion of the number of occurrences of the initial emotional fusion result in the number of occurrences of all emotional fusion results in the sliding time window is greater than the first proportion threshold, the start condition of the emotion intervention module is triggered.
[0083] For the convenience of our further understanding, the duration size of the time window can be denoted as SlidingWindow1, and the proportion threshold is denoted as R start , in this embodiment, optionally, the size SlidingWindow1 of the time window can be set to 30s, and the proportion threshold R start can be set to 0.8.
[0084] Even further:
[0085] 1) When confirming the real-time emotional fusion result, use the real-time observation time as the time reference point, and record the emotional fusion result of this reference point as E n , that is, the initial emotional fusion result is E n . Count the sequence of emotional fusion results {E n-SlidingWindow1 、E n-SlidingWindow1+1 、……、En}, determine the proportion of the occurrence times of various types of emotions, and record them as R Angry 、R Happy 、R Neutral 、R Sad 、R Surprise 。
[0086] 2) If the proportion of the occurrence times of the initial emotion fusion result within the sliding time window is greater than or equal to R start 。At this time, the initial emotion fusion result triggers the start condition of the emotion intervention module, and then the content of step S17 is executed.
[0087] Of course, in step S15, when judging whether the initial emotion fusion result triggers the start condition of the emotion intervention module, it may also include the situation where the initial emotion fusion result does not trigger the start condition of the emotion intervention module. Optionally, step S13 can be returned.
[0088] In this embodiment, after step S17 is executed, steps S18 and S19 will be executed.
[0089] Step S18: Judge whether the real-time emotion fusion result triggers the start condition; it should be noted that in this step, the logic of judging whether the real-time emotion fusion result triggers the start condition is the same as the logic of the initial emotion fusion result triggering the start condition described above, and will not be elaborated here.
[0090] Step S19: If so, judge whether the priority of the real-time emotion fusion result is greater than the priority of the initial emotion fusion result.
[0091] In the present invention, the "priority" is defined according to the needs of the user, and the priority is specifically: Angry > Sad > Surprised > Happy > Neutral.
[0092] Furthermore, after step S19 is executed, step S201 can be selectively executed.
[0093] Step S201 is specifically: if the priority of the real-time emotion fusion result is not greater than the priority of the initial emotion fusion result, continue to execute the current emotion intervention task.
[0094] For example, when the real-time emotion fusion result is Sad and the initial emotion fusion result is Angry, then continue to execute the current emotion intervention task.
[0095] After step S201 is executed, step S202 will be executed.
[0096] Step S202 specifically includes: obtaining the real-time emotion fusion result and determining whether the real-time emotion fusion result triggers the stop condition of the emotion intervention module.
[0097] In the present invention, the stop rule for the driver's negative emotion intervention is as follows: after the intervention task for a certain negative emotion is triggered, continuously determine whether the real-time emotion fusion result of the driver turns for the better, and whether it turns from a negative emotion to a neutral or positive emotion. If the driver's emotion turns for the better, such as changing from Angry to Neutral, the intervention can be stopped.
[0098] It can be understood that in the present invention, the stop condition for triggering the emotion intervention module is slightly different from the triggering start condition. The difference is that the trigger stop can only be triggered during the execution of the emotion intervention task.
[0099] In the process described above, it can be understood that the emotion fusion result when triggering the start of the emotion intervention module includes: positive emotion, neutral emotion or negative emotion.
[0100] Corresponding to the type of the emotion fusion result for the trigger condition of the emotion intervention module, the stop conditions for triggering the emotion intervention module in the present invention include the following:
[0101] First, when the emotion fusion result at the time of triggering the start condition of the emotion module is a positive emotion or a neutral emotion, the stop condition is: the duration of executing the emotion intervention task is a predetermined duration.
[0102] Generally speaking, in the first trigger stop condition, after executing the emotion intervention task for the predetermined duration, the execution of the emotion intervention task is stopped. The predetermined duration is set by the system according to the user's setting. Optionally, the duration can also be set as the playing duration of the music segment in the emotion intervention plan.
[0103] Second, it needs to be judged in combination with the emotion fusion result at the time of triggering the start condition of the emotion module. Specifically, when the emotion fusion result at the time of triggering the start condition of the emotion module is a negative emotion, the stop condition for triggering the emotion module in this case is: by judging whether the "real-time emotion fusion result" improves, and whether the sum of the proportions of the number of occurrences of neutral emotion and / or positive emotion in the sliding time window is greater than or equal to R stop , to determine whether to trigger the stop condition.
[0104] Here, the definition of the "sliding time window" can refer to the description above.
[0105] In the second stop condition, more specifically, when the "real-time emotion fusion result" improves, by judging whether the sum of the proportions of the number of occurrences of Neutral emotion and Happy emotion in the sliding time window is greater than or equal to Rstop (Here, the size of the sliding time window is denoted as SlidingWindow2, and the proportional threshold is denoted as R stop , and the proportional threshold here can be denoted as the second proportional threshold). If the sum of the proportions of the occurrences of Neutral emotion and Happy emotion within the sliding time window ≥ R stop , then the stop condition is triggered.
[0106] More specifically described as:
[0107] 1) Here, the real-time emotion fusion result is still denoted as E n , and count the proportions of the occurrences of the emotion fusion results of various categories in the sequence segment {E n-SlidingWindow2 , E n-SlidingWindow2+1 , ……, E n}, denoted as R Angry , R Happy , R Neutral , R Sad , R Surprise , etc.;
[0108] 2) In this embodiment, when the priority of the "real-time emotion fusion result" after performing the emotion intervention task is lower than the "initial emotion fusion result" before performing the emotion intervention task, and, R Happy + R Neutral > R stop , that is, the sum of the proportions of the occurrences of Neutral and Happy emotions within the sliding time window is greater than or equal to R stop , then the intervention task is stopped.
[0109] It should be noted that the improvement of the "real-time emotion fusion result" specifically means that the priority of the "real-time emotion fusion result" after performing the emotion intervention task is lower than the "initial emotion fusion result" before performing the emotion intervention task. In this case, it indicates that the emotion has changed for the better after the intervention.
[0110] In this embodiment, after performing step S202, step S203 will be executed. Step S203 is specifically: If so, stop executing the current emotion intervention task.
[0111] Of course, after performing step S202, step S204 may also be executed. Step S204 is specifically: If not, continue to execute the current emotion intervention task.
[0112] In this embodiment, the emotion intervention plan corresponding to the initial emotion fusion result includes multiple plans from the X1st intervention plan to the X n th intervention plan, where n is a natural number;
[0113] The process of "continuing to execute the current emotion intervention task" in step S204 includes:
[0114] Sequentially execute multiple intervention tasks from the X1st intervention plan to the X n intervention plan according to a predetermined duration until all the foregoing intervention tasks are traversed and executed, that is, the execution duration of each intervention plan is the predetermined duration.
[0115] Specifically, when the real-time emotion fusion result does not trigger the stop condition of the emotion intervention module, the current emotion intervention task will continue to be executed. The emotion intervention plan executed by the current emotion intervention task can optionally be composed of multiple emotion intervention plans from X1 to X n During the execution of the emotion intervention task, there is still no trigger of the stop condition of the emotion intervention module after executing the emotion intervention plan X1 to the emotion intervention plan X n In this case, as long as the emotion intervention plans from X1 to X n are traversed and executed, the execution of the current emotion intervention task can be stopped.
[0116] In this embodiment, after step S19 is executed, step S211 may also be executed. Step S211 is specifically: if it is greater than, then execute a new emotion intervention task according to the emotion intervention plan corresponding to the real-time emotion fusion result.
[0117] It can be understood that the emotion intervention task executed according to the emotion intervention plan corresponding to the real-time emotion fusion result is also the new emotion intervention task.
[0118] In some embodiments, referring to Figure 3 , optionally, after step S211 is executed, step S2111 may also be turned to. Step S2111 is specifically: during the execution of the new emotion intervention task, judge whether the real-time emotion fusion result triggers the stop condition of the emotion intervention module. The logic of triggering the stop condition of the emotion intervention module here is the same as that described above and will not be repeated here.
[0119] After step S2111 is executed, step S2112 may be executed. Step S2112 is specifically: if triggered, stop executing the new emotion intervention task.
[0120] Of course, after step S2111 is executed, step S2113 may also be executed. Step S2113 is specifically: if not triggered, continue to execute the new emotion intervention task.
[0121] As described above, the new emotion intervention task is also the emotion intervention task executed according to the emotion intervention plan corresponding to the real-time emotion fusion result. The foregoing emotion intervention plan includes the Y1st intervention plan to the Y nMultiple programs including an intervention program, where n is a natural number;
[0122] The process of "if not triggered, continue to execute a new emotion intervention task" described in step S2113 specifically includes:
[0123] Execute the emotion intervention tasks of multiple programs from the Y1st intervention program to the Yth intervention program in sequence according to a predetermined duration until all the foregoing intervention tasks are traversed and executed. The logic here is similar to the above and will not be elaborated here. n In the present invention, the emotion intervention module includes a music material library that matches the emotion fusion result triggering the start condition of the emotion intervention module. Each of the X1st to Xth intervention programs or the Y1st to Yth intervention programs includes playing a music segment from the music material library that matches its respective emotion fusion result, and the predetermined duration is the playing duration of the music segment.
[0124] In the present invention, the emotion intervention module includes a music material library that matches the emotion fusion result triggering the start condition of the emotion intervention module. Each of the X1st to Xth intervention programs or the Y1st to Yth intervention programs includes playing a music segment from the music material library that matches its respective emotion fusion result, and the predetermined duration is the playing duration of the music segment. n intervention programs or the Y1st to Y n intervention programs includes playing a music segment from the music material library that matches its respective emotion fusion result, and the predetermined duration is the playing duration of the music segment.
[0125] The matching relationship between the emotion fusion result and the music material library can be referred to in the following table:
[0126] Emotions triggering intervention Music material library Angry Neutral or Happy Sad Neutral or Happy Fear Neutral Neutral None Happy None
[0127] With such a setting, it can not only meet the use of a specific emotion fusion result with its corresponding music library to make the effect of emotion intervention better, and moreover, the playing time of the music segment can control the execution time of each of the X1st to X n and Y1st to Y n Each program's execution time, and the playing duration of the music segment is also the "predetermined duration" mentioned above.
[0128] It should be noted that in this invention patent, the emotion intervention program refers to a set of control methods for in-vehicle fragrance, audio, and ambient light systems. Different programs represent different control methods. When the driver shows or exhibits a certain emotion or physiological state, a specific program (pre-plan) will be triggered. The different control methods have two meanings. One is that the in-vehicle fragrance, audio, and ambient light systems each have independent multiple sets of programs; the other is the different combinations between the control programs of the in-vehicle fragrance, audio, and ambient light systems. For example, ambient light program 1, fragrance program 1, and audio program 1 are one combination, and ambient light program 2, fragrance program 2, and audio program 1 are another combination. Each program must include the control programs of the in-vehicle fragrance, audio, and ambient light systems.
[0129] In this invention patent, an emotion intervention task refers to the in-vehicle fragrance, audio, and ambient light systems performing specific tasks according to a certain scheme. For example, the ambient light system performs tasks such as flashing the light color, alternating flashing and breathing, the audio system performs the task of playing music, and the fragrance system performs the task of spraying essential oil in gear 1, etc. The tasks have some attributes, such as the number of task executions, the duration of task execution, the start time of the task, the end time of the task, the task status (running, stopped), etc. Since the in-vehicle fragrance, audio, and ambient light systems each perform a certain task, the execution duration of the tasks is usually different. For example, the execution of an ambient light task usually takes 5 - 10 seconds each time, while the execution of an audio task usually takes 60 - 90 seconds each time.
[0130] It can be understood that the emotion intervention scheme in the present invention may include an independent operation scheme of any one of the in-vehicle fragrance, audio, and ambient light systems or a combined operation scheme of more than one.
[0131] In this embodiment, preferably, when the emotion intervention scheme includes an in-vehicle fragrance operation scheme, the gear of the in-vehicle fragrance is set according to a predetermined gear setting rule, where the gear setting rule is: setting the gear of the in-vehicle fragrance operation based on the emotion fusion result, emotion confidence level, or the A value in the PAD when triggering the start of the emotion intervention task.
[0132] Specifically, the "gear setting rule is to set the gear of the in-vehicle fragrance operation based on the emotion fusion result, emotion confidence level, or the A value in the PAD when triggering the start of the emotion intervention task" includes:
[0133] When the emotion fusion result is Happy or Neutral, the in-vehicle fragrance is turned off;
[0134] When the emotion fusion result is Fear, set the gear according to the emotion confidence level threshold of Fear;
[0135] When the emotion fusion result is Angry or Sad, set the gear according to the emotion confidence level of Angry or Sad, or according to the absolute value of the A value in the PAD when Angry or Sad is triggered.
[0136] Referring to the following table, various situations of setting the gear of the in-vehicle fragrance operation are shown:
[0137]
[0138] To sum up, the embodiments of the present invention achieve the following technical effects:
[0139] First, triggering the start of the emotion intervention task and stopping the emotion intervention task based on the time series of emotion confidence level, emotion fusion result, and triggering rule, thereby providing more accurate and personalized emotion intervention services for the driver;
[0140] Second, setting the gear of the in-vehicle fragrance according to a predetermined gear setting rule can scientifically control the operation of the in-vehicle fragrance and improve the accuracy and interference effect of mood intervention.
[0141] Obviously, the above-described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0142] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0143] It should be noted that the terms "first", "second", etc. in the description, claims and drawings of the present application are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein.
[0144] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An emotion intervention method based on audio and video dual-modal emotion recognition, characterized in that: The steps include: Start the emotion intervention module and obtain the initial emotion fusion results; Determine whether the current task is being performed; If not, it is determined whether the initial emotion fusion result triggers the start-up condition of the emotion intervention module, wherein the start-up condition includes: the initial emotion fusion result appears continuously in the time series and the number of occurrences exceeds a predetermined threshold; or, in a predetermined time interval starting from the current time and tracing back on the time series, the number of occurrences of the initial emotion fusion result accounts for the number of occurrences of all emotion fusion results The ratio is greater than a first ratio threshold; If so, the emotion intervention task is executed according to the preset emotion intervention plan corresponding to the initial emotion fusion result and the emotion intervention task is set as the current emotion intervention task, and the current emotion intervention task is continued to be executed.
2. The emotion intervention method based on audio and video dual-modal emotion recognition according to claim 1, characterized in that: The “determining whether the current task is being executed” also includes: If in, continue to perform the current emotion intervention task.
3. The emotion intervention method based on audio and video dual-modal emotion recognition according to claim 1 or 2, characterized in that: The process of "continuing to perform the current emotion intervention task" also includes: Determine whether the real-time emotion fusion result triggers the start condition; If so, determining whether the priority of the real-time emotion fusion result is greater than the priority of the initial emotion fusion result; If it is greater than, a new emotion intervention task is executed according to the emotion intervention plan corresponding to the real-time emotion fusion result; If not, continue to execute the current emotion intervention task.
4. The emotion intervention method based on audio and video dual-modal emotion recognition according to claim 3 is characterized in that: The process of "continuing to perform the current emotion intervention task" also includes: Obtain the real-time emotion fusion result and determine whether the real-time emotion fusion result triggers the stop condition of the emotion intervention module, wherein the stop condition includes: When the initial emotion fusion result of the triggering start condition is a positive emotion or a neutral emotion, the stopping condition is: the duration of executing the emotion intervention task is a predetermined duration; or, when the initial emotion fusion result of the triggering start condition is a negative emotion, the stopping condition is: the priority of the real-time emotion fusion result is lower than the priority of the initial emotion fusion result, and the sum of the proportions of the occurrence times of Neutral and Happy emotions in the sliding time window is greater than the second proportion threshold; If yes, stop executing the current emotion intervention task; If not, continue to perform the current emotion intervention task.
5. The emotion intervention method based on audio and video dual-modal emotion recognition according to claim 4 is characterized in that: The emotional intervention plans corresponding to the initial emotional fusion results include intervention plans X1 to X2. n Multiple plans including intervention plans, where n is a natural number; The process of "continuing to perform the current emotion intervention task" includes: Execute intervention plan X1 to X1 in sequence according to the scheduled time. n The multiple intervention tasks including the intervention plan are performed until all the aforementioned intervention tasks are traversed and executed.
6. The emotion intervention method based on audio and video dual-modal emotion recognition according to claim 4, characterized in that: The process of "executing a new emotion intervention task according to the emotion intervention scheme corresponding to the real-time emotion fusion result" includes: During the execution of the new emotion intervention task, determining whether the real-time emotion fusion result triggers the stop condition of the emotion intervention module; If triggered, stop executing the new emotion intervention task; If not triggered, continue with the new emotion intervention task.
7. The emotion intervention method based on audio and video dual-modal emotion recognition according to claim 6, characterized in that: The emotional intervention plans corresponding to the real-time emotional fusion results include the Y1 intervention plan to the Y n Multiple plans including intervention plans, where n is a natural number; The process of "if not triggered, continue to perform a new emotion intervention task" includes: Execute intervention plan Y1 to Y2 in sequence according to the scheduled time. n The emotional intervention tasks of multiple plans including the intervention plan are carried out until all the aforementioned intervention tasks are traversed and executed.
8. The emotion intervention method based on audio and video dual-modal emotion recognition according to claim 5 or 7, characterized in that: The emotion intervention module includes a music material library that matches the emotion fusion result of the start condition that triggers the emotion intervention module, and the X1 intervention scheme to the X n Intervention plan or intervention plan Y1 to Y n The intervention plans all include playing a music clip included in the music material library that matches the respective emotion fusion results, and the predetermined duration is the playing duration of the music clip.
9. The emotion intervention method based on audio and video dual-modal emotion recognition according to claim 8, characterized in that: The emotion intervention scheme includes any one of the independent operation schemes of the vehicle fragrance, audio and ambient light systems or more than one combined operation scheme.
10. The emotion intervention method based on audio and video dual-modal emotion recognition according to claim 9, characterized in that: When the emotion intervention scheme includes a vehicle-mounted fragrance operation scheme, the gear of the vehicle-mounted fragrance is set according to a predetermined gear setting rule, wherein the predetermined gear setting rule is: The gear of the vehicle fragrance operation is set based on the emotion fusion result that triggers the activation of the emotion module, the confidence of the emotion fusion result, or the A value in the PAD when triggered.
11. The emotion intervention method based on audio and video dual-modal emotion recognition according to claim 10, characterized in that: "Setting the gear of the vehicle fragrance operation based on the emotion fusion result that triggers the activation of the emotion module, the confidence of the emotion fusion result, or the absolute value of the A value in the PAD when triggered" specifically includes: When the emotion fusion result is Happy or Neutral, the car fragrance is turned off; When the emotion fusion result is Fear, the gear is set according to the emotion confidence threshold of Fear; When the emotion fusion result is Angry or Sad, the gear is set according to the emotion confidence of Angry or Sad, or according to the absolute value of the A value in the PAD when Angry or Sad is triggered.