A driver emotion recognition method
Patent Information
- Application Number
- CN202510195526.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2026-08-21
AI Technical Summary
然而在现有技术中,多种模态的权重仅取决于驾驶场景,而忽视了其他因素(例如,不同驾驶员之间的个体差异)的影响,导致情绪状态的识别不够准确
[0020] According to the plan, by intervening in the driver's emotional state in a timely manner, the risk of traffic accidents caused by the driver's unstable emotional state can be reduced.
Smart Images

Figure CN122618596A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for driver emotion recognition, and more specifically, to a more accurate method for driver emotion recognition. Background Technology
[0002] A driver's emotional state is crucial while driving. An unstable emotional state can lead to a lack of concentration, increasing the risk of traffic accidents. Therefore, it is necessary to monitor the driver's emotional state to alert them when abnormalities occur and to proactively intervene.
[0003] Emotional values representing a driver's emotional state can be obtained based on data from multiple modalities (e.g., voice modality, driving behavior modality, physiological modality, and visual modality) to determine whether the driver's emotional state is abnormal. However, in existing technologies, the weights of multiple modalities depend only on the driving scenario, ignoring the influence of other factors (e.g., individual differences between different drivers), resulting in inaccurate identification of emotional states.
[0004] Therefore, it is hoped that a driver emotion recognition method can be proposed to improve the shortcomings of existing technologies. Summary of the Invention
[0005] According to a first aspect of the present invention, a driver emotion recognition method is proposed, comprising: collecting historical data; determining the corresponding scenario based on the historical data; for each of multiple scenarios, determining the weight of each of multiple modalities representing the driver's emotional state; determining the normal emotional range for each of the multiple scenarios based on the historical data and the weights; acquiring real-time data; determining the corresponding scenario based on the real-time data; determining a real-time emotion value based on the weight of the corresponding modality for that scenario and the real-time data; if the real-time emotion value is within the normal emotional range of that scenario, then determining that the driver's emotional state is normal; if the real-time emotion value is not within the normal emotional range of that scenario, then determining that the driver's emotional state is abnormal.
[0006] According to this scheme, since the normal emotional range for the corresponding scenario is determined by historical data, it varies depending on the different characteristics of different drivers, making the identification of the driver's emotional state more accurate.
[0007] In some schemes, historical data may include sentiment scores corresponding to multiple modalities in multiple scenarios. The sentiment value in each scenario is equal to the weighted average of the sentiment scores of each modality in that scenario. The normal sentiment range in that scenario is determined based on multiple sentiment values in that scenario.
[0008] In some solutions, the normal emotional range , where μ is the average of multiple emotion values, σ is the standard deviation of multiple emotion values, and k is the volatility coefficient.
[0009] In some schemes, the volatility coefficient depends on the scenario type.
[0010] According to this scheme, different fluctuation coefficients can be determined depending on the type of scenario, making the identification of the driver's emotional state more accurate.
[0011] In some schemes, the weights of each modality in each scenario can be updated based on an autoregressive method.
[0012] In some schemes, the weights of each modality in each scenario can be dynamically adjusted based on the recognition of the occurrence of a specific event.
[0013] According to this scheme, the weight of each modality in a specific scenario can be adjusted to more accurately represent the degree of influence of that modality on the driver's emotional state, making the identification of the driver's emotional state more accurate.
[0014] In some schemes, multiple modalities may include voice modality, driving behavior modality, physiological modality, and visual modality.
[0015] In some schemes, a speech modality may also include multiple speech submodalities, which include speech rate, volume, intonation, and semantic features.
[0016] In some solutions, multiple scenarios can include urban congestion, highways, and plateau sections.
[0017] In some schemes, the weight of physiological modalities in high-altitude road scenarios can be higher than that in urban congestion and highway scenarios; and / or the weight of visual modalities in high-altitude road scenarios can be lower than that in urban congestion and highway scenarios; and / or the weight of speech modalities in high-altitude road scenarios can be lower than that in urban congestion and highway scenarios.
[0018] In some scenarios, the fluctuation coefficient in urban congestion scenarios can be higher than that in highway scenarios.
[0019] According to a second aspect of the present invention, a driver emotion intervention method is proposed, comprising: acquiring the driver's emotional state based on the driver emotion recognition method described in the first aspect of the present invention; if the emotional state is abnormal, intervening with the driver; if the emotional state is normal, not intervening with the driver.
[0020] According to the plan, by intervening in the driver's emotional state in a timely manner, the risk of traffic accidents caused by the driver's unstable emotional state can be reduced. Attached Figure Description
[0021] Figure 1 A schematic diagram of a driver emotion recognition method according to an embodiment of the present invention is shown;
[0022] Figure 2 A schematic diagram of a driver emotion intervention method according to an embodiment of the present invention is shown;
[0023] Figure 3 A schematic diagram of normal emotional ranges according to an embodiment of the present invention is shown;
[0024] Figure 4 A schematic diagram of a driver emotion recognition system according to an embodiment of the present invention is shown.
[0025] Figure Labels
[0026] 10. Emotion Recognition System
[0027] 11 Data Collection Module
[0028] 12 Scene Recognition Module
[0029] 13 Calculation Module
[0030] 14. Emotion Judgment Module Detailed Implementation
[0031] To make the objectives, solutions, and advantages of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Unless otherwise stated, the terms used herein have their ordinary meanings in the art. The same reference numerals in the drawings represent the same parts.
[0032] As indicated in this application and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" are not specifically singular and may include plural forms. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of steps and elements that have been identified, and these steps and elements do not constitute an exclusive list; the method or apparatus may also include other steps or elements.
[0033] While this application makes various references to certain modules of the systems according to embodiments of this application, any number of different modules can be used and run on the driver terminal and / or server. The modules are merely illustrative, and different aspects of the systems and methods may use different modules.
[0034] Flowcharts are used in this application to illustrate the operations performed by the system according to embodiments of this application. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, various steps can be processed in reverse order or simultaneously as needed. Furthermore, other operations can be added to these processes, or one or more steps can be removed from them.
[0035] A driver's emotional state while driving is crucial to driving safety. An unstable emotional state can lead to a loss of concentration and even, in an attempt to vent emotions, a dangerous act of ramming into crowds, thus increasing the risk of traffic accidents. Therefore, it is necessary to monitor drivers' emotional states in real time so that timely intervention can be implemented when abnormalities occur. Figure 1 A schematic diagram of a driver emotion recognition method according to an embodiment of the present invention is shown. The method generally scores the driver's emotional state in a specific scenario, and then compares the score with the normal range of emotional state to determine whether the driver's emotional state is abnormal.
[0036] To improve the accuracy of driver emotional state assessment, the driver's emotional state is quantified. For example, the driver's level of anger is scored based on various detected information (e.g., 0 points represents no anger at all, and 10 points represents the most severe level of anger). Information that may characterize the driver's emotional state can include the following four modalities: vocal modality, driving behavior modality, physiological modality, and visual modality. It should be understood that the above four modalities are merely exemplary, and the present invention is not intended to limit specific modal types, but can include any type of information capable of characterizing the driver's emotional state.
[0037] Speech modality refers to the driver's voice. The driver's emotional state can be inferred to some extent from their speech patterns (e.g., vocal features and semantic information). Speech modality can also include several submodalities, such as speech rate, volume, tone, and semantic features. Speech rate indicates the speed at which the driver speaks; generally, faster speech indicates a higher level of anger. Volume indicates the loudness of the driver's speech; generally, louder speech indicates a higher level of anger. Tone indicates the frequency variation of the driver's voice; generally, higher frequency speech indicates a higher level of anger. Semantic features include the content of the driver's speech (e.g., vocabulary and expressions); generally, the more negative words (e.g., terrible, annoyed, angry, etc.) the driver uses in their speech, the higher their level of anger. Anger scores obtained from the submodalities of the aforementioned speech modality can be weighted and averaged to obtain the anger score corresponding to the speech modality. For example, if the weight of speech rate is 20%, the anger score based on speech rate is 5; if the weight of volume is 25%, the anger score based on volume is 4; if the weight of intonation is 20%, the anger score based on intonation is 3; and if the weight of semantic features is 35%, the anger score based on semantic features is 8. Thus, the anger score based on the speech modality is 5*20%+4*25%+3*20%+8*35%=5.4.
[0038] Driving behavior modality refers to a driver's driving behavior, specifically their operation of the steering wheel, accelerator, and brake. Generally, if a driver uses significantly more force and frequency than necessary to operate the driving equipment, it may indicate a higher level of anger. For example, a driver repeatedly and forcefully honking the horn may indicate that they are in a state of extreme anger.
[0039] Physiological modalities refer to a driver's physiological indicators, specifically heart rate, blood pressure, etc. Generally, higher heart rate and blood pressure indicate a higher level of anger.
[0040] Visual modality refers to the driver's facial expressions. Generally, the more serious the driver's facial expression, the higher the level of anger the driver may be.
[0041] The anger scores obtained from the four modalities mentioned above can be weighted and averaged to obtain the driver's anger score in the corresponding scenario. For example, the weight of the voice modality is 10%, and the anger score obtained based on the voice modality is 5; the weight of the driving behavior modality is 30%, and the anger score obtained based on the driving behavior modality is 4; the weight of the physiological modality is 40%, and the anger score obtained based on the physiological modality is 3; the weight of the visual modality is 20%, and the anger score obtained based on the visual modality is 8. Thus, the driver's anger score is 5*10%+4*30%+3*40%+8*20%=4.5.
[0042] The weights corresponding to the different modalities and submodals are jointly determined by the scenario type and user characteristics. Specifically, an initial modal weight distribution is set based on the scenario type, and then the modal weight distribution is dynamically adjusted based on user behavior. Furthermore, the modal weight distribution is updated on a rolling basis based on historical data. The details of how to set the weights will be described in detail below.
[0043] Back Figure 1 The driver emotion recognition method includes the following steps:
[0044] (1) Collect historical data
[0045] Historical data primarily includes various data influencing driver emotion scores and data assisting in determining the scenario type. The data influencing driver emotion scores are data across different modalities, as described above, such as the driver's speaking volume, heart rate, facial expressions, and steering wheel pressure. Data assisting in determining the scenario type may include GPS map data, road condition data acquired by cameras, and altitude data.
[0046] It should be understood that historical data refers to data collected before the moment the driver's emotional state was determined. For example, if the current time is December 1, 2024, historical data refers to data collected before December 1, 2024. Specifically, historical data can be data collected within 30 days prior to December 1, 2024 (i.e., from November 1, 2024 to November 30, 2024). The 30-day time span selected for historical data is merely illustrative; any other suitable time span (e.g., 10 days, 20 days, or 60 days, etc.) can be selected depending on the specific circumstances.
[0047] (2) Identify historical scenes
[0048] The scenario type corresponding to the collected historical data is determined. The scenario type refers to the external environment in which the vehicle is driving (e.g., urban congestion, highway, and plateau). Specifically, if GPS map data collected at 9:00 AM on November 10, 2024, indicates that the vehicle is driving in an urban congestion area, then the scenario type at that time is identified as urban congestion; if GPS map data collected at 2:00 PM on November 15, 2024, indicates that the vehicle is driving on a highway, then the scenario type at that time is identified as highway; if altitude data collected at 12:00 PM on November 20, 2024, indicates that the vehicle is driving in a plateau area, then the scenario type at that time is identified as plateau.
[0049] It should be understood that the types of scenarios involved in this invention are not limited to these, and may also include situations such as slippery roads due to heavy rain, and low visibility due to dense fog. Drivers exhibit different modal responses and emotional states in different scenario types, therefore, it is necessary to analyze the driver's emotional state separately for each scenario type. For example, in a congested urban scenario, a higher heart rate may indicate a higher level of anger; however, in a high-altitude road scenario, a higher heart rate may simply be due to altitude sickness and is unrelated to the driver's emotional state.
[0050] (3) Assigning initial weights
[0051] As mentioned above, drivers exhibit different modal responses in different scenario types. Therefore, it is necessary to set corresponding weight distributions for different modalities for specific scenario types.
[0052] In urban congestion scenarios, drivers are likely to argue with other drivers, passengers, or even pedestrians on congested roads. Therefore, a higher weight can be assigned to the speech modality, such as 40%. In addition, drivers are likely to use negative words during arguments, so a higher weight can be assigned to the semantic feature submodalities within the speech modality accordingly.
[0053] In a highway scenario, drivers are less likely to argue with other people, so a lower weight can be assigned to the speech modality, such as 10%. Furthermore, drivers are also less likely to use negative words, so a correspondingly lower weight can be assigned to the semantic feature sub-modalities within the speech modality.
[0054] In high-altitude driving scenarios, drivers' physiological characteristics are more unstable compared to those in plains areas and are more susceptible to emotional influences. Therefore, a higher weight can be assigned to the physiological modality, such as 45%. Furthermore, drivers at high altitudes are more prone to facial expressions due to intense sunlight exposure. In other words, a driver's serious expression at high altitudes is likely not due to anger but rather to the intense sunlight. Therefore, a lower weight can be assigned to the visual modality, such as 15%.
[0055] The weight distributions for the different scenarios described above can be stored in a standard scenario weight template library, which users can use as the initial weight distribution. If a specific event occurs during use, the weight distribution in the template library can be adjusted accordingly. For example, if driver A is identified as frequently arguing with other people in urban congestion scenarios, the weight of the speech modality in urban congestion scenarios can be increased accordingly.
[0056] (4) Calculate the normal emotional range
[0057] For each specific type of scenario, a weighted average of emotion scores across different modalities is used to obtain a large number of different time points. It can be assumed that the driver's emotion scores follow a normal distribution; fluctuations around the average emotion score can be considered normal, while fluctuations significantly deviating from the average can be considered abnormal. Therefore, a normal emotion range... σ is the average of multiple emotion values, σ is the standard deviation of multiple emotion values, and k is the volatility coefficient. k can be any real number greater than 0, and is generally between 1 and 3. For example, when k=1, the normal emotion range contains approximately 68% of normal emotion fluctuations, and the remaining approximately 32% of emotion values outside the normal range can be considered abnormal. Alternatively, when k=2, the normal emotion range contains approximately 95% of normal emotion fluctuations, and the remaining approximately 5% of emotion values outside the normal range can be considered abnormal. Alternatively, when k=2, the normal emotion range contains approximately 99.7% of normal emotion fluctuations, and the remaining approximately 0.3% of emotion values outside the normal range can be considered abnormal.
[0058] The fluctuation coefficient k can be determined based on different scenario types to make the identification of the driver's emotional state more accurate. For example, the fluctuation coefficient in urban congestion scenarios can be higher than that in highway scenarios. Specifically, the fluctuation coefficient can be set to 2 in urban congestion scenarios and 1 in highway scenarios. This is because drivers' emotional states are more prone to fluctuation in urban congestion scenarios compared to highway scenarios, so setting a higher fluctuation coefficient can more accurately indicate whether the driver's emotional state is abnormal. It should be understood that fluctuations in a driver's emotional state do not necessarily mean that their emotional state is abnormal; "abnormal" refers to the driver's emotional state fluctuation exceeding their usual fluctuation level.
[0059] Specifically, in the context of urban congestion, the emotional values obtained at 10 different times were 3, 6, 5, 8, 7, 4, 5, 9, 1, and 2. The average of the above emotional values was 5, the standard deviation was 2.58, and the fluctuation coefficient was set to 1. Then, the normal emotional range A = [2.42, 7.58].
[0060] Specifically, in the scenario of driving on a highway, the emotional values at 10 different times were obtained as 4, 6, 5, 7, 6, 4, 5, 8, 3, and 7. The average of the above emotional values was 5.5, the standard deviation was 1.58, and the fluctuation coefficient was set to 2. Then the normal emotional range A = [2.34, 8.66].
[0061] It should be understood that the above calculation process for the normal emotional range is only illustrative. In fact, the sample size used to calculate the normal emotional range can be much larger than 10, such as 10,000 emotional value samples.
[0062] (5) Obtain real-time data
[0063] Real-time data primarily includes various data influencing driver emotion scores and data assisting in determining the scenario type. The data influencing driver emotion scores are data across different modalities, as described above, such as the driver's speaking volume, heart rate, facial expressions, and steering wheel pressure. Data assisting in determining the scenario type can include GPS map data, road condition data acquired by cameras, and altitude data.
[0064] It should be understood that real-time data refers to data collected at the moment the driver's emotional state is being assessed. For example, if the current time is 1:30 PM on December 1, 2024, then real-time data refers to data collected at 1:30 PM on December 1, 2024. The substantive content of real-time data and historical data is the same; the only difference is the chronological order in which they are collected relative to the moment the driver's emotional state is being assessed. For instance, data collected at 1:30 PM on December 1, 2024, is considered real-time data when assessing the driver's emotional state at that time; however, it becomes historical data when assessing the driver's emotional state on December 2, 2024.
[0065] (6) Identify real-time scenes
[0066] The process involves determining the scenario type corresponding to the acquired real-time data. The scenario type refers to the external environment in which the vehicle is driving (e.g., urban congestion, highways, and plateau roads). The process of identifying real-time scenarios from real-time data is similar to the process of identifying historical scenarios from historical data described above, and therefore will not be repeated here.
[0067] (7) Determine the weights
[0068] Based on the initial weight allocation described above and the dynamic adjustment of weight distribution based on the occurrence of specific events, the modal weight distribution corresponding to the real-time scene can be determined. For example, in the standard scene weight template library, the initial weight of the voice modality in the highway scene is set to 10%. After detecting that the driver is arguing with other people on the highway, the weight of the voice modality can be increased to 15%.
[0069] Preferably, the modal weight distribution for each scenario is updated based on an autoregressive method, which is described by the following equation:
[0070]
[0071] in, It is the weight corresponding to the real-time data. The weights are based on historical data. It is the regression coefficient.
[0072] For example, + ,in In a highway scenario, the weight of the voice modality is used to determine the driver's emotional state on the 5th. It is the weight of the 4-day speech modality. It is the weight of the 3-day speech modality. It is the weight of the 2-day speech modality. It can be 0.4. It can be 0.3. The value can be 0.3. It should be understood that the above formula is merely illustrative, and the historical data used in the autoregressive method can be historical data from any other time period, such as within 3 days, 5 days, or 10 days. Furthermore, the regression coefficient is also illustrative; generally, the closer the regression coefficient is to the current judgment date, the larger it is, because more recent historical data has a greater impact on the current judgment.
[0073] Preferably, the regression coefficients can be dynamically adjusted based on driver feedback. For example, for + In the autoregressive model, if the driver approves of the emotional judgments on days 4 and 3 but disapproves of the emotional judgment on day 2, and this disapproval is identified through direct voice feedback from the driver, or by turning off or modifying relevant functions (e.g., turning off sound alerts, adjusting the direction or temperature of the air conditioning vents, turning off or adjusting the audio equipment), then the regression coefficient for day 2 can be lowered accordingly, and the regression coefficients for days 3 and 4 can be increased accordingly, thus modifying the autoregressive model to... + Then, the weight distribution for this scenario is redefined based on the modified autoregressive model. User feedback is used to refine the weight distribution, making it more reasonable. It should be understood that although the weights assigned to the speech modalities in the highway scenario have been described in detail above, this invention is not limited to this; the weight allocation for any modality in any scenario can be determined based on the above method.
[0074] (8) Calculate real-time sentiment value
[0075] For each specific type of scenario, the emotional value at the judgment moment is obtained by weighted averaging of the emotional scores under different modalities.
[0076] (9) Determine the emotional state
[0077] If the real-time emotion value falls within the normal emotion range for that scenario, the driver's emotional state is determined to be normal. If the real-time emotion value does not fall within the normal emotion range for that scenario, the driver's emotional state is determined to be abnormal. Figure 3 As shown, the horizontal axis represents the time within a day, and the vertical axis represents the sentiment value. The solid line represents the average historical sentiment value, and the dashed line represents the normal fluctuation range of historical sentiment values (i.e., k times the standard deviation). Real-time sentiment values within the envelope of the dashed line are considered normal, while real-time sentiment values outside the envelope of the dashed line are considered abnormal.
[0078] (10) User feedback
[0079] If the user believes the emotional state assessment is accurate, the relevant data (especially the weight distribution of different modalities in a specific scenario) is saved for later retrieval. If the user believes the emotional state assessment is inaccurate, the weight distribution of different modalities is adjusted based on the modified autoregressive model method described above.
[0080] Figure 2 The diagram illustrates a driver emotion intervention method according to an embodiment of the present invention. The driver's emotional state is obtained based on the aforementioned emotion recognition method. If the emotional state is abnormal, intervention is initiated; if the emotional state is normal, no intervention is performed. Emotional intervention measures may include, for example, first gently reminding the driver to pay attention to driving safety while suggesting deep breathing and relaxation. If the driver continues to exhibit high levels of anxiety / anger, such as repeated angry verbal expressions and a continuously rising heart rate, the system will further intervene by playing soothing audio in the vehicle's background music, adjusting the air conditioning system, or using a voice assistant for interventional dialogue to help the driver alleviate negative emotions.
[0081] Figure 4 A schematic diagram of a driver emotion recognition system 10 according to an embodiment of the present invention is shown. The emotion recognition system 10 includes a data collection module 11, a scene recognition module 12, a calculation module 13, and an emotion judgment module 14. These modules can be individual modules or integrated together. The data collection module 11 is configured to collect historical data and real-time data (e.g., the driver's speaking volume, the driver's heart rate, the driver's facial expressions, and the force with which the driver manipulates the steering wheel). The scene recognition module 12 determines the corresponding scene type based on the data collected by the data collection module 11 (e.g., congested city, highway, and plateau road). The calculation module 13 assigns weights to different modalities based on specific scene types, and calculates historical and real-time emotion values based on the assigned weights and the emotion scores corresponding to different modalities. The emotion judgment module 14 compares the real-time emotion value with the normal fluctuation range of the historical emotion value. If the real-time emotion value is outside the normal fluctuation range of the historical emotion value, it indicates that the driver's emotional state is abnormal, and vice versa.
[0082] Therefore, according to this application, by setting corresponding weight values for multiple modalities for each different scenario, the driver's emotional state in each scenario can be determined more accurately, thereby improving the accuracy of emotion recognition. Furthermore, by updating the aforementioned weight values using a regression method, the accuracy of judging and recognizing the emotional state in each scenario is further improved. According to another aspect of the invention, a non-volatile computer-readable storage medium is also provided, on which computer-readable instructions are stored. When a vehicle controller reads these instructions, the controller can instruct corresponding modules or devices to execute the method described above.
[0083] The program portion of a technology can be considered a "product" or "artifact" existing in the form of executable code and / or related data, and is involved in or implemented through a computer-readable medium. Tangible, permanent storage media can include memory or storage used by any computer, processor, or similar device or related module. For example, various semiconductor memories, tape drives, disk drives, or any similar device capable of providing storage functionality for software.
[0084] All software, or parts thereof, may sometimes communicate via networks, such as the Internet or other communication networks. Such communication can load software from one computer device or processor to another. For example, loading software from a server or host computer in a traffic light recognition system to a hardware platform of a computer environment, or another computer environment implementing the system, or a system with similar functionality related to providing the information needed for traffic light recognition. Therefore, another medium capable of transmitting software elements can also be used as a physical connection between local devices, such as light waves, radio waves, electromagnetic waves, etc., propagated through cables, fiber optic cables, or air. Physical media used for carrier waves, such as cables, wireless connections, or fiber optic cables, can also be considered as media carrying software. In this context, unless limited to tangible "storage" media, the term "readable medium" for a computer or machine refers to the medium involved in the execution of any instructions by the processor.
[0085] Furthermore, those skilled in the art will understand that aspects of this application can be described and illustrated through several patentable types or situations, including any new and useful combination of processes, machines, products, or substances, or any new and useful improvements thereof. Accordingly, aspects of this application can be implemented entirely by hardware, entirely by software (including firmware, resident software, microcode, etc.), or by a combination of hardware and software. All of the above hardware or software may be referred to as a “data block,” “module,” “engine,” “unit,” “component,” or “system.” Furthermore, aspects of this application may manifest as a computer product located on one or more computer-readable media, the product including computer-readable program code.
[0086] Unless otherwise defined, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in a common dictionary shall be interpreted as having a meaning consistent with their meaning in the context of the relevant art, and not as being interpreted in an idealized or highly formalized sense, unless expressly defined herein.
[0087] The above description is illustrative of the invention and should not be construed as limiting it. Although several exemplary embodiments of the invention have been described, those skilled in the art will readily understand that many modifications can be made to the exemplary embodiments without departing from the novel concept and advantages of the invention. Therefore, all such modifications are intended to be included within the scope of the invention as defined in the claims. It should be understood that the above description is illustrative of the invention and should not be construed as limiting it to the specific embodiments disclosed, and modifications to the disclosed embodiments and other embodiments are intended to be included within the scope of the appended claims. The invention is defined by the claims and their equivalents.
Claims
1. A driver emotion recognition method, comprising: Collect historical data; Based on the historical data, the corresponding scenario is determined, and for each of the multiple scenarios, the weight of each of the multiple modalities representing the driver's emotional state is determined. The normal emotional range for each of the multiple scenarios is determined based on the historical data and the weights. Obtain real-time data; The corresponding scenario is determined based on the real-time data, and the real-time sentiment value is determined based on the weight of the corresponding modality of the scenario and the real-time data. If the real-time emotion value is within the normal emotion range for the scene, the driver's emotional state is determined to be normal; if the real-time emotion value is not within the normal emotion range for the scene, the driver's emotional state is determined to be abnormal.
2. The driver emotion recognition method according to claim 1, wherein the historical data includes emotion scores corresponding to multiple modalities in multiple scenarios, the emotion value in each scenario is equal to the weighted average of the emotion scores of each modality in that scenario, and the normal emotion range in that scenario is determined based on multiple emotion values in that scenario.
3. The driver emotion recognition method according to claim 2, wherein the normal emotion range , in, μ is the average of the multiple emotion values, σ is the standard deviation of the multiple emotion values, and k is the volatility coefficient.
4. The driver emotion recognition method according to claim 3, wherein the fluctuation coefficient depends on the scene type.
5. The driver emotion recognition method according to claim 1, wherein the weights of each modality in each scenario are updated based on an autoregressive method, the autoregressive method being described by the following formula: , in, It is the weight corresponding to the real-time data. The weights are based on historical data. It is the regression coefficient.
6. The driver emotion recognition method according to claim 5, wherein the weight of each modality in each scenario is dynamically adjusted based on the occurrence of a specific event, and / or the regression coefficient is dynamically adjusted based on the driver's feedback.
7. The driver emotion recognition method according to claim 4, wherein the multiple modalities include one or more of a speech modality, a driving behavior modality, a physiological modality, and a visual modality, wherein the speech modality further includes multiple speech submodalities, wherein the speech submodalities include one or more of speech rate, volume, tone, and semantic features.
8. The driver emotion recognition method according to claim 7, wherein the multiple scenarios include urban congestion, highways and plateau sections.
9. The driver emotion recognition method according to claim 8, wherein... The weight of the physiological modality in the high-altitude road section scenario is higher than the weight of the physiological modality in the urban congestion and highway scenarios; and / or The weight of the visual modality in the high-altitude road section scenario is lower than the weight of the visual modality in the urban congestion and highway scenarios; and / or The weight of the voice modality in the high-altitude road section scenario is lower than the weight of the voice modality in the urban congestion and highway scenarios.
10. The driver emotion recognition method according to claim 9, wherein the fluctuation coefficient in the urban congestion scenario is higher than the fluctuation coefficient in the highway scenario.