Road rage identification and mitigation method and system based on personalized intelligent voice interaction
By combining road condition information with the driver's personalized factors into a voice interaction system, road rage can be identified and alleviated, solving the problem that existing technologies fail to effectively consider road conditions and personalized factors, and achieving a more targeted mitigation effect.
Patent Information
- Application Number
- CN202411862295.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2026-07-10
- Estimated Expiration
- 2044-12-17
Smart Images

Figure CN119673214B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent driving and driving safety technology, and in particular relates to a method and system for identifying and alleviating road rage based on personalized intelligent voice interaction. Background Technology
[0002] Cars have become a common household item, serving as a basic means of transportation. However, the increasing pressures of modern life and work can easily lead to pent-up emotions that erupt at any moment, particularly among drivers who are prone to irritability and irrational behavior. "Road rage" is one such example. Road rage refers to aggressive or angry behavior by motorists, primarily manifested in the use of insulting language or gestures, deliberately driving in a way that threatens safety, and engaging in actions that endanger the life or health of others.
[0003] Road traffic safety is a crucial component of public safety. With the increase in vehicles and the growing complexity of road traffic conditions, the impact of "road rage" on road traffic safety is becoming increasingly prominent. Therefore, detecting and intervening in road rage among drivers is of paramount importance.
[0004] The existing technology includes a Chinese invention patent application with publication number CN 117549902 A entitled "A Driver Road Rage Detection Device and Detection Method Thereof". This patent acquires information on the driver's grip pressure and steering wheel tapping pressure through a pressure detection module, the angular velocity of the steering wheel rotation through a direction detection module, the driver's heart rate through a heart rate detection module, and the driver's voice decibel level through a sound pressure level detection module. A processing module processes and analyzes the information acquired from these modules to determine whether the driver is experiencing road rage. An alarm module receives the information from the processing module and issues corresponding warning signals, including voice signals, seat vibration, and forced stopping.
[0005] While the aforementioned technical solutions can detect road rage to some extent, they do not consider potential factors in road condition information that could trigger road rage during the detection process. Furthermore, the warning signals do not take into account the driver's personality. Simple voice signals, seat vibrations, and forced stops may exacerbate road rage symptoms and further negatively impact the driver. Summary of the Invention
[0006] To overcome the shortcomings of the prior art, this invention provides a method and system for identifying and alleviating road rage based on personalized intelligent voice interaction. It combines road condition information to identify whether there are driving scenarios that lead to road rage, and takes into account the driver's personality traits to match personalized voice and voice materials, and generate personalized voice interaction content to alleviate the driver's negative emotions of road rage, achieving a better and more targeted alleviation effect.
[0007] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:
[0008] The first aspect of this invention provides a method for identifying and alleviating road rage based on personalized intelligent voice interaction.
[0009] A road rage identification and mitigation method based on personalized intelligent voice interaction includes the following steps:
[0010] Obtain road condition information and detect whether there are driving scenarios in the road condition information that could lead to road rage;
[0011] When road condition information contains driving scenarios that could lead to road rage, the system acquires the driver's real-time voice and language characteristics to identify whether the driver is experiencing negative emotions related to road rage.
[0012] Acquire voice feature information of drivers during daily driving to determine the drivers' personality traits;
[0013] Personalized voices are matched based on the driver's personality traits, and voice materials are matched based on road condition information. When the driver is under the influence of negative emotions, personalized voice interaction content is generated based on the personalized voices and voice materials to alleviate the driver's negative emotions of road rage.
[0014] A second aspect of the present invention provides a road rage recognition and mitigation system based on personalized intelligent voice interaction.
[0015] A road rage recognition and mitigation system based on personalized intelligent voice interaction includes:
[0016] The road rage factor detection module is configured to: acquire road condition information and detect whether there are driving scenarios in the road condition information that could lead to road rage.
[0017] The road rage emotion recognition module is configured to: when there are driving scenarios in the road condition information that could lead to road rage, acquire the driver's real-time voice and language feature information, and identify whether the driver is in a negative emotional state of road rage.
[0018] The personality trait determination module is configured to: acquire the driver's voice feature information during daily driving to determine the driver's personality traits;
[0019] The personalized voice interaction module is configured to: match a personalized voice based on the driver's personality traits, match voice materials based on road condition information, and generate personalized voice interaction content based on the personalized voice and voice materials when the driver is under the influence of negative emotions, so as to alleviate the driver's negative emotions of road rage.
[0020] A third aspect of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps of the road rage identification and mitigation method based on personalized intelligent voice interaction as described in the first aspect of the present invention.
[0021] The fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the road rage recognition and mitigation method based on personalized intelligent voice interaction as described in the first aspect of the present invention.
[0022] The above one or more technical solutions have the following beneficial effects:
[0023] This invention provides a method and system for identifying and alleviating road rage based on personalized intelligent voice interaction. When detecting a driver's road rage, the system combines road condition information to identify whether there are driving scenarios that could lead to road rage. If such scenarios exist, the system further identifies whether the driver is experiencing negative emotions related to road rage based on the driver's real-time voice and language features. Through these technical means, road rage can be identified more effectively.
[0024] When soothing and alleviating road rage in drivers, the system takes into account the drivers' individual factors and personality traits, matches personalized voices and voice materials, and generates personalized voice interaction content to alleviate drivers' negative emotions about road rage, achieving better and more targeted relief results.
[0025] The method of this invention combines scenarios that are likely to cause negative driving, real-time fusion perception of the driver, and long-term fusion perception to determine the risk of road rage in drivers and effectively alleviate it, thereby reducing the probability of road rage and improving the driving experience while ensuring safe driving.
[0026] This invention directly acquires surrounding road condition information through an autonomous driving sensor ADAS system, which is more accurate and real-time compared to other systems that use GPS to acquire road information; the voice system can provide different feedback to alleviate road rage emotions based on different personality traits.
[0027] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0028] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0029] Figure 1 This is an overall flowchart of Embodiment 1 of the present invention.
[0030] Figure 2 This is a flowchart illustrating the process of obtaining surrounding road condition information via ADAS in Embodiment 1 of the present invention.
[0031] Figure 3 A flowchart for determining the driver's current emotions is provided for Embodiment 1 of the invention.
[0032] Figure 4 The flowchart for determining the personality characteristics of a driver is shown in Embodiment 1 of the present invention. Detailed Implementation
[0033] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0034] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.
[0035] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0036] Example 1
[0037] To address driving safety issues caused by road rage, and to ensure the comprehensiveness and accuracy of road rage identification while effectively calming drivers' negative emotions with personalized voice and language based on their personality traits, thus preventing further adverse stimulation, this invention proposes a method for mitigating road rage based on intelligent voice interaction and ADAS information. This method combines scenarios that easily trigger negative driving, real-time fusion perception of the driver, and long-term fusion perception to assess the risk of road rage and effectively alleviate it, thereby reducing the probability of road rage occurring.
[0038] Overall, this embodiment first acquires surrounding road condition information through ADAS; then determines the driver's current emotion based on intelligent fusion perception; determines the driver's personality traits based on the results of long-term intelligent fusion perception; and determines a personalized driver emotion mitigation plan based on the current road conditions and driver personality traits acquired by the ADAS system. This effectively reduces the probability of road rage.
[0039] The technical solution of this embodiment will be further explained below with reference to the accompanying drawings.
[0040] like Figure 1 As shown, the method for identifying and alleviating road rage based on personalized intelligent voice interaction includes the following steps:
[0041] Obtain road condition information and detect whether there are driving scenarios in the road condition information that could lead to road rage;
[0042] When road condition information contains driving scenarios that could lead to road rage, the system acquires the driver's real-time voice and language characteristics to identify whether the driver is experiencing negative emotions related to road rage.
[0043] Acquire voice feature information of drivers during daily driving to determine the drivers' personality traits;
[0044] Personalized voices are matched based on the driver's personality traits, and voice materials are matched based on road condition information. When the driver is under the influence of negative emotions, personalized voice interaction content is generated based on the personalized voices and voice materials to alleviate the driver's negative emotions of road rage.
[0045] In some embodiments, obtaining road condition information and detecting whether there are driving scenarios in the road condition information that could lead to road rage includes:
[0046] Road condition information is obtained through the ADAS system, including traffic congestion, the speed of surrounding vehicles, whether other vehicles have cut in front of them, the distance to surrounding vehicles, and whether surrounding vehicles frequently change lanes.
[0047] The system detects whether there are driving scenarios in the road condition information that could lead to road rage. These scenarios include being cut off by another vehicle, the vehicle in front braking suddenly, and traffic jams lasting longer than a set time.
[0048] In some embodiments, acquiring road condition information through an ADAS system includes collecting information about the vehicle's surroundings using multiple sensors and predicting potential hazards. Specifically, collecting information about the vehicle's surroundings using multiple sensors includes:
[0049] The camera captures visual information about the vehicle's surroundings, including road signs, lane lines, pedestrians, and obstacles.
[0050] Millimeter-wave radar can be used to detect the distance, speed, and azimuth of target objects.
[0051] Depth information of the surrounding environment is obtained through lidar;
[0052] Close-range detection is performed using ultrasonic radar;
[0053] Predicting potential hazardous situations, specifically including:
[0054] Based on information collected by sensors, obstacles around the vehicle are detected and tracked;
[0055] Through image processing technology and algorithms, the system can identify the vehicle's current driving status and lane position, and issue a warning when the vehicle deviates from the lane;
[0056] It detects obstacles in front of the vehicle and predicts the risk of collision in order to issue an early warning;
[0057] The system detects vehicles ahead using sensors and automatically adjusts its speed to maintain a safe distance from the vehicle in front.
[0058] In some embodiments, acquiring the driver's real-time voice and language feature information to identify whether the driver is experiencing road rage includes:
[0059] Acquire the driver's voice characteristics, including volume, speech rate, pitch, intensity, formants, and tone; acquire the driver's language characteristics, including profanity and uncivilized language; acquire other driver-related information, including interjections, sighs, sounds of patting objects, and sounds played from a mobile phone;
[0060] While considering the driver's personality, age, and voice characteristics, the system comprehensively considers the aforementioned voice characteristics, language characteristics, and other driver-related information, and matches them with the corresponding characteristics of typical anger / irritability / tension to determine whether the driver is already in a negative emotional state of road rage.
[0061] It is understandable that, in order to facilitate matching, it is necessary to pre-set an information database of features corresponding to anger / irritability / tension. The similarity between the three types of information—voice features, language features, and other driver-related information—and the corresponding features in the database is comprehensively considered. When the similarity is greater than a set threshold, it can be determined that the current driver is in a negative emotional state of road rage.
[0062] In some embodiments, acquiring voice feature information of the driver during daily driving to determine the driver's personality characteristics specifically includes:
[0063] The voice feature information of the driver during daily driving is obtained and input into the personality trait model to obtain the personality trait prediction results. The personality trait prediction results include five personality traits: openness, conscientiousness, extraversion, agreeableness, and neuroticism.
[0064] The personality trait model is used to determine personality trait information based on speech feature information.
[0065] In some embodiments, personalized voices are matched based on the driver's personality traits, and voice materials are matched based on road condition information, specifically including:
[0066] Considering the driver's personality traits, select a voice from the existing voice library that matches the driver's personality and emotional preferences as the personalized voice. The personalized voice needs to be polite and friendly in tone and consistent with the vehicle's product image.
[0067] Based on road condition information, voice materials that match the current road condition information are selected from the existing voice material library as matching voice materials. The matching voice materials contain both positive and negative content, and the subject is environmental description. Positive content is praise-based, while negative content is prompt-based, warning-based, and command-based.
[0068] In some embodiments, personalized voice interaction content is generated based on personalized voice and speech materials to alleviate road rage in drivers, specifically including:
[0069] Considering that the driver belongs to a specific category among the five major personality traits—openness, conscientiousness, extraversion, agreeableness, and neuroticism—personalized voice interaction content is generated based on personalized voice and speech materials, including:
[0070] For drivers with open-minded personality traits, use more direct language to remind them;
[0071] For drivers with a strong sense of responsibility, from a driving safety perspective, use encouraging language to remind them.
[0072] For drivers with extroverted personality traits, use more direct or explicit language to remind them than for drivers with open personality traits.
[0073] More specifically, such as Figures 1-4 As shown, this embodiment may include the following steps:
[0074] Step S1:
[0075] The main purpose of using ADAS systems to obtain road and traffic information is to detect potential factors that could lead to road rage in a given scenario. These potential factors include traffic congestion, the speed of surrounding vehicles, whether other vehicles have cut in front of you, the distance to surrounding vehicles, and whether surrounding vehicles are frequently changing lanes.
[0076] The ADAS system, short for Advanced Driver Assistance System, is an in-vehicle system that uses sensors such as cameras and radar to acquire environmental data. This data is then processed and analyzed by hardware and software systems to identify and assess traffic and driving conditions, providing information and assistance to the driver. When a potential hazard is anticipated, it will warn the driver and may even take proactive intervention measures such as emergency braking.
[0077] In this embodiment, the road condition information mainly refers to driving scenarios that are prone to causing negative emotions in people. It is generally identified through sensors used in autonomous driving. The road condition information includes, but is not limited to:
[0078] 1) His car cut in front of me;
[0079] 2) The car in front braked suddenly, almost causing a rear-end collision;
[0080] 3) Long-term traffic jams.
[0081] Furthermore, the specific steps for using ADAS to acquire road and traffic condition information mainly include three stages: environmental perception, computational analysis, and control execution, as detailed below. Figure 2 As shown.
[0082] (1) Environmental perception
[0083] The main function of environmental perception is to perceive the environment around the vehicle. At this stage, information about the vehicle's surroundings is collected through different types of vehicle sensors, such as traffic congestion, the speed of surrounding vehicles, whether other vehicles have cut in front of the vehicle, the distance to surrounding vehicles, and whether surrounding vehicles frequently change lanes.
[0084] These sensors include cameras, millimeter-wave radar, lidar, and ultrasonic radar.
[0085] Further:
[0086] The camera captures visual information around the vehicle, including road signs, lane lines, pedestrians, obstacles, etc., and converts this information into digital signals for the system to process.
[0087] Millimeter-wave radar detects the distance, velocity, and azimuth of target objects by transmitting and receiving electromagnetic waves at the millimeter level.
[0088] LiDAR obtains depth information about the surrounding environment by emitting a laser beam and measuring the reflection time;
[0089] Ultrasonic radar is mainly used for short-range detection, such as measuring the distance between a vehicle and an obstacle in a parking assistance system.
[0090] During the environmental sensing process, the data collected by the sensors will be transmitted to the electronic control unit or domain controller for processing.
[0091] (2) Operational Analysis
[0092] The primary function of computational analysis is to analyze the information obtained from environmental perception in order to predict potential hazards. At this stage, the electronic control unit (ECU) or domain controller processes and analyzes the data collected by sensors, employing various algorithms and models to identify obstacles around the vehicle, determine the vehicle's driving status, and assess road conditions. Based on this information, the system makes corresponding decisions and issues commands.
[0093] The computational analysis phase includes:
[0094] Obstacle detection and tracking: Based on information collected by sensors, the system can detect and track obstacles around the vehicle.
[0095] Lane departure detection and warning: Through image processing technology and algorithms, the system can identify the vehicle's current driving status and lane position, and issue a warning when the vehicle deviates from the lane.
[0096] Forward collision detection and warning: Using sensors such as radar or cameras, the system can detect obstacles in front of the vehicle and predict the risk of collision, so as to issue an early warning.
[0097] Adaptive cruise control: The system can detect vehicles in front of the vehicle using sensors and automatically adjust the vehicle speed to maintain a safe distance from the vehicle in front.
[0098] (3) Control execution
[0099] The primary function of control execution is to control the vehicle through actuators based on the results of computational analysis. Actuators can control the vehicle's acceleration, braking, steering, and other functions to ensure safe driving. At this stage, the electronic control unit or domain controller sends corresponding control commands to the actuators to achieve various control functions.
[0100] The control execution phase includes:
[0101] Automatic parking: The system can automatically control the vehicle's steering, acceleration, braking, etc., to achieve automatic parking.
[0102] Automatic braking: When sensors detect obstacles ahead, the system can automatically brake the vehicle to avoid the risk of a collision.
[0103] Automatic acceleration: The system can automatically control the acceleration and deceleration of the vehicle based on the driving status of the vehicle in front, in order to maintain a safe distance from the vehicle in front.
[0104] Automatic steering: By detecting the vehicle's driving status and road conditions through sensors, the system can automatically control the vehicle's steering to ensure that the vehicle is driving on the correct road.
[0105] Step S2:
[0106] In this embodiment, the driver's current emotion is determined based on intelligent fusion perception, which mainly includes identifying the current emotional state based on voice and language feature information.
[0107] The intelligent fusion perception includes, but is not limited to:
[0108] 1) Driver's facial expressions and head movement information extracted by the camera;
[0109] 2) Driver's speech and voice information extracted by the microphone;
[0110] 3) Driver information extracted by motion sensors such as steering, braking, and acceleration sensors;
[0111] 4) Information about the driver's daily use of the car, such as the force and angle at which the car doors are opened and closed.
[0112] Furthermore, such as Figure 3 As shown, if it is determined that there are many potential factors causing road rage due to the current road conditions, the driver's current state of anger, tension, or irritability can be identified mainly by the voice and language characteristics information provided by the microphone.
[0113] The voice feature information may include volume, speech rate, pitch, intensity, formants, tone, etc. While taking into account voice characteristics such as personality and age, the driver's feature information is matched with typical voice features of anger / irritability / tension to determine whether the driver is currently experiencing road rage.
[0114] The language feature information primarily includes the detection of profanity. The voice and language detection may also include other driver-related information, such as interjections, sighs, patting objects, and the presence of any sounds from a mobile phone. This type of information can be used to determine whether the driver is under the influence of negative emotions.
[0115] Step S3:
[0116] like Figure 4As shown, the driver's personality characteristics are determined based on the results of long-term intelligent fusion perception. This mainly includes identifying the driver's personality traits based on voice feature information and a personality trait model, wherein the personality trait model is used to determine personality trait information based on voice feature information.
[0117] The car is equipped with a microphone, and the computer equipment uses the microphone to acquire multiple frames of voice data from the driver while driving.
[0118] To obtain the correlation between a driver's voice characteristics and personality while driving, and considering that a car may contain other occupants besides the driver in the driver's seat, such as passengers in the front passenger seat or rear seats, it is necessary to determine whether the acquired voice data is the driver's voice data. If the voice data is not the driver's voice data, voice data is reacquired until the driver's voice data is obtained. In addition, if the voice data includes the driver's voice data but also the voice data of other occupants, it is still necessary to reacquire voice data until only the driver's voice data is obtained, thereby improving the quality of the acquired voice data.
[0119] The computer device stores the driver's voiceprint characteristics in advance; correspondingly, the steps for the computer device to determine whether the acquired voice data is the driver's voice data can be as follows:
[0120] Extract voiceprint features from the speech data and determine whether the extracted voiceprint features match the driver's voiceprint features. If the extracted voiceprint features match the driver's voiceprint features, determine that the speech data is the driver's speech data. If the extracted voiceprint features do not match the driver's voiceprint features, determine that the speech data is not the driver's speech data.
[0121] Based on multi-frame speech data, multiple sets of voice features of the driver during the driving process are determined. Each set of voice features includes volume, pitch, intensity, tone, formants, speech rate, effective speech length, and number of pauses.
[0122] Because it is necessary to find the correlation between the driver's voice characteristics while driving and personality, it is necessary to solve for the characteristic values of the driver's voice while driving, such as pitch, intensity, formants, effective speech length, speech rate and number of pauses, and then solve for the maximum, minimum, variance, mean and median of these characteristics, so as to obtain the driver's voice characteristics.
[0123] Sound characteristics include pitch; this is achieved through the following steps:
[0124] For each frame of speech data, determine the fundamental frequency value of the speech data;
[0125] Pitch is determined by the vibration frequency of the sound source; in this step, the computer device performs autocorrelation calculation on each frame of speech data, and estimates the fundamental frequency value of each frame of speech data based on the results of the autocorrelation calculation.
[0126] Based on the fundamental frequency value of the speech data, the driver's pitch, determined from the speech data, is calculated using the following formula:
[0127]
[0128] Where h represents the driver's pitch determined based on the voice data, and f represents the fundamental frequency value of the voice data.
[0129] Sound characteristics include sound intensity; sound intensity is the energy carried by sound propagation and is determined by the amplitude of the sound. This embodiment calculates it using frequency domain Fourier transform and energy spectrum. A Fourier transform is performed on each frame of data, converting the time-domain signal into a frequency-domain signal. The amplitude of the frequency-domain signal is squared to calculate the energy spectrum of each frame. The energy spectrum represents the energy of each frequency component. The sound intensity of each frame is then calculated based on the energy spectrum.
[0130] Accordingly, this step can be:
[0131] For each frame of speech data, a Fourier transform is performed to obtain the frequency domain signal, which includes multiple frequencies. For each frequency in the frequency domain signal, the square of the frequency amplitude is determined to obtain the frequency's energy spectrum. Based on the energy spectrum of each frequency, the sound pressure level of the speech data is determined. Based on the sound pressure level of the speech data, the driver's voice intensity determined from the speech data is determined using the following formula:
[0132]
[0133] Where Lp represents the driver's voice intensity determined based on voice data, p represents the sound pressure level, and p0 represents the reference sound pressure level.
[0134] Sound features include formants; formants are broad peaks or local maxima in the frequency spectrum. Formants are regions of relatively concentrated energy in the sound spectrum, reflecting the physical characteristics of the vocal tract. Extracting formants hinges on estimating the maxima within the natural speech spectrum envelope. Formant information is contained within the frequency envelope; therefore, methods for extracting formant parameters include cepstral method, LPC (Linear Predictive Coding), and HHT (Yellow Transform). In this embodiment, the cepstral method is used as an example to determine formants; correspondingly, this can be achieved through the following steps:
[0135] Pre-emphasis processing is performed on multi-frame speech data;
[0136] The pre-emphasis processed multi-frame speech data is windowed and framed to obtain multi-frame second speech data.
[0137] The frame length of each frame of second speech data in the multi-frame second speech data is no greater than the frame length of the frame-segmented processing; and the frame length of the frame-segmented processing can be set and changed as needed.
[0138] The frequency domain signal of the multiple frames of second speech data is obtained by performing a Fourier transform on the data using the following formula:
[0139]
[0140] Where Xi(n) represents the frequency domain signal of the second speech data in the i-th frame, xi(n) represents the second speech data in the i-th frame, N represents the frame length of the frame-by-frame processing, and n represents the number of frames of the second speech data;
[0141] The frequency domain signal of multiple frames of second speech data is subjected to inverse Fourier transform to obtain the inverted sequence of multiple frames of second speech data;
[0142] Based on the DAP sequence of multiple frames of second speech data, the envelope of the frequency domain signal of the multiple frames of second speech data is determined by the following formula:
[0143]
[0144] Where Hi(k) represents the envelope of Xi(n), hi(n) represents the product of the window function and the cepstral sequence, N represents the frame length of the frame-segmented processing, and n represents the number of frames of the second speech data;
[0145] For each second speech data, the maximum value is determined on the envelope of the frequency domain signal of the second speech data to obtain the driver's formant determined based on the second speech data.
[0146] Sound features include effective speech length; correspondingly, the effective speech length can be obtained by determining the difference between the duration of multiple frames of speech data and the duration of silence.
[0147] Voice features include speech rate, which is the vocabulary capacity contained per unit time; correspondingly, speech rate can be obtained by determining the ratio of the total duration to the number of words in multiple frames of speech data.
[0148] The sound features include the number of pauses; correspondingly, it can be: removing the time period at the start of silence and the time period at the end of silence from multiple frames of speech data, and determining the number of times the silence duration in the remaining speech segment is longer than a preset duration as the number of pauses.
[0149] After completing the above steps, based on multiple sets of sound features, determine the maximum value, minimum value, variance, mean, and median of each sound feature included in the multiple sets of sound features to obtain the sound features of the driver during the driving process.
[0150] A set of sound features includes pitch, intensity, formants, effective speech length, speech rate, and number of pauses. In this step, the computer device determines the maximum, minimum, variance, mean, and median of pitch, intensity, effective speech length, speech rate, and number of pauses for multiple sets of sound features, thus obtaining the sound features of the driver during driving.
[0151] Next, sample data is generated based on voice features, with labels for the sample data representing pre-determined driver personality traits. Computer equipment generates multiple sample data sets based on the voice features of multiple drivers, and then trains the model. Based on the sample data, a personality trait model is trained using machine learning algorithms, and this model is used to identify the user's personality traits.
[0152] Computer equipment constructs multiple machine learning models based on multiple machine learning algorithms, determines the quality parameter values of multiple machine learning models, selects the target machine learning model with the best quality from multiple machine learning models based on the quality parameter values of multiple machine learning models, and then trains the target machine learning model based on sample data. The model obtained after training is the personality trait model.
[0153] Among them, multiple machine learning algorithms can be selected from random forest, Naive Bayes, support vector machine, decision tree, and neural network algorithms.
[0154] The steps for a computer device to determine the quality parameter values of multiple machine learning models can be as follows:
[0155] For each machine learning model, the computer device determines the accuracy, precision, recall, and overall score of the machine learning model, and then performs a weighted sum of the accuracy, precision, recall, and overall score to obtain the quality parameter value of the machine learning model.
[0156] In this embodiment of the application, since the machine learning model with the best quality is selected for training, the accuracy of the personality trait model trained based on the machine learning model can be improved.
[0157] Furthermore, the personality trait model needs to obtain a training dataset, and the model is trained based on the training dataset to obtain the personality trait model.
[0158] Furthermore, the driver's voice feature information is input into the personality trait model to obtain the corresponding personality trait prediction results. The personality trait prediction results should include five personality traits: openness, conscientiousness, extraversion, agreeableness, and neuroticism. These can be used as a factor in selecting the voice and voice materials for subsequent voice interaction.
[0159] Furthermore, suitable sounds are selected from the existing sound library, and corresponding materials are selected from the existing voice materials according to different road scenarios.
[0160] Furthermore, the in-vehicle terminal equipment should have a pre-set default voice language setting, which can be set to a female voice with a polite and friendly tone that matches the vehicle's product image.
[0161] Furthermore, based on the driver's personality traits and road condition information obtained by the ADAS system, sound and voice materials that are appropriate for the current situation are selected.
[0162] The audio material contains both positive and negative content, and both should be primarily descriptive of the environment; they can be used in combination.
[0163] Furthermore, the definition of positive content is praise-based, for example:
[0164] 1. "Not cutting in line is the right thing to do; you are an excellent driver who obeys traffic rules."
[0165] 2. "You have helped protect the safety of the road driving environment";
[0166] 3. "Your driving makes us feel safe and secure";
[0167] 4. "You are the guardian and practitioner of road safety."
[0168] Furthermore, negative content can be defined as suggestive, warning, or commanding, and can include the driver's full name (XXX) as a reminder. For example:
[0169] 1. "XXX, please slow down before turning ahead";
[0170] 2. "XXX, the roads are slippery due to heavy rain, so please maintain a safe distance between vehicles."
[0171] 3. "XXX, you should stay focused and concentrate on driving. Anger only increases the risk of a car accident."
[0172] 4. “XXX, the drivers and passengers of surrounding vehicles can see your bad driving behavior and will condemn it,” etc.
[0173] Furthermore, in terms of voice, the personality traits of drivers are taken into consideration, and voices that match their character and emotional preferences are selected, while also allowing drivers to switch voice interaction voices at any time.
[0174] Step S4:
[0175] Based on the current road conditions and driver personality traits obtained by the ADAS system, a driver emotion mitigation plan is determined. This mainly includes alleviating road rage and negative emotions while driving through voice interaction. The voice interaction aims to achieve the following objectives:
[0176] 1) Alleviate road rage and prevent drivers from engaging in dangerous driving behaviors;
[0177] 2) Shift the driver's attention to driving rather than venting their emotions. Over time, this will help the driver learn to view bad road conditions more objectively and rationally, and better adjust their mindset.
[0178] 3) Suppress profanity to prevent it from amplifying the negative impact of profanity on driving anger.
[0179] In one embodiment, based on the situation of being cut off by another vehicle, under the same negative driving emotional atmosphere:
[0180] For drivers with open-minded personality traits, the voice interaction will take into account their thinking style and will remind them in more direct language, such as: "You are currently in a state of anger or impatience. Please calm down and maintain a good mood while driving."
[0181] For drivers with a responsible personality, the voice interaction will take driving safety into consideration and will use encouraging language, such as: "You are a very responsible driver. Please pay attention to the safety of yourself and others and give way to vehicles that cut in front of you in time."
[0182] For drivers with extroverted personality traits, voice interaction will use more direct or explicit language to alleviate the driver's current emotions, such as: "Beautiful lady / handsome gentleman, please don't get angry because the car behind you cut in. Anger will ruin your pleasant mood."
[0183] For drivers with agreeable personality traits, the voice interaction will use friendly, cooperative, and supportive language to alleviate the driver's current emotions. For example: "Thank you for always being considerate of others and patient. Your friendly behavior not only makes the road more harmonious but also ensures driving safety. Please continue to influence those around you with your good behavior."
[0184] For drivers with neurotic personality traits, the voice interaction will use soothing language to help them manage their current emotions. For example: "You can try taking deep breaths and slowing down to give yourself some space to adjust your mood. We understand that your current emotions are normal, but focusing on safe driving will make the journey smoother and more comfortable."
[0185] When detecting road rage in drivers, this invention combines road condition information to identify whether there are driving scenarios that could lead to road rage. If such scenarios exist, the invention further identifies whether the driver is experiencing negative emotions related to road rage based on the driver's real-time voice and language features. Through these technical means, road rage can be identified in a more targeted manner.
[0186] When soothing and alleviating road rage in drivers, the system takes into account the drivers' individual factors and personality traits, matches personalized voices and voice materials, and generates personalized voice interaction content to alleviate drivers' negative emotions about road rage, achieving better and more targeted relief results.
[0187] This invention combines scenarios that easily trigger negative driving behavior with real-time and long-term fusion perception of the driver to assess the risk of road rage and effectively mitigate it, reducing the probability of road rage and improving the driving experience while ensuring driver safety. The ADAS (Advanced Driver Assistance Systems) system directly acquires surrounding road condition information through autonomous driving sensors, which is more accurate and real-time compared to systems that use GPS to obtain road information; the voice system can provide different feedback to alleviate road rage emotions based on different personality traits.
[0188] Example 2
[0189] This embodiment discloses a road rage recognition and mitigation system based on personalized intelligent voice interaction.
[0190] A road rage recognition and mitigation system based on personalized intelligent voice interaction includes:
[0191] The road rage factor detection module is configured to: acquire road condition information and detect whether there are driving scenarios in the road condition information that could lead to road rage.
[0192] The road rage emotion recognition module is configured to: when there are driving scenarios in the road condition information that could lead to road rage, acquire the driver's real-time voice and language feature information, and identify whether the driver is in a negative emotional state of road rage.
[0193] The personality trait determination module is configured to: acquire the driver's voice feature information during daily driving to determine the driver's personality traits;
[0194] The personalized voice interaction module is configured to: match a personalized voice based on the driver's personality traits, match voice materials based on road condition information, and generate personalized voice interaction content based on the personalized voice and voice materials when the driver is under the influence of negative emotions, so as to alleviate the driver's negative emotions of road rage.
[0195] Example 3
[0196] The purpose of this embodiment is to provide a computer-readable storage medium.
[0197] A computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps in the road rage recognition and mitigation method based on personalized intelligent voice interaction as described in Embodiment 1 of this disclosure.
[0198] Example 4
[0199] The purpose of this embodiment is to provide an electronic device.
[0200] An electronic device includes a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps in the road rage recognition and mitigation method based on personalized intelligent voice interaction as described in Embodiment 1 of this disclosure.
[0201] The steps and methods involved in the apparatuses of Embodiments 2, 3, and 4 above correspond to those in Embodiment 1. For specific implementation details, please refer to the relevant description section of Embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood as including any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and enabling the processor to perform any of the methods in this invention.
[0202] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computer devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computer device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. The present invention is not limited to any particular combination of hardware and software.
[0203] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A method for identifying and alleviating road rage based on personalized intelligent voice interaction, characterized in that, Includes the following steps: Obtain road condition information and detect whether there are driving scenarios in the road condition information that could lead to road rage; When road condition information contains driving scenarios that could lead to road rage, the system acquires the driver's real-time voice and language characteristics to identify whether the driver is experiencing negative emotions related to road rage. Acquire voice feature information of drivers during daily driving to determine the drivers' personality traits; The voice feature information of the driver during daily driving is obtained and input into the personality trait model to obtain the personality trait prediction results. The personality trait prediction results include five personality traits: openness, conscientiousness, extraversion, agreeableness, and neuroticism. Personalized voices are matched based on the driver's personality traits, and voice materials are matched based on road condition information. When the driver is under the influence of negative emotions, the system considers the specific category of the five personality traits of openness, conscientiousness, extraversion, agreeableness, and neuroticism. Personalized voice interaction content is generated based on the personalized voice and voice materials to alleviate the driver's negative emotions of road rage.
2. The method for identifying and alleviating road rage based on personalized intelligent voice interaction as described in claim 1, characterized in that, Acquire road condition information and detect whether there are driving scenarios in the road condition information that could lead to road rage, specifically including: Road condition information is obtained through the ADAS system, including traffic congestion, the speed of surrounding vehicles, whether other vehicles have cut in front of them, the distance to surrounding vehicles, and whether surrounding vehicles frequently change lanes. The system detects whether there are driving scenarios in the road condition information that could lead to road rage. These scenarios include being cut off by another vehicle, the vehicle in front braking suddenly, and traffic jams lasting longer than a set time.
3. The method for identifying and alleviating road rage based on personalized intelligent voice interaction as described in claim 2, characterized in that, ADAS systems acquire road condition information, including collecting information about the vehicle's surroundings through multiple sensors and predicting potential hazards. Specifically, the collection of information about the vehicle's surroundings through multiple sensors includes: The camera captures visual information about the vehicle's surroundings, including road signs, lane lines, pedestrians, and obstacles. Millimeter-wave radar can be used to detect the distance, speed, and azimuth of target objects. Depth information of the surrounding environment is obtained through lidar; Close-range detection is performed using ultrasonic radar; or, Predicting potential hazardous situations, specifically including: Based on information collected by sensors, obstacles around the vehicle are detected and tracked; It identifies the vehicle's current driving status and lane position, and issues a warning when the vehicle deviates from its lane; Detect obstacles in front of the vehicle and predict collision risks, issuing early warnings; It detects vehicles in front of it and automatically adjusts its speed to maintain a safe distance from the vehicle in front.
4. The method for identifying and alleviating road rage based on personalized intelligent voice interaction as described in claim 1, characterized in that, Acquire real-time voice and language feature information of the driver to identify whether the driver is experiencing road rage, specifically including: Acquire the driver's voice characteristics, including volume, speech rate, pitch, intensity, formants, and tone; acquire the driver's language characteristics, including profanity and uncivilized language; acquire other driver-related information, including interjections, sighs, sounds of patting objects, and sounds played from a mobile phone; While considering the driver's personality, age, and voice characteristics, the system comprehensively considers the aforementioned voice characteristics, language characteristics, and other driver-related information, and matches them with the corresponding characteristics of typical anger / irritability / tension to determine whether the driver is already in a negative emotional state of road rage.
5. The method for identifying and alleviating road rage based on personalized intelligent voice interaction as described in claim 1, characterized in that, The driver's voice feature information during daily driving is obtained to determine the driver's personality characteristics, wherein the personality trait model is used to determine personality trait information based on the voice feature information.
6. The method for identifying and alleviating road rage based on personalized intelligent voice interaction as described in claim 1, characterized in that, Personalized voices are matched based on the driver's personality traits, and voice materials are matched based on road condition information, specifically including: Considering the driver's personality traits, select a voice from the existing voice library that matches the driver's personality and emotional preferences as the personalized voice. The personalized voice needs to be polite and friendly in tone and consistent with the vehicle's product image. Based on road condition information, voice materials that match the current road condition information are selected from the existing voice material library as matching voice materials. The matching voice materials contain both positive and negative content, and the subject is environmental description. Positive content is praise-based, while negative content is prompt-based, warning-based, and command-based.
7. The method for identifying and alleviating road rage based on personalized intelligent voice interaction as described in claim 5, characterized in that, Personalized voice interaction content is generated based on personalized voice and voice materials to alleviate road rage in drivers, specifically including: Personalized voice interaction content is generated based on personalized voice and speech materials, including: For drivers with open-minded personality traits, use more direct language to remind them; For drivers with a strong sense of responsibility, from a driving safety perspective, use encouraging language to remind them. For drivers with extroverted personality traits, use more direct or explicit language to remind them than for drivers with open personality traits.
8. A road rage recognition and mitigation system based on personalized intelligent voice interaction, characterized in that, The method for identifying and alleviating road rage based on personalized intelligent voice interaction as described in any one of claims 1-7 includes: The road rage factor detection module is configured to: acquire road condition information and detect whether there are driving scenarios in the road condition information that could lead to road rage. The road rage emotion recognition module is configured to: when there are driving scenarios in the road condition information that could lead to road rage, acquire the driver's real-time voice and language feature information, and identify whether the driver is in a negative emotional state of road rage. The personality trait determination module is configured to: acquire the driver's voice feature information during daily driving to determine the driver's personality traits; The personalized voice interaction module is configured to: match a personalized voice based on the driver's personality traits, match voice materials based on road condition information, and generate personalized voice interaction content based on the personalized voice and voice materials when the driver is under the influence of negative emotions, so as to alleviate the driver's negative emotions of road rage.
9. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the road rage recognition and mitigation method based on personalized intelligent voice interaction as described in any one of claims 1-7.
10. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the road rage recognition and mitigation method based on personalized intelligent voice interaction as described in any one of claims 1-7.
Citation Information
Patent Citations
Driver road rage detection device and detection method thereof
CN117549902A
Driver safety awareness assessment method based on virtual driving and EEG detection
CN108268887A
Safe driving method and system based on emotion recognition
CN112370037A