Safety warning method, device and storage medium based on in-vehicle voice environment monitoring

By using microphone arrays and seat sensors in the vehicle to monitor the sound and seat status in the vehicle, combined with the sentiment analysis network model, the comprehensive risk score of the voice environment in the vehicle is calculated and warning signals are issued, which solves the problem that the existing vehicle voice system cannot effectively monitor and early warning of the voice environment in the vehicle that may affect driving safety, real-time monitoring and safety warning of the voice environment in the vehicle is achieved.

CN118447878BActive Publication Date: 2025-05-16CHONGQING SELIS PHOENIX INTELLIGENT INNOVATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410499523.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-24
Publication Date
2025-05-16
Estimated Expiration
2044-04-24

AI Technical Summary

Technical Problem

The existing vehicle voice system has shortcomings in comprehensive monitoring and early warning of the interior environment, and cannot effectively identify and deal with in-car voice environments that may affect driving safety, such as long-term communication between drivers and passengers or emotional excitement.

Method used

The sound signals in the vehicle are collected through a microphone array installed inside the vehicle, combined with the seat sensor to obtain seat occupancy information, calculate differential sound signals and perform time and frequency domain processing, use optimization algorithms to determine the location of the sound source, extract voice characteristics and predict emotional state through the emotional analysis network model, calculate comprehensive risk scores and issue corresponding levels of early warning signals.

Benefits of technology

Real-time and accurate identification and monitoring of the voice environment in the car can effectively warn of potential safety hazards and improve driving safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118447878B_ABST
    Figure CN118447878B_ABST
Patent Text Reader

Abstract

The present application provides a safety warning method, device and storage medium based on in-vehicle voice environment monitoring. The method includes: collecting in-vehicle sound signals and obtaining seat occupancy information; calculating differential sound signals and performing time domain and frequency domain processing; using an optimization algorithm to calculate and determine the sound source position, and determining the target object based on the sound source position; preprocessing the sound signal of the target object and performing speech segmentation, and extracting the target speech features from the speech segments; using an emotion analysis network model to extract local features in the preprocessed sound signal, and learning the time series features in the local features, and predicting the emotional state category based on the time series features; calculating the comprehensive risk score of the target object based on the target speech features and emotional state category, and issuing a warning signal of the corresponding level based on the comprehensive risk score. The present application accurately identifies and monitors the in-vehicle voice environment, and effectively warns of safety hazards in the vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of new energy vehicle technology, and in particular to a safety warning method, device and storage medium based on in-vehicle voice environment monitoring. Background Art

[0002] In the development of vehicle safety systems, the monitoring and safety warning system of the in-vehicle voice environment is an important component. Although in-vehicle voice recognition technology has been widely used in the market, mainly used to process the driver's voice commands and perform related in-vehicle control functions, these systems have obvious shortcomings in the comprehensive monitoring and warning of the in-vehicle environment.

[0003] Current in-vehicle voice systems focus primarily on the recognition and response of voice commands, often ignoring the complexity of the in-vehicle environment and the impact that environment may have on driving safety. For example, these systems cannot effectively recognize and process long conversations between drivers and passengers, emotional agitation, or other behavior patterns that may affect driving safety, such as quarrels and other emotionally charged communication behaviors. These situations may cause the driver to lose focus and even abnormal driving behavior, greatly increasing the risk of accidents.

[0004] In addition, existing technologies for in-car sound processing usually only focus on single indicators such as noise level or speech clarity, and lack a comprehensive assessment of the in-car speech environment. For example, although some systems try to identify nervous or excited emotional states by extracting emotional features of speech, these recognitions usually rely on simple sound features and are difficult to accurately capture complex human emotional changes, so their effectiveness in practical applications is limited. Summary of the invention

[0005] In view of this, the embodiments of the present application provide a safety warning method, device and storage medium based on in-vehicle voice environment monitoring to solve the problem that the prior art lacks a comprehensive assessment of the in-vehicle voice environment, resulting in the inability to effectively warn of possible safety hazards.

[0006] In a first aspect of an embodiment of the present application, a safety warning method based on in-vehicle voice environment monitoring is provided, comprising: collecting sound signals in the vehicle using a microphone array installed inside the vehicle, and obtaining seat occupancy information using a seat sensor; calculating a differential sound signal based on the collected sound signal, and performing time domain and frequency domain processing on the differential sound signal; calculating and determining the sound source position using a predetermined optimization algorithm based on the processed differential sound signal and occupancy information, and determining the target object based on the sound source position, wherein the occupancy information is used to adjust the weight of the differential sound signal; preprocessing the sound signal of the target object, segmenting the preprocessed sound signal into speech segments, and extracting target speech features from the speech segments; extracting local features from the preprocessed sound signal using an emotion analysis network model, learning time series features from the local features, and predicting the emotional state category based on the time series features; calculating a comprehensive risk score for the target object based on the target speech features and emotional state category corresponding to the target object, and issuing a warning signal of a corresponding level based on the comprehensive risk score.

[0007] According to a second aspect of an embodiment of the present application, a safety warning device based on in-vehicle voice environment monitoring is provided, comprising: an acquisition module, configured to acquire sound signals in the vehicle using a microphone array installed inside the vehicle, and to obtain seat occupancy information using a seat sensor; a processing module, configured to calculate a differential sound signal based on the acquired sound signal, and to perform time domain and frequency domain processing on the differential sound signal; a determination module, configured to calculate and determine a sound source position using a predetermined optimization algorithm based on the processed differential sound signal and occupancy information, and to determine a target object based on the sound source position, wherein the occupancy information is used to adjust the weight of the differential sound signal; an extraction module, configured to preprocess the sound signal of the target object, and to perform speech segmentation on the preprocessed sound signal, and to extract target speech features from the speech segments; a prediction module, configured to extract local features from the preprocessed sound signal using an emotion analysis network model, and to learn time series features from the local features, and to predict the emotional state category based on the time series features; and a warning module, configured to calculate a comprehensive risk score of the target object based on the target speech features and emotional state category corresponding to the target object, and to issue a warning signal of a corresponding level based on the comprehensive risk score.

[0008] According to a third aspect of an embodiment of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the steps of the above method are implemented when the processor executes the computer program.

[0009] According to a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, which stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0010] At least one of the above technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects:

[0011] The sound signals in the car are collected by using a microphone array installed inside the vehicle, and the seat occupancy information is obtained by using a seat sensor; the differential sound signal is calculated based on the collected sound signal, and the differential sound signal is processed in the time domain and frequency domain; according to the processed differential sound signal and occupancy information, a predetermined optimization algorithm is used to calculate and determine the sound source position, and the target object is determined according to the sound source position, wherein the occupancy information is used to adjust the weight of the differential sound signal; the sound signal of the target object is preprocessed, and the preprocessed sound signal is segmented into speech paragraphs, and the target speech features are extracted from the speech paragraphs; the local features in the preprocessed sound signal are extracted using an emotion analysis network model, and the time series features in the local features are learned, and the emotional state category is predicted according to the time series features; the comprehensive risk score of the target object is calculated according to the target speech features and emotional state category corresponding to the target object, and a warning signal of the corresponding level is issued according to the comprehensive risk score. This application can accurately identify and monitor the voice environment in the car in real time, and effectively warn of safety hazards in the car. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0013] Figure 1 It is a structural diagram of the in-vehicle voice environment monitoring and safety warning system provided by an embodiment of the present application;

[0014] Figure 2 It is a flowchart of a safety warning method based on in-vehicle voice environment monitoring provided by an embodiment of the present application;

[0015] Figure 3 It is a structural schematic diagram of a safety warning device based on in-vehicle voice environment monitoring provided by an embodiment of the present application;

[0016] Figure 4 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0017] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.

[0018] In the existing technology, the comprehensive evaluation and safety warning system of the in-vehicle voice environment is not very perfect. The in-vehicle voice recognition systems currently on the market mainly focus on the recognition and execution of voice commands, while the monitoring and safety warning functions of the in-vehicle voice environment are relatively weak. When detecting the driver's voice commands, the existing in-vehicle voice system does not take into account the real-time situation of the in-vehicle environment and cannot effectively warn of possible safety hazards.

[0019] Current in-vehicle systems focus primarily on the clarity and accuracy of voice commands, often ignoring other factors that affect the in-vehicle voice environment. For example, the system may not be able to effectively identify the noise level in the driving environment, the conversations between passengers, and their emotional fluctuations, which are important factors that affect driving safety. The following are behaviors that are likely to affect driving safety during driving:

[0020] (1) The driver and passengers communicate for a long time, causing the main driver to lose concentration;

[0021] (2) The driver and the passengers may become emotionally agitated during the communication process (such as quarreling), which may lead to abnormal behavior of the driver or the passengers (such as grabbing the steering wheel);

[0022] (3) The driver and passengers communicate with each other for a long time, causing the driver to be distracted while driving.

[0023] Existing systems can usually only detect a single indicator, such as noise level or clarity of the driver's voice, and lack a comprehensive assessment of the overall condition of the in-vehicle environment. In addition, the ability to warn of possible safety hazards is limited, and it is unable to provide accurate and timely safety tips. Although the existing technology proposes to extract speech features for emotion recognition, the recognition results of such a single feature are often not accurate and comprehensive enough. Therefore, a more intelligent and comprehensive solution is needed to evaluate the in-vehicle voice environment and provide safety warnings to make up for the shortcomings of the existing technology.

[0024] Therefore, these limitations and shortcomings in the existing technology reveal an important technical need, that is, to develop an intelligent system that can comprehensively evaluate the in-vehicle voice environment and provide effective safety warnings based on it. Such a system should not only be able to process and recognize standard voice commands, but also be able to comprehensively consider environmental noise, passenger communication content and their emotional state, so as to more comprehensively ensure driving safety.

[0025] In view of the problems existing in the prior art, the present application provides a method and system for comprehensive assessment of vehicle voice environment and safety warning based on sound source positioning and sound analysis, aiming to solve the limitations of in-vehicle voice environment assessment and safety warning functions in the prior art. The system is mainly composed of a sound source positioning module, a sound processing and analysis module, a risk scoring module and a warning issuance module. By using the sound source positioning technology (SSL) to accurately capture the location information of each sound source in the car, and combining the sound analysis technology to analyze the volume, volume change rate, communication duration and emotional state of each sound source, the system can comprehensively evaluate the in-vehicle voice environment and issue a corresponding level of safety warning based on the evaluation results. This solution realizes real-time monitoring and analysis of the in-vehicle voice environment and improves driving safety by comprehensively utilizing the quality improvement of the in-vehicle sound signal, sound source positioning, emotional analysis and risk assessment.

[0026] Before describing the technical solution of the embodiment of the present application in detail, the system architecture involved in the technical solution of the present application in the actual scenario is first explained in combination with the drawings and specific embodiments. Figure 1 Schematic diagram of the structure of the in-vehicle voice environment monitoring and safety warning system provided by the embodiment of the present application. Figure 1 As shown, the in-vehicle voice environment monitoring and safety warning system may specifically include the following components:

[0027] Microphone array: An array of multiple high-quality microphones used to capture sound signals inside the car. These microphones can receive sounds from different directions and provide the necessary data for sound source localization.

[0028] Sound processing unit: A specialized hardware module used for preprocessing of sound signals, such as denoising, sound amplification, filtering, etc., to improve the accuracy of sound analysis.

[0029] Cockpit chip (CPU): The central processing unit is used to process sound signals and perform computationally intensive tasks such as sound source localization algorithms, sound analysis, and emotion analysis. The performance of the CPU directly affects the response speed and accuracy of the system.

[0030] Deep learning accelerators: Hardware accelerators optimized for deep neural network models, such as GPUs (graphics processing units) or TPUs (tensor processing units). These accelerators can significantly increase the processing speed of sentiment analysis, enabling the system to respond in real time.

[0031] All of the above modules are connected to the CPU through circuits, and the CPU issues warnings to the driver through the vehicle's central control display and audio system.

[0032] The contents of the technical solution of the present application are described in detail below with reference to the accompanying drawings and specific embodiments.

[0033] Figure 2 It is a flow chart of a safety warning method based on in-vehicle voice environment monitoring provided in an embodiment of the present application. Figure 2 The safety warning method based on in-vehicle voice environment monitoring can be executed by the cockpit chip. Figure 2 As shown, the safety warning method based on in-vehicle voice environment monitoring may specifically include:

[0034] S201, collecting sound signals inside the vehicle using a microphone array installed inside the vehicle, and obtaining seat occupancy information using a seat sensor;

[0035] S202, calculating a differential sound signal based on the collected sound signal, and performing time domain and frequency domain processing on the differential sound signal;

[0036] S203, calculating and determining the sound source position using a predetermined optimization algorithm according to the processed differential sound signal and the occupancy information, and determining the target object according to the sound source position, wherein the occupancy information is used to adjust the weight of the differential sound signal;

[0037] S204, preprocessing the sound signal of the target object, segmenting the preprocessed sound signal into speech segments, and extracting target speech features from the speech segments;

[0038] S205, extracting local features from the preprocessed sound signal using an emotion analysis network model, learning time series features from the local features, and predicting the emotion state category based on the time series features;

[0039] S206, calculating a comprehensive risk score of the target object according to the target voice features and emotional state category corresponding to the target object, and issuing a warning signal of a corresponding level according to the comprehensive risk score.

[0040] In some embodiments, a microphone array installed inside a vehicle is used to collect sound signals inside the vehicle, including:

[0041] At least one microphone is installed at each cockpit position inside the vehicle, and the microphones are formed into a microphone array. The microphone array is used to collect sound signals inside the vehicle in real time, wherein each microphone is used to receive sound at a corresponding cockpit position from a different direction.

[0042] Specifically, the embodiment of the present application installs at least one high-sensitivity microphone in each cabin position of the vehicle, including the driver's seat, the co-driver's seat and the back seat. Each microphone is configured to collect sound from a specific direction to cover every corner inside the vehicle. These microphones together form a microphone array that can collect various sound signals in the car in real time, including but not limited to conversations, music, external noise and any abnormal sounds.

[0043] In one example, the microphones may be arranged as follows:

[0044] Driver's seat: Two microphones are installed near the dashboard and steering wheel of the driver's seat, mainly used to capture the driver's instructions and communication sounds.

[0045] Passenger seat: A microphone is installed on the top of the passenger seat and in front of the dashboard to ensure that the voice of the passenger can be effectively captured.

[0046] Rear seats: Microphones are installed in the middle of the roof and near the head of each rear seat to collect the communication sounds of the rear passengers.

[0047] Furthermore, each microphone is highly directional and sensitive, accurately capturing sounds coming from its designated direction. In addition, the microphone is equipped with advanced noise suppression, effectively isolating and extracting human voice signals in high background noise environments. The microphone array is connected to the central processing unit (CPU) via a high-speed data line, and the CPU is responsible for processing and analyzing the collected sound data.

[0048] The system monitors and records the sound environment in the car in real time. By analyzing the characteristics of voice data, such as volume, frequency, pitch and duration, the system can identify different speech patterns such as normal communication, tense or intense emotional changes, etc. In this way, when the system detects possible safety risks, such as disputes between the driver and passengers or long-term distraction, it can issue a warning in time to remind the driver or automatically take measures to restore the vehicle to a safe state.

[0049] In some embodiments, the sound source position is calculated and determined using a predetermined optimization algorithm based on the processed differential sound signal and occupancy information, including:

[0050] The processed differential sound signal is weighted by using the occupancy information to obtain a weighted differential sound signal, and an optimization objective function is constructed with the goal of minimizing the distance between the weighted differential sound signal and the sound source position;

[0051] When solving the optimization objective function, based on the sound propagation differences between the sound source position and each pair of microphones, an optimization algorithm is used to find the sound source position with the smallest difference in the sound signal received at each position after propagating in space, so as to determine the final sound source position.

[0052] Specifically, sound source localization can identify the position of the speaker and process the sounds from different areas separately, which is very useful for accurately evaluating the communication behavior of each passenger and its potential impact on driving safety. In order to improve the accuracy of sound source localization in the car, the embodiment of the present application proposes a method for determining the location of the sound source by combining differential positioning and position occupancy. Specifically, the microphone is installed at a position in the car, and each position corresponds to a microphone. The differential signals received by different microphones are calculated and then processed in the time domain or frequency domain, and then weighted according to whether the position in the car is occupied (hereinafter referred to as the weighted signal). The distance from the predetermined sound source position to the microphone is calculated and subtracted from the weighted signal. On this basis, each microphone is traversed to perform the same operation on the sound position, and the weighted weight is still determined by whether it is occupied.

[0053] In one example, the process of calculating and determining the sound source position using the optimization algorithm provided in the embodiment of the present application is as follows:

[0054] 1) Sound signal collection and positioning algorithm:

[0055] N microphones are set at fixed positions in each cabin of the vehicle, and the sound signals collected are represented by x i (t), where i = 1, 2, ..., N, and t represents time.

[0056] 2) Differential sound signal calculation:

[0057] For each pair of microphones i and j, the differential sound signal is calculated as: ij (t) = x i (t)-x j (t).

[0058] 3) Differential sound signal processing:

[0059] The differential sound signal is processed in the time and frequency domains, such as filtering and noise reduction, to enhance signal quality and extract sound source characteristics.

[0060] 4) Seat occupancy information acquisition:

[0061] Sensors are installed on vehicle seats to detect whether the seats are occupied and obtain seat occupancy information. Assume there are N seats, and use vector S = s1, s2, ..., s N Indicates that s i =1 means seat i is occupied, s i =0 means seat i is not occupied.

[0062] 5) Sound source localization algorithm design:

[0063] Assume that the sound source position is Y = (x, y, z), where x, y, and z represent the horizontal and vertical positions of the sound source in the car, respectively. Consider the seat occupancy information and adjust the weight of each microphone. If the seat where microphone i is located is occupied, the weight of its differential sound signal is reduced, otherwise the original weight is maintained. Let the adjusted differential sound signal be w ij d ij (t), where w ij Represents the weight of the differential sound signal. The adjusted differential sound signal and seat occupancy information are used again to minimize the distance between the adjusted differential sound signal and the sound source position as the objective function to solve the sound source position.

[0064] In one example, the optimization objective function can be defined as:

[0065]

[0066] The physical meaning of optimizing the objective function is to minimize the distance between the adjusted differential sound signal and the actual sound source position by adjusting the sound source position Y. The purpose of this is to find an optimal sound source position to minimize the difference between the differential sound signal and the actual sound source position, thereby improving the accuracy of sound source localization. The microphones at each position are traversed and all differential signals are calculated. The weight information is determined by whether the current seat is occupied. By minimizing the objective function, the specific position of the sound source in the car can be determined.

[0067] 6) Output of sound source localization results:

[0068] According to the sound source position Y obtained by the optimization algorithm, the position of the sound source in the car is determined.

[0069] Through the steps of the above embodiments, accurate positioning of the sound source in the car can be achieved, and it is ensured that the adjusted differential sound signal can correctly participate in the process of sound source positioning. The system can accurately locate the position of the sound source in the car, which is crucial for further sound analysis and driving safety assessment. The entire process takes into account the physical propagation characteristics of the sound signal and the usage status of the seats in the car to ensure the accuracy and practicality of sound source positioning. Doing so can better consider whether the seat is occupied, providing a smarter and safer driving experience for the driver and passengers.

[0070] In some embodiments, the optimization objective function is expressed using the following formula:

[0071]

[0072] Where Y represents the location of the sound source, w ij represents the weight of the differential sound signal, w ij d ij (t) represents the weighted differential sound signal, dist(Y,x i ,y i ,z) represents the distance from the sound source position Y to the position of microphone i.

[0073] Specifically, the purpose of subtracting the distance between the differential sound signal and the sound source position is to measure the difference in sound propagation between the sound source position and each pair of microphones. This difference can be understood as the difference in the sound signal received at different positions after propagating through space (delay difference and amplitude difference). By calculating the difference, it is possible to determine which positions are more likely to be the sound source positions. In a physical sense, the subtraction operation can help find the most suitable sound source position in space at a distance from all microphone positions. Because sound propagation has specific physical laws, the time and sound intensity of sound arriving at different positions will be affected by distance. Therefore, by comparing the difference between the sound signals received at different positions and the theoretically expected sound signals, the position of the sound source can be inferred.

[0074] In some embodiments, the sound signal of the target object is preprocessed, and the preprocessed sound signal is segmented into speech segments, and the target speech features are extracted from the speech segments, including:

[0075] The sound signal of the target object is filtered and denoised by using a digital filter, and the sound signal after filtering and denoising is segmented into speech segments by using a predetermined speech segmentation method;

[0076] Calculate the average volume of the speech segment, and determine the volume change rate based on the rapid change of the volume in the continuous speech segment or the preset time window. Determine the communication duration based on the time of continuous speech activity, and use the average volume, volume change rate and communication duration as the target speech features.

[0077] Specifically, the embodiment of the present application designs a complex system in the sound processing and analysis part to improve the quality and accuracy of in-vehicle voice monitoring, and then evaluate the risks that may affect driving safety. The following are the implementation process and steps of sound processing and analysis in the embodiment of the present application:

[0078] 1) Sound signal processing:

[0079] Use digital filters to process sound data, By setting the cut-off frequency f c To eliminate background noise, the voice signal in the car can be clearly captured. The processed sound signal provides a cleaner and easier to analyze data input, which is helpful for subsequent sound segmentation and feature extraction.

[0080] 2) Sound segmentation and feature extraction:

[0081] Use speech segmentation methods (such as energy-based and zero-crossing rate methods) to segment speech segments, and extract basic features from each segment, such as volume average, volume fluctuation, communication duration, etc. The above features are mainly used for emotion prediction, and basic emotion prediction rules are defined. For example, if the volume of a speech segment suddenly rises and is accompanied by a fast speaking speed (which can be judged by the increase in the number of speech segments in a short period of time), it is judged as a possible negative emotion.

[0082] 2.1) Calculate the average RMS (root mean square) value of each speech segment as the volume indicator.

[0083]

[0084] Among them, x i Represents the i-th sample of the audio signal, and N represents the total number of samples. The RMS value can reflect the volume. Calculate the volume score S vol , if RMS>T vol , then S vol =k vol ·(RMS-T vol ), where k vol is a moderating factor used to adjust the sensitivity of the risk score.

[0085] 2.2) Define volume change rate: volume change rate R vol , which can be used to measure the degree of rapid change between two consecutive speech segments or a certain time window, is defined as follows:

[0086]

[0087] Among them, V current and V previous Represents the average RMS volume value of the current and previous time windows respectively, and Δt is the time difference between the two time windows. vol , set the threshold R threshold To classify the risk level of volume change rate, for example:

[0088] Normal change: R vol ≤R threshold

[0089] Quick Change: Rvol >R threshold

[0090] Volume change risk score S vol_change : The risk score based on the volume change rate can be calculated by S vol A similar method determines that if R vol Exceeding the threshold R threshold , it is considered to be a high risk:

[0091] S vol_change =n·(R vol -R threshold )

[0092] Here, n represents a weight factor used to adjust the sensitivity of the risk score.

[0093] 2.3) Temporal risk assessment of continuous speech activities, temporal level classification:

[0094] Short-term exchange: 0 ≤ T <T short

[0095] Medium AC: T short ≤T <T long

[0096] Long time communication: T>T long

[0097] Where T represents the duration of continuous communication, T short and T long Indicates the pre-set time threshold. Risk score S for continuous communication time dur It can be defined as: dur =m·(TT short ), m is the weight factor. This means that continuous voice activity will generate certain risks only when the communication time is greater than the predetermined minimum time (for example, one hour).

[0098] In some embodiments, the emotion analysis network model is used to extract local features from the preprocessed sound signal, and the time series features in the local features are learned, and the emotion state category is predicted according to the time series features, including:

[0099] The preprocessed sound signal is converted into a Mel-spectrogram, and the Mel-spectrogram is input into the convolution layer of the emotion analysis network model. The convolution layer is used to extract local features in the Mel-spectrogram. The local features are used to characterize the changes in specific sound frequencies.

[0100] The local features are input into the long short-term memory network, which is used to learn the time series features in the local features. The time series features are input into the fully connected layer, which is used to predict the probability of each emotional state to obtain the final emotional state category.

[0101] Specifically, to achieve sentiment analysis, the embodiment of the present application adopts a deep learning architecture (i.e., sentiment analysis network model) that combines a convolutional neural network (CNN) and a recurrent neural network (RNN, especially a long short-term memory network LSTM). Among them, CNN is used to extract local features in sound signals, while LSTM is used to capture dynamic features that change over time in sound data.

[0102] Feature extraction: Convert the sound signal into a Mel-spectrogram, which is a method of converting the frequency of the sound signal into a scale that matches the human ear's perception and can effectively capture the characteristics of the sound. The mathematical formula can be expressed as: M = Mel (x (t)), where x (t) is the given sound signal, and Mel (.) represents the function that converts the sound signal into a Mel-spectrogram.

[0103] CNN layers: Use convolutional layers to process the mel-spectrograms and extract high-level sound features. These convolutional layers can identify local patterns in the sound signal, such as changes in specific frequencies, which are very useful for understanding emotional content.

[0104] LSTM layer: The output of the CNN layer is passed to the LSTM layer to learn the time series features in the sound signal. LSTM can handle long-term dependency problems and is particularly effective in capturing dynamic changes in emotions.

[0105] Output layer: Use one or more fully connected layers to transform the output of the LSTM layer into emotion classification results. These layers can use the softmax function to predict the probability of each emotional state.

[0106] In one example, the risk score S of the affective state is defined as emo as follows:

[0107]

[0108] Among them, p i represents the probability of the model predicting the i-th emotional state, w emo Represents the risk weight associated with this emotional state.

[0109] In some embodiments, a comprehensive risk score of the target object is calculated based on the target voice features and emotional state category corresponding to the target object, and a warning signal of a corresponding level is issued based on the comprehensive risk score, including:

[0110] According to the scoring thresholds corresponding to the emotional state categories and each target speech feature, the emotional state categories and each target speech feature are scored respectively;

[0111] The scores are weighted according to the preset emotional state categories and the score weights corresponding to each target voice feature to obtain a comprehensive risk score;

[0112] Determine the warning level corresponding to the comprehensive risk score, and issue a warning signal corresponding to the warning level according to the warning level, wherein the warning signal includes a visual and / or sound signal.

[0113] Specifically, the risk scoring module is used to analyze voice features and emotional states and calculate comprehensive risk scores. Among them, voice features and emotional state analysis: the sound data collected by multiple microphones inside the vehicle is first processed by denoising and feature extraction to extract target voice features such as volume, volume change rate, and communication duration. At the same time, the emotional state of the voice, such as calmness, tension, or anger, is analyzed through the emotion analysis network model (combining convolutional neural networks and recurrent neural networks).

[0114] Further, a comprehensive risk score is calculated: for each target voice feature and emotional state, a separate score is assigned based on a preset scoring threshold. For example, a higher volume or a fast rate of volume change may indicate a higher risk. Each score is weighted by a preset weight based on its importance to safety. The weight setting is based on an analysis of past accident data and behavioral research to ensure that the score reflects the actual safety risk. All weighted scores are summed up to obtain a comprehensive risk score for each sound source.

[0115] Furthermore, the warning issuance module is used to determine the warning level and issue the warning signal. The warning level is determined based on the comprehensive risk score. A higher score will trigger a higher level of warning. The warning level is divided into low, medium and high levels, and each level corresponds to different warning measures.

[0116] Furthermore, the issuance of warning signals: Low-level warnings may include simple dashboard visual prompts, such as warning lights on. Intermediate warnings may include visual and sound prompts to warn the driver of potential risks. Advanced warnings may further include measures that automatically trigger the vehicle safety system, such as automatic deceleration or activation of the emergency braking system, while providing clear instructions to the driver through the central control display and voice system.

[0117] Through the method of the above embodiment, the safety monitoring efficiency of the vehicle interior can be significantly improved, and the safety accidents caused by internal communication can be effectively prevented or mitigated through comprehensive analysis and real-time response to the voice environment in the vehicle. The design of the system allows flexible adjustment of the scoring threshold and weight according to actual conditions to ensure the accuracy and adaptability of the early warning system.

[0118] According to the technical solution provided in the embodiment of the present application, the technical solution of the present application comprehensively applies sound source localization technology, sound analysis technology, and deep neural network to perform sentiment analysis and risk scoring mechanism, which brings the following significant beneficial effects compared with the prior art:

[0119] 1. Improve accuracy: A positioning algorithm that combines the occupancy information of the vehicle is proposed to help improve the accuracy of identifying the location of each sound source in the vehicle. Combined with the high-precision analysis of emotions by deep neural networks, this system can more accurately assess the safety risks of the in-vehicle voice environment and provide accurate safety warnings for drivers.

[0120] 2. Enhanced real-time performance and response speed: The system design allows real-time processing of voice data in the car, quickly identifying potential safety risks and issuing warnings immediately, which significantly improves response speed and real-time performance.

[0121] 3. Improve safety: By comprehensively considering risk scores based on multiple factors such as volume, communication duration, and emotional state, the system can more comprehensively assess potential safety threats and help drivers take timely measures to avoid or reduce the occurrence of safety accidents.

[0122] 4. Save cost and space: The technical solution adopted by the system can be implemented through software upgrades to enhance the existing vehicle system without the need for a large amount of additional hardware investment, thus saving cost and space.

[0123] The following is an embodiment of the device of the present application, which can be used to execute the embodiment of the method of the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the method of the present application.

[0124] Figure 3 Schematic diagram of the structure of the safety warning device based on in-vehicle voice environment monitoring provided by the embodiment of the present application. Figure 3 As shown, the safety warning device based on in-vehicle voice environment monitoring includes:

[0125] The acquisition module 301 is configured to collect the sound signals in the vehicle using a microphone array installed inside the vehicle, and obtain the seat occupancy information using a seat sensor;

[0126] The processing module 302 is configured to calculate a differential sound signal based on the collected sound signal, and perform time domain and frequency domain processing on the differential sound signal;

[0127] A determination module 303 is configured to calculate and determine the sound source position using a predetermined optimization algorithm according to the processed differential sound signal and occupancy information, and determine the target object according to the sound source position, wherein the occupancy information is used to adjust the weight of the differential sound signal;

[0128] The extraction module 304 is configured to pre-process the sound signal of the target object, segment the pre-processed sound signal into speech segments, and extract the target speech features from the speech segments;

[0129] The prediction module 305 is configured to extract local features from the preprocessed sound signal using the emotion analysis network model, learn the time series features from the local features, and predict the emotion state category according to the time series features;

[0130] The warning module 306 is configured to calculate a comprehensive risk score of the target object according to the target voice features and emotional state category corresponding to the target object, and issue a warning signal of a corresponding level according to the comprehensive risk score.

[0131] In some embodiments, Figure 3 The acquisition module 301 installs at least one microphone at each cabin position inside the vehicle, and the microphones form a microphone array, and use the microphone array to collect the sound signals inside the vehicle in real time, wherein each microphone is used to receive the sound of the corresponding cabin position from different directions.

[0132] In some embodiments, Figure 3 The determination module 303 performs weighted processing on the processed differential sound signal using the occupancy information to obtain a weighted differential sound signal, and constructs an optimization objective function with the goal of minimizing the distance between the weighted differential sound signal and the sound source position; when solving the optimization objective function, based on the sound propagation difference between the sound source position and each pair of microphones, an optimization algorithm is used to find the sound source position with the smallest difference in the sound signal received at each position after propagating in the space, so as to determine the final sound source position.

[0133] In some embodiments, the optimization objective function is expressed using the following formula:

[0134]

[0135] Where Y represents the location of the sound source, w ij represents the weight of the differential sound signal, w ij d ij (t) represents the weighted differential sound signal, dist(Y,x i ,y i ,z) represents the distance from the sound source position Y to the position of microphone i.

[0136] In some embodiments, Figure 3The extraction module 304 uses a digital filter to filter and denoise the sound signal of the target object, and uses a predetermined speech segmentation method to segment the filtered and denoised sound signal into speech segments; calculates the average volume of the speech segments, and determines the volume change rate based on the rapid change degree of the volume in the continuous speech segments or the preset time window, determines the communication duration according to the time of the continuous speech activity, and uses the volume average, volume change rate and communication duration as the target speech features.

[0137] In some embodiments, Figure 3 The prediction module 305 converts the preprocessed sound signal into a Mel-spectrogram, and inputs the Mel-spectrogram into the convolution layer of the emotion analysis network model, and uses the convolution layer to extract local features in the Mel-spectrogram, and the local features are used to characterize the changes in specific sound frequencies; the local features are input into the long short-term memory network, and the long short-term memory network is used to learn the time series features in the local features, and the time series features are input into the fully connected layer, and the fully connected layer is used to predict the probability of each emotional state to obtain the final emotional state category.

[0138] In some embodiments, Figure 3 The early warning module 306 scores the emotional state category and each target voice feature according to the scoring thresholds corresponding to the emotional state category and each target voice feature; weights the scores according to the preset scoring weights corresponding to the emotional state category and each target voice feature to obtain a comprehensive risk score; determines the early warning level corresponding to the comprehensive risk score, and issues an early warning signal corresponding to the early warning level according to the early warning level, wherein the early warning signal includes a visual and / or sound signal.

[0139] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0140] Figure 4 Schematic diagram of the structure of the electronic device 4 provided in the embodiment of the present application. Figure 4 As shown, the electronic device 4 of this embodiment includes: a processor 401, a memory 402, and a computer program 403 stored in the memory 402 and executable on the processor 401. When the processor 401 executes the computer program 403, the steps in the above-mentioned method embodiments are implemented. Alternatively, when the processor 401 executes the computer program 403, the functions of the modules / units in the above-mentioned device embodiments are implemented.

[0141] Exemplarily, the computer program 403 may be divided into one or more modules / units, which are stored in the memory 402 and executed by the processor 401 to complete the present application. The one or more modules / units may be a series of computer program instruction segments capable of completing specific functions, which are used to describe the execution process of the computer program 403 in the electronic device 4.

[0142] The electronic device 4 may be a desktop computer, a notebook, a PDA, a cloud server, or other electronic device. The electronic device 4 may include, but is not limited to, a processor 401 and a memory 402. Those skilled in the art will appreciate that Figure 4 It is only an example of the electronic device 4 and does not constitute a limitation of the electronic device 4. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the electronic device may also include input and output devices, network access devices, buses, etc.

[0143] The processor 401 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc.

[0144] The memory 402 may be an internal storage unit of the electronic device 4, for example, a hard disk or memory of the electronic device 4. The memory 402 may also be an external storage device of the electronic device 4, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 4. Further, the memory 402 may also include both an internal storage unit and an external storage device of the electronic device 4. The memory 402 is used to store computer programs and other programs and data required by the electronic device. The memory 402 may also be used to temporarily store data that has been output or is to be output.

[0145] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.

[0146] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0147] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0148] In the embodiments provided in the present application, it should be understood that the disclosed devices / computer equipment and methods can be implemented in other ways. For example, the device / computer equipment embodiments described above are only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation. Multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0149] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0150] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0151] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. The computer program may include computer program code, which may be in source code form, object code form, executable file or some intermediate form. Computer-readable media may include: any entity or device capable of carrying computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc.

[0152] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A safety warning method based on in-vehicle voice environment monitoring, characterized in that: include: The microphone array installed inside the vehicle is used to collect the sound signals inside the vehicle, and the seat occupancy information is obtained using the seat sensor; Calculating a differential sound signal between each pair of microphones based on the collected sound signal, and performing time domain and frequency domain processing on the differential sound signal; Calculating and determining the sound source position using a predetermined optimization algorithm according to the processed differential sound signal and the occupancy information, and determining the target object according to the sound source position, wherein the occupancy information is used to adjust the weight of the differential sound signal; Preprocessing the sound signal of the target object, segmenting the preprocessed sound signal into speech segments, and extracting target speech features from the speech segments; Extracting local features from the preprocessed sound signal using an emotion analysis network model, learning time series features from the local features, and predicting the emotion state category based on the time series features; Calculating a comprehensive risk score of the target object according to the target voice features and emotional state category corresponding to the target object, and issuing a warning signal of a corresponding level according to the comprehensive risk score; The step of calculating and determining the sound source position by using a predetermined optimization algorithm according to the processed differential sound signal and the occupancy information includes: Using the occupancy information, weighted processing is performed on the processed differential sound signal to obtain a weighted differential sound signal, and an optimization objective function is constructed with the goal of minimizing the distance between the weighted differential sound signal and the sound source position; When solving the optimization objective function, based on the sound propagation difference between the sound source position and each pair of microphones, an optimization algorithm is used to find the sound source position with the smallest difference in the sound signal received at each position after propagating in space, so as to determine the final sound source position.

2. The method according to claim 1, characterized in that The method of collecting the sound signal inside the vehicle by using a microphone array installed inside the vehicle includes: At least one microphone is installed at each cockpit position inside the vehicle, and the microphones are formed into a microphone array, and the microphone array is used to collect sound signals inside the vehicle in real time, wherein each microphone is used to receive sound at a corresponding cockpit position from a different direction.

3. The method according to claim 1, characterized in that The optimization objective function is expressed by the following formula: Where Y represents the location of the sound source, w ij represents the weight of the differential sound signal, d ij (t) represents the differential sound signal, w ij d ij (t) represents the weighted differential sound signal, dist(Y,x i ,y i ,z) represents the distance from the sound source position Y to the position of microphone i, x i ,y i ,z represents the coordinates of the location of microphone i.

4. The method according to claim 1, characterized in that: The preprocessing of the sound signal of the target object, segmentation of the preprocessed sound signal into speech segments, and extraction of target speech features from the speech segments include: Filtering and denoising the sound signal of the target object using a digital filter, and performing speech segmentation on the filtered and denoised sound signal using a predetermined speech segmentation method; The volume average value of the speech segment is calculated, and the volume change rate is determined based on the rapid change degree of the volume in the continuous speech segment or the preset time window. The communication duration is determined according to the time of continuous speech activity, and the volume average value, the volume change rate and the communication duration are used as target speech features.

5. The method according to claim 1, characterized in that The method of extracting local features from the preprocessed sound signal using the emotion analysis network model, learning time series features from the local features, and predicting the emotion state category according to the time series features includes: Converting the preprocessed sound signal into a Mel-spectrogram, and inputting the Mel-spectrogram into a convolutional layer of the emotion analysis network model, and using the convolutional layer to extract local features in the Mel-spectrogram, wherein the local features are used to characterize changes in specific sound frequencies; The local features are input into a long short-term memory network, which is used to learn the time series features in the local features. The time series features are input into a fully connected layer, which is used to predict the probability of each emotional state to obtain the final emotional state category.

6. The method according to claim 1, characterized in that The step of calculating a comprehensive risk score of the target object according to the target voice features and emotional state category corresponding to the target object, and issuing a warning signal of a corresponding level according to the comprehensive risk score, includes: Scoring the emotional state category and each of the target speech features respectively according to the scoring thresholds corresponding to the emotional state category and each of the target speech features; The scores are weighted according to the preset emotional state categories and the score weights corresponding to each of the target voice features to obtain the comprehensive risk score; Determine a warning level corresponding to the comprehensive risk score, and issue a warning signal corresponding to the warning level according to the warning level, wherein the warning signal includes a visual and / or sound signal.

7. A safety warning device based on in-vehicle voice environment monitoring, characterized in that: include: A collection module is configured to collect sound signals inside the vehicle using a microphone array installed inside the vehicle, and to obtain seat occupancy information using a seat sensor; A processing module, configured to calculate a differential sound signal between each pair of microphones based on the collected sound signal, and perform time domain and frequency domain processing on the differential sound signal; a determination module configured to calculate and determine the sound source position using a predetermined optimization algorithm according to the processed differential sound signal and the occupancy information, and determine the target object according to the sound source position, wherein the occupancy information is used to adjust the weight of the differential sound signal; An extraction module is configured to preprocess the sound signal of the target object, segment the preprocessed sound signal into speech segments, and extract target speech features from the speech segments; A prediction module is configured to extract local features from the preprocessed sound signal using an emotion analysis network model, learn time series features from the local features, and predict the emotion state category according to the time series features; An early warning module is configured to calculate a comprehensive risk score of the target object according to the target voice features and emotional state category corresponding to the target object, and issue an early warning signal of a corresponding level according to the comprehensive risk score; Among them, the determination module is used to use the occupancy information to perform weighted processing on the processed differential sound signal to obtain a weighted differential sound signal, and to construct an optimization objective function with the goal of minimizing the distance between the weighted differential sound signal and the sound source position; when solving the optimization objective function, based on the sound propagation difference between the sound source position and each pair of microphones, an optimization algorithm is used to find the sound source position with the smallest difference in the sound signal received at each position after propagating in the space, so as to determine the final sound source position.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Small space sound source azimuth detecting device and method thereof

    CN108490384A

  • Sound source direction positioning method and device, voice equipment and voice system

    CN111025233A