Vehicle head-up display method and system for hearing-disabled driver
By obtaining and identifying sound signals from complex driving environments and converting them into text information to display real-time display, the problem that drivers with hearing disabilities cannot perceive the sounds of driving environments is solved, and driving safety and comfort are improved.
Patent Information
- Application Number
- CN202510630232.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-05-16
AI Technical Summary
Drivers with hearing disabilities cannot perceive sound information about the surrounding driving environment, which affects driving safety and comfort.
By obtaining the sound signals of complex driving environments, a sound recognition model for hearing-disabled drivers is constructed, and combined with the vehicle head-up display system, the sound signals are converted into text information to display in real time.
It realizes that drivers with hearing disabilities can better perceive the surrounding driving environment, improve driving safety, and solves the problem of being unable to perceive sound information.
Smart Images

Figure CN120156307A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent driving, and specifically relates to a vehicle head-up display method and system for hearing-impaired drivers. Background Art
[0002] With the progress of society and the continuous improvement of civilization, ensuring the travel rights of disabled people has received extensive attention from society. As a special driving group, hearing-impaired drivers are unable to accurately and effectively hear the sound information of the surrounding driving environment, such as horn sounds and sirens of special vehicles, which even become an important factor affecting safe travel. For drivers with a low degree of hearing disability, they can still wear hearing aids or cochlear implants and other assistive hearing devices. However, severely hearing-impaired drivers cannot hear any sounds at all and are unable to perceive and understand the surrounding driving environment. Especially in case of an emergency, they are likely to miss important sound signals, posing a severe challenge to driving safety and convenience.
[0003] Vehicle head-up display technology can display key traffic information in front of the driver's field of vision, effectively reducing the number of times the driver looks down at the dashboard, achieving effective human-machine interaction, and improving driving safety and comfort. However, traditional vehicle head-up display technology still cannot solve the problem of the lack of hearing-dimensional information for hearing-impaired drivers. With the continuous development of large models such as speech recognition and machine translation, the sound signal data of the driving environment obtained by the sound sensors in the vehicle can be converted into text information and displayed on the in-vehicle display screen, enabling hearing-impaired drivers to comprehensively perceive the surrounding driving environment. In addition, the complex head-up display information also brings a heavy burden to the driver. Therefore, combining speech recognition, machine translation with head-up display technology to develop a head-up display system for hearing-impaired drivers is not only a need for technological innovation but also an important manifestation of social responsibility. It is of great significance for improving driving safety, providing an autonomous and comfortable driving experience for hearing-impaired drivers, and providing important support for realizing a more equal and safe travel environment and promoting the comprehensive development of intelligent transportation systems.
[0004] There are few patents on vehicle head-up display systems and methods for hearing-impaired drivers in the fields of intelligent cockpits and human-machine interaction. Chinese Patent CN 119065139 A discloses a head-up display control system and a display method for a head-up display, which projects key information onto the windshield, reduces the driver's line-of-sight transfer during driving, and improves driving efficiency. Chinese Patent CN 110531551 A discloses a liquid crystal display device and a head-up display system, which solves the problem of heat generation caused by long-term irradiation of the display module and can enhance the image brightness. Chinese Patent CN 206773941 U discloses a vehicle head-up display device for red-green color-blind and color-weak people, which converts the traffic light image information into a signal that can be recognized by red-green color-blind drivers and projects it onto the windshield of the vehicle. The above three patents can transmit driving information to the windshield and reduce the driver's perception load. However, due to the lack of sound signals, the needs of hearing-impaired drivers still cannot be met. Therefore, it is still very crucial to develop a head-up display system and method for hearing-impaired drivers. Summary of the Invention
[0005] The present invention provides a vehicle head-up display method for hearing-impaired drivers, enabling hearing-impaired drivers to better perceive the surrounding driving environment, improving driving safety, and solving the problem that hearing-impaired drivers cannot perceive the sound information of the surrounding driving environment.
[0006] The technical solution of the present invention is described in conjunction with the accompanying drawings as follows: In a first aspect, the present invention provides a vehicle head-up display method for hearing-impaired drivers, including the following steps: S1. Obtain sound signals for complex driving environments; S2. Construct a sound recognition model for hearing-impaired drivers; S3. Construct a vehicle head-up display system for hearing-impaired drivers.
[0007] Further, the specific method of S1 is as follows: S11. Obtain sound signals based on sensors; S12. Obtain driving sound signals of web crawlers; S13. Obtain sound signals based on generative adversarial networks; Further, the specific method of S11 is as follows: S111. Arrange vehicle sound sensors; Install a sound sensor on each side of the vehicle head to collect sound signals from the front and sides; install a sound sensor on each side of the vehicle rear to obtain the vehicle horn sounds from the rear and the side-rear. S112. Define the acquisition method; According to the Nyquist sampling theorem, the sampling rate is set to twice the highest frequency of the signal, and all frequency sound signals have the same sensitivity; multiple channels are adopted, and it is ensured that the sound sensor has the same sensitivity to sound signals in different directions; S113. Digitally process the sound signal; Sample the analog signal at fixed time intervals to discretize the continuous signal; convert the amplitude value of each sampling point into a digital value with finite precision; convert the quantized digital value into a binary code, and the coding method is PCM; S114. Compress and store the data of the sound signal; Adopt the method of lossless compression and local storage.
[0008] Furthermore, the specific method of S12 is as follows: S121. Grab sound data based on keywords; Using web crawler technology, grab driving sound signal data from the Internet based on keywords; first, determine the target website and data source, analyze the HTML structure of the target web page, and locate the position where the target data is located; extract the links, titles, descriptions, and tags of the target data from the search result page; S122. Download sound data based on multi-threading; Use the method of thread pool to implement multi-threading; the sound signal web page parsing module and the download module each maintain a thread pool; according to the remaining capacity of the thread pool, allocate a new thread, and when the web page parsing task or download task is completed, destroy the thread, and at this time the current remaining capacity is incremented by 1; when it is detected that the remaining capacity of the thread pool is 0, wait for the parsing task or download task to complete and release the resources, and then allocate resources to the new thread; S123. Store and classify the sound signal; Store the metadata information of the audio file in a database or file; convert the downloaded audio file into a unified format; classify the audio file, and label each segment according to the timestamp; finally, upload the data to the cloud storage service.
[0009] Furthermore, the specific method of S13 is as follows: S131. Preprocess the sound signal; Adopt a noise reduction algorithm to remove the environmental noise in the S11 and S12 audio files; divide the driving scene sound signal into short-time frames, perform Fourier transform on each frame of the signal to obtain the spectrum; subtract the noise spectrum from the spectrum of the noisy signal, and finally convert the processed spectrum back to the time-domain signal; then perform standardization or normalization processing on the noise-reduced sound data, and scale the waveform data to between [-1, 1]; S132. Build a voice generation model; The voice generation model includes a generator and a discriminator. The generator is built using one-dimensional transposed convolution, and the discriminator is built using one-dimensional convolution. The input of the generator is a random noise vector, and the output is the generated voice waveform data. The generator includes multiple fully connected layers or convolutional layers. First, the random noise data is mapped to an intermediate dimension through the fully connected layer. The deconvolution layer is used to gradually upsample the data to the dimension of the target waveform data. The output layer outputs the generated voice waveform data, with the same dimension as the real data. The input of the discriminator is the real driving sound signal data or the fake data generated by the generator, and the output is a scalar representing the probability that the input data is real data. The binary cross-entropy loss function is used to alternately train the generator and the discriminator, enabling the discriminator to correctly distinguish between real data and fake data, and enabling the fake data generated by the generator to deceive the discriminator into thinking that the fake data is real. Finally, data of the same type as the real driving sound waveform data is generated; S133. Conduct reliability analysis on the voice signal data generated in S132; Post-process the generated voice signal. Use objective metrics such as signal-to-noise ratio and log spectral distance to evaluate the quality of the generated signal. The higher the signal-to-noise ratio, the better the signal quality. The smaller the log spectral distance, the closer the generated signal is to the target signal. Also, evaluate the quality of the generated signal through listening tests to ensure that the timbre, rhythm, and dynamic range of the driving sound signal conform to the auditory characteristics of the real scenario; S134. Integrate the voice signal data for complex driving scenarios; Integrate the reliable driving sound data generated in S11, S12, and S132 L , L =( L 1, L 2,..., L i ,..., L N ) for a total of N pieces of driving sound signal data. Among them, the data length of each voice waveform signal data L i is different.
[0010] Furthermore, the specific method of S2 is as follows: S21. Build a driving scenario voice signal dataset; S22. Build a driving sound recognition model that integrates time-frequency domain features; S23. Conduct performance test and evaluation on the driving sound recognition model.
[0011] Furthermore, the specific method of S21 is as follows: S211. Extract data features; Extract features from the equal-length sound waveform data integrated by S134, including time-domain feature extraction and frequency-domain feature extraction; the time-domain features include mean, variance, and peak value; use the fast Fourier transform to convert the time-domain sound signal into a frequency-domain sound signal, expressed as a two-dimensional array of frequency and time, calculate the power spectrum from the complex matrix of the Fourier transform, and perform matrix multiplication on the power spectrum and the Mel filter bank to obtain the Mel spectrogram, where the horizontal axis is time and the vertical axis is the Mel frequency; S212, label the data; Label the collected driving sound data; classify and organize the driving sound data according to the scenario and sound category, and assign one or more category labels to the sound in each time period; S213, divide the training set, test set, and validation set; Divide the driving sound data into equal-length sound wavelength data according to the sliding window method, and the sliding window is w , and the sliding step is s ; The calculation formula for the number of samples in each sound data is shown in formula (1): (1) In the formula, is the length of the i-th sound data sample; N The calculation formula (2) for the total number of samples of the driving sound signal data is: (2) Finally, divide the samples into the training set, test set, and validation set according to the ratio.
[0012] Furthermore, the specific method of S22 is as follows: S221, time-domain feature representation learning based on the attention mechanism; Extract time series features from the time-domain waveform signal; the sound waveform signal data with timestamps X ={ x 1, x 2,…, x T}, where x t is the waveform signal at time t , and T is the time step; the original sound waveform signal , and the waveform signal after position encoding is expressed by formula (3): (3) Extract features through the encoder. The encoder consists of multiple layers of multi-head self-attention mechanisms and feed-forward neural networks, and outputs the feature representation result H t, ; (4) S222. CNN-based frequency-domain feature representation learning; Perform a short-time Fourier transform on X to obtain a spectrogram , where F is the frequency dimension, is the number of time frames; Obtain the intermediate feature representation result P through multi-layer 2D convolution and pooling operations. The calculation formula (5) is: (5) Finally, obtain the frequency-domain feature representation result H f , ; S223. Fuse features based on the cross-attention mechanism; Through the cross-attention mechanism, perform fusion on the time-domain feature representation result H t obtained by S221, and the frequency-domain feature representation result H f to obtain the time-frequency domain fusion representation result H fused . The calculation formula (6) is: (6) The calculation formula (7) of CrossAttention is: (7) Through the learnable weight matrix W Q , perform a linear transformation on H t to generate the corresponding Q vector. Through the weight matrices W K and W V perform a linear transformation on H f to generate the corresponding K , V vector; S224. Generate traffic signal recognition text; Through the transformer decoder, generate the word probability distribution through linear projection and Softmax; The decoder consists of l layers, including self-attention, cross-attention, and feed-forward networks; For the lThe layer takes the output of the previous layer as input, and the initial input is the embedding of the target sequence. Y = y 1, y 2,…, y (t-1) , plus the positional encoding; For the processing of the self-attention sublayer, the decoding result Zs is obtained through residual connection and layer normalization. The calculation formula (8) of Zs is: (8) where , , , W Qs , W Ks and W Vs are learnable parameters; For the processing of the cross-attention sublayer, the decoding result Zc is obtained. The calculation formula (9) of Zc is: (9) where , , , W Qc , W Kc and W Vc are learnable parameters; Then, through the feed-forward network sublayer, the calculation formula (10) is: (10) After residual connection and layer normalization, the representation result of the l th layer is obtained. The calculation formula (11) is: (11) The word probability distribution is generated through linear projection and Softmax. The calculation formula (12) is: (12) Finally, the text output of the sound signal is realized.
[0013] Furthermore, the specific method of S23 is as follows: S231. Conduct offline test and evaluation on the model; Taking the mean absolute error, root mean square error, and mean absolute percentage error as the evaluation indicators of the driving sound recognition model, evaluate the difference between the predicted value and the true value of the sound signal recognition model; The calculation formulas (13), (14), and (15) of MAE, RMSE, and MAPE are as follows: (13) (14) (15) In the formula, and are the real voice signal text and the predicted voice signal text respectively; S232. Conduct in-loop test and evaluation on the model; Deploy the trained model to the vehicle hardware; ensure that the model can process voice data in real time when running on the hardware and give feedback in a timely manner; and verify the real-time recognition ability of the model for key voices in the real driving environment.
[0014] Furthermore, the specific method of S3 is as follows: S31. Define the vehicle head-up display system architecture; S32. Obtain the vehicle head-up display information; S33. Design a vehicle head-up display strategy for hearing-impaired drivers.
[0015] Furthermore, the specific method of S31 is as follows: S311. Design system indicators and parameters; Define the head-up display system indicators, including the resolution and contrast of the display information, horizontal and vertical field of view angles, brightness, and chromaticity uniformity; S312. Compose the head-up display system; The head-up display system includes a MEMS scanning mirror, a laser, a laser diode driver, and a projection screen; S313. Generate MEMS laser projection imaging; Combining microelectromechanical systems and laser scanning technology, using red, green, and blue primary color lasers as light sources, collimate, focus, and modulate the lasers through optical elements; after the laser beam passes through the MEMS micromirror, it is scanned according to the pixel information of the image data, and finally a complete image is formed on the screen.
[0016] Furthermore, the specific method of S32 is as follows: S321. Convert the voice signal into text information; Deploy the driving voice recognition model that combines time-frequency domain features constructed in S22 on the vehicle hardware. The driving voice signal data obtained by the vehicle-mounted voice sensor is input into the vehicle control unit VCU through the CAN bus, and after processing, the corresponding text information is output; S322. Obtain the vehicle driving state information; The acceleration and angular velocity are obtained through an inertial navigator, and the velocity and position information are obtained through integral processing; the voltage, current and temperature of the battery are monitored through a battery management system to obtain the battery capacity and cruising range; S323. Generate vehicle collision warning information; The original data of the target object obtained by lidar, millimeter wave radar, ultrasonic radar and camera are processed and fused, and combined with the vehicle's own speed and acceleration to calculate the relative distance d and relative speed v rel , and predict the collision time ttc , ttc = d / v rel , and generate vehicle collision warning information; S324. Integrate the vehicle head-up display information; Integrate the voice signal text information, driving state information and vehicle collision warning information output by S321, S322 and S323; the VCU sends the information to the head-up display system controller according to a fixed data format and message format, and the head-up display system controller receives the signal, converts the received signal into a visual graph, and projects it on the screen through an optical projection system.
[0017] Further, the specific method of S33 is as follows: S331. Design the head-up display interface; Project the display information to different positions on the display screen according to different priorities; take the driving sound signal as the first priority and project it to the center of the screen where the driver is looking, the display information of the second priority is displayed below the center of the screen, and the display information of the fourth priority and the third priority are respectively displayed on the left and right sides; S332. Determine the priority of the vehicle head-up display information; Put the voice signal information that hearing-impaired drivers cannot obtain in the first priority to help drivers understand the driving environment; the display information of the second priority includes collision warning information, vehicle battery power shortage prompt information, lane departure information, and fault information; the third priority displays navigation information, including lane information, lane line information, traffic light information and traffic sign information; the fourth priority displays external environment interaction information, including phone call display information, weather information, and vehicle interior temperature information.
[0018] In a second aspect, the present invention also provides a vehicle head-up display system for hearing-impaired drivers, which is used for the vehicle head-up display method for hearing-impaired drivers, and includes: An acquisition module, which is used to acquire voice signals for complex driving environments; The first construction module is used to construct a voice recognition model for hearing-impaired drivers; The second construction module is used to construct a vehicle head-up display system for hearing-impaired drivers.
[0019] The beneficial effects of the present invention are as follows: 1. The vehicle head-up display method and system based on hearing-impaired drivers according to the present invention include three steps: obtaining voice signal in a complex driving environment, constructing a voice recognition model for hearing-impaired drivers, and constructing a vehicle head-up display system for hearing-impaired drivers. For obtaining voice signal in a complex driving environment, driving voice signal data is obtained through sensors and crawler methods, and vehicle driving voice signal data is generated according to the generative adversarial network for data augmentation. 2. The vehicle head-up display method and system for hearing-impaired drivers according to the present invention construct a voice recognition model for hearing-impaired drivers to recognize voice signals into text information. 3. The vehicle head-up display method and system for hearing-impaired drivers according to the present invention design a head-up display interface and display the text information in the interface in real time, realizing friendly interaction between humans and machines, enabling hearing-impaired drivers to better perceive the surrounding driving environment and improving driving safety. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0021] Figure 1 It is a schematic flowchart of the present invention; Figure 2 It is a schematic overall architecture diagram of step S1 in the present invention; Figure 3 It is a schematic overall architecture diagram of step S2 in the present invention; Figure 4 It is a schematic overall architecture diagram of step S3 in the present invention; Figure 5 It is an algorithm example diagram of a driving voice recognition model for fusing time-frequency domain features in S22 of the present invention; Figure 6 It is a schematic structural diagram of the hardware platform built in S22 of the present invention; Figure 7 It is an example diagram of the priority of the vehicle head-up display human-computer interaction interface in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. In addition, it should be noted that for the convenience of description, only the parts related to the present invention rather than all the structures are shown in the drawings.
[0023] Embodiment 1: Refer to Figures 1 - 7 , this embodiment provides a vehicle head-up display method for hearing-impaired drivers, including the following steps: S1. Obtain sound signals for complex driving environments, specifically as follows: S11. Obtain sound signals based on sensors; S111. Arrange vehicle sound sensors; Install a sound sensor on each side of the vehicle head to collect sound signals in the front and side directions, such as the horn sound of oncoming vehicles and the sound of driving tires. Install a sound sensor on each side of the vehicle tail to obtain the horn sounds of vehicles behind and at the side and rear, such as the sirens of special vehicles like ambulances, police cars, and fire trucks, etc.; S112. Define the acquisition method; By setting a high signal-to-noise ratio, a wide dynamic range, and appropriate directivity to accurately capture the target sound signal; according to the Nyquist sampling theorem, set the sampling rate to twice the highest frequency of the signal, and all frequency sound signals have the same sensitivity; adopt multiple channels and ensure that the sound sensor has the same sensitivity to sound signals in different directions; S113. Digitally process the sound signals; Sample the analog signal at fixed time intervals to discretize the continuous signal; convert the amplitude value of each sampling point into a digital value with finite precision to ensure the dynamic range and precision of the signal; convert the quantized digital value into a binary code for easy storage and transmission, and the coding method used is PCM; S114. Compress and store the data of the sound signals; Adopt lossless compression to completely retain the original sound quality and avoid feature loss caused by compression. Adopt the local storage method to ensure fast data access speed, low latency, and high real-time requirements for scenarios.
[0024] S12. Obtain the driving sound signals of the web crawler; S121. Crawl the sound data based on keywords; Using web crawler technology, capture driving sound signal data from the Internet based on keywords; First, determine the target website and data source, and select a suitable crawler framework and programming language according to the complexity of the target website; Analyze the HTML structure of the target web page and locate the position where the target data is located; Extract information such as the link, title, description, and tags of the target data from the search result page; S122. Download sound data based on multi-threading; Use the method of thread pool to implement multi-threading and improve the crawler efficiency; The sound signal web page parsing module and the download module each maintain a thread pool; According to the remaining capacity of the thread pool, allocate a new thread. When the web page parsing task or the download task is completed, destroy the thread, and at this time the current remaining capacity is incremented by 1; When it is detected that the remaining capacity of the thread pool is 0, wait for the parsing task or the download task to complete and then release the resources, and then allocate resources to the new thread; S123. Store and classify the sound signals; Store the metadata information of the audio file (such as file name, download time, source website, etc.) in a database or file for convenient subsequent management and analysis; Convert the downloaded audio files into a unified format; Classify the audio files and label each segment according to the timestamp, such as the siren sound of a characteristic vehicle, the horn sounds of vehicles in the opposite and side directions, the horn sounds of vehicles in the rear and side-rear directions, etc. Finally, upload the data to the cloud storage service.
[0025] S13. Obtain sound signals based on the adversarial generative network; S131. Preprocess the sound signals; Adopt a noise reduction algorithm to remove the environmental noise in the S11 and S12 audio files and improve the signal quality; Divide the driving scenario sound signals into short-time frames, perform Fourier transform on each frame of the signal to obtain the spectrum; Subtract the noise spectrum from the spectrum of the noisy signal, and finally convert the processed spectrum back to the time-domain signal; Then perform standardization or normalization processing on the noise-reduced sound data and scale the waveform data between [-1, 1]; S132. Build a sound generation model; The sound generation model includes a generator and a discriminator. The generator is constructed using one-dimensional transposed convolution, and the discriminator is constructed using one-dimensional convolution. The input of the generator is a random noise vector, and the output is the generated sound waveform data. The generator includes multiple fully connected layers or convolutional layers. First, the random noise data is mapped to an intermediate dimension through a fully connected layer. Then, an anti-convolution layer is used to gradually upsample the data to the dimension of the target waveform data. The output layer outputs the generated sound waveform data, with the same dimension as the real data. The input of the discriminator is the real driving sound signal data or the fake data generated by the generator, and the output is a scalar representing the probability that the input data is real data. The binary cross-entropy loss function is used to alternately train the generator and the discriminator, enabling the discriminator to correctly distinguish between real data and fake data, and enabling the fake data generated by the generator to deceive the discriminator into believing that the fake data is real. Finally, data identical to the real driving sound waveform data is generated. S133. Perform reliability analysis on the sound signal data generated in S132 to verify the effectiveness of the data. Perform post-processing (such as noise reduction, smoothing, etc.) on the generated sound signal to further improve the signal quality. Use objective metrics such as signal-to-noise ratio and logarithmic spectral distance to evaluate the quality of the generated signal. And evaluate the quality of the generated signal through listening tests to ensure that the timbre, rhythm, and dynamic range of the driving sound signal conform to the auditory characteristics of the real scenario. S134. Integrate the sound signal data for complex driving scenarios. Integrate the reliable driving sound data generated in S11, S12, and S132 L , L =( L 1, L 2,..., L i ,..., L N ) for a total of N pieces of driving sound signal data. Among them, the data length of each sound waveform signal data L i is different.
[0026] S2. Construct a sound recognition model for hearing-impaired drivers as follows: S21. Construct a driving scenario sound signal data set. S211. Extract data features. Extract features from the equal-length sound waveform data integrated with S134, including time-domain feature extraction and frequency-domain feature extraction; the time-domain features include mean, variance, and peak value; use the fast Fourier transform to convert the time-domain sound signal into a frequency-domain sound signal, expressed as a two-dimensional array of frequency and time, calculate the power spectrum from the complex matrix of the Fourier transform, and perform matrix multiplication on the power spectrum and the Mel filter bank to obtain the Mel spectrogram, where the horizontal axis is time and the vertical axis is the Mel frequency; S212. Label the data; Label the collected driving sound data; classify and organize the driving sound data according to the scenario and sound category, and assign one or more category labels to the sound in each time period; for example, the siren of a special vehicle in the surrounding environment, the horn sound of the vehicle behind and on the rear side, and the emergency braking sound of the surrounding vehicles.
[0027] S213. Divide the training set, test set, and validation set; Divide the driving sound data into equal-length sound wavelength data according to the sliding window method, and the sliding window is w , and the sliding step is s ; the calculation formula for the number of samples in each sound data is shown in formula (1): (1) In the formula, is the length of the i-th sound data sample; N The calculation formula (2) for the total number of samples of the driving sound signal data is: (2) Finally, divide the samples into the training set, test set, and validation set according to a ratio. For example: the training set is 70% for model training, the test set is 20% for parameter tuning and model selection, and the test set is 10% for the final evaluation of the model.
[0028] S22. Build a driving sound recognition model that fuses time-frequency domain features; S221. Time-domain feature representation learning based on the attention mechanism; Extract time series features from the time-domain waveform signal; the sound waveform signal data with timestamps X ={ x 1, x 2,…, x T}, where x t is the waveform signal at time t , and T is the time step; the original sound waveform signal , and the waveform signal after position encoding is expressed by formula (3): (3) Extract features through an encoder, which consists of multiple layers of multi-head self-attention mechanisms and feed-forward neural networks, and output the feature representation results H t , ; (4) S222. Frequency-domain feature representation learning based on CNN; Perform short-time Fourier transform on X to obtain a spectrogram , where F is the frequency dimension, is the number of time frames; obtain the intermediate feature representation results through multiple layers of 2D convolution and pooling operations P , and the calculation formula (5) is: (5) Finally, obtain the frequency-domain feature representation results through dimension transformation and flattening operations H f , ; S223. Fuse features based on cross-attention mechanism; Fuse the time-domain feature representation results H t obtained by S221 through the cross-attention mechanism, H f and the frequency-domain feature representation results of S222 H fused to obtain the time-frequency domain fusion representation results (6) The calculation formula (7) of CrossAttention is: (7) Through the learnable weight matrix W Q ,perform a linear transformation on H t to generate the corresponding Q vector, and perform a linear transformation on W K and W V on H f to generate the corresponding K , V vector; S224. Generate text for vehicle signal recognition; Through the Transformer decoder, the word probability distribution is generated through linear projection and Softmax; the decoder consists of l layers, including self-attention, cross-attention, and feed-forward networks; for the l th layer of the decoder, the input is the output of the previous layer, and the initial input is the embedding of the target sequence Y = y 1, y 2,…, y (t-1) , plus the positional encoding; For the processing of the self-attention sublayer, the decoding result Zs is obtained through residual connection and layer normalization. The calculation formula (8) of Zs is: (8) Where , , , W Qs , W Ks and W Vs are learnable parameters; For the processing of the cross-attention sublayer, the decoding result Zc is obtained. The calculation formula (9) of Zc is: (9) Where , , , W Qc , W Kc and W Vc are learnable parameters; Then, through the feed-forward network sublayer, the calculation formula (10) is: (10) After residual connection and layer normalization, the representation result of the l th layer is obtained. The calculation formula (11) is: (11) The word probability distribution is generated through linear projection and Softmax. The calculation formula (12) is: (12) Finally, the text output of the sound signal is realized.
[0029] S23. Perform performance test and evaluation on the driving sound recognition model; S231. Perform offline test and evaluation on the model; Offline testing and evaluation of the model. In this step, the Mean Absolute Error (MAE), Root Mean Square Error (RMSE), and Mean Absolute Percentage Error (MAPE) are used as evaluation metrics for the driving sound recognition model to evaluate the difference between the predicted value and the true value of the sound signal recognition model. The calculation formulas (13), (14), and (15) of MAE, RMSE, and MAPE are as follows: (13) (14) (15) Wherein, and are the true sound signal text and the predicted sound signal text respectively; S232. Conduct in-loop testing and evaluation of the model; Deploy the trained model to in-vehicle hardware; ensure that the model can process sound data in real time when running on the hardware and provide timely feedback; and verify the real-time recognition ability of the model for key sounds (such as ambulance sirens, collision sounds, abnormal tire noise, emergency braking sounds) in the real driving environment.
[0030] S3. Build a vehicle head-up display system for hearing-impaired drivers; S31. Define the architecture of the vehicle head-up display system; S311. Design system indicators and parameters; Define the indicators of the head-up display system, including optical performance parameters such as the resolution and contrast of the displayed information, horizontal and vertical field of view angles, brightness, and chromaticity uniformity. It is necessary to meet the test under temperature and humidity conditions to ensure the reliability and safety in various environments; S312. Compose the head-up display system; The head-up display system includes a MEMS scanning mirror, a laser, a laser diode driver, and a projection screen; the shape and brightness of the laser beam can be flexibly adjusted through the driver, and the light source provided by the laser is controlled by the scanner to achieve high-resolution image projection.
[0031] S313. Generate MEMS laser projection imaging; Combining microelectromechanical systems and laser scanning technology, using red, green, and blue (RGB) three primary color lasers as light sources, the laser is collimated, focused, and modulated through optical elements (such as lenses, mirrors, etc.). After passing through the MEMS micromirror, the laser beam is scanned according to the pixel information of the image data, and finally a complete image is formed on the screen; S32. Obtain the vehicle head-up display information; S321. Convert the sound signal into text information; Deploy the driving sound recognition model that fuses time-frequency domain features constructed in S22 on the vehicle hardware. Obtain the driving sound signal data through the in-vehicle sound sensor and input it into the vehicle control unit (VCU) via the CAN bus. After processing, output the corresponding text information; S322. Obtain the vehicle driving status information; Obtain the acceleration and angular velocity through the inertial navigator, and perform integral processing to obtain the speed and position information; Monitor the voltage, current, and temperature of the battery through the battery management system, and calculate the battery capacity and driving range as the driving status information; S323. Generate the vehicle collision warning information; Obtain the original data of the target object through the lidar, millimeter-wave radar, ultrasonic radar, and camera. After data processing and fusion, combine the vehicle's own speed and acceleration, calculate the relative motion relationship with the target object, and predict the collision time to generate the vehicle collision warning information; S324. Integrate the vehicle head-up display information; Integrate the sound signal text information, driving status information, and vehicle collision warning information output by S321, S322, and S323; The VCU sends the information to the head-up display system controller according to the fixed data format and message format. The head-up display system controller receives the signal, converts the received signal into a visual graph, and projects it onto the screen through the optical projection system.
[0032] S33. Design a vehicle head-up display strategy for hearing-impaired drivers.
[0033] S331. Design the head-up display interface; Project the display information to different positions on the display screen according to different priorities; Take the driving sound signal as the first priority and project it to the center of the screen where the driver is looking. The display information of the second priority is displayed below the center of the screen, and the display information of the fourth priority and the third priority are displayed on the left and right sides respectively; S332. Determine the priority of the vehicle head-up display information; S4. The driver assists in driving according to the display information of the vehicle head-up display system.
[0034] Place the sound signal information that a hearing - impaired driver cannot obtain at the first priority to help the driver understand the driving environment around the vehicle. The second - priority display information includes collision warning information, vehicle battery power shortage prompt information, lane departure information, and fault information. The third - priority display information is navigation information, including lane information, lane line information, traffic light information, and traffic sign information. The fourth - priority display information is external environment interaction information, including incoming call display information, weather information, and vehicle interior temperature information.
[0035] Embodiment 2: This embodiment provides a vehicle head - up display system for hearing - impaired drivers, which is used to implement the vehicle head - up display method for hearing - impaired drivers described in Embodiment 1, and includes: An acquisition module, which is used to acquire sound signals for a complex driving environment; A first construction module, which is used to construct a voice recognition model for hearing - impaired drivers; A second construction module, which is used to construct a vehicle head - up display system for hearing - impaired drivers.
[0036] In summary, the present invention enables hearing - impaired drivers to better perceive the surrounding driving environment, improves driving safety, and solves the problem that hearing - impaired drivers cannot perceive the sound information of the surrounding driving environment.
[0037] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A vehicle head-up display method for hearing-impaired drivers, Its characteristics include the following steps: S1. Acquire sound signals for complex driving environments; S2. Build a voice recognition model for hearing-impaired drivers; S3. Build a vehicle head-up display system for hearing-impaired drivers.
2. A vehicle head-up display method for hearing-impaired drivers according to claim 1, characterized in that: The specific method of S1 is as follows: S11, obtaining a sound signal based on a sensor; S12, obtaining a driving sound signal of the network crawler; S13. Obtain a sound signal based on a generative adversarial network.
3. A vehicle head-up display method for hearing-impaired drivers according to claim 2, Its characteristic is that the specific method of S11 is as follows: S111. Arrange vehicle sound sensors; A sound sensor is installed on each side of the front of the vehicle to collect sound signals from the front and sides; a sound sensor is installed on each side of the rear of the vehicle to obtain the horn sound of the vehicle behind and behind the side; S112, define the collection method; According to the Nyquist sampling theorem, the sampling rate is set to twice the highest frequency of the signal, and the sound signals of all frequencies have the same sensitivity; multi-channel is used, and the sound sensor is guaranteed to have the same sensitivity to sound signals from different directions; S113, digitally processing the sound signal; The analog signal is sampled at fixed time intervals to discretize the continuous signal; the amplitude value of each sampling point is converted into a digital value with finite precision; the quantized digital value is converted into a binary code, and the encoding method used is PCM; S114, compressing and storing the data of the sound signal; Use lossless compression and local storage.
4. The vehicle head-up display method for hearing-impaired drivers according to claim 2, characterized in that: The specific method of S12 is as follows: S121, capturing sound data based on keywords; Using web crawler technology, the driving sound signal data is captured from the Internet based on keywords. First, the target website and data source are determined, the HTML structure of the target web page is analyzed, and the location of the target data is located. Extract the link, title, description, and tags of the target data from the search results page; S122, downloading multi-threaded sound data; The multithreading is realized by using the thread pool method; a thread pool is maintained respectively by using the sound signal web page parsing module and the downloading module; According to the remaining capacity of the thread pool, a new thread is allocated. When the web page parsing task or downloading task is completed, the thread is destroyed, and the current remaining capacity is increased by 1. When the remaining capacity of the thread pool is detected to be 0, resources are released after the parsing task or downloading task is completed, and then resources are allocated to the new thread. S123, storing and classifying the sound signals; Store metadata information of audio files in a database or file; Convert the downloaded audio files into a unified format; classify the audio files and label each segment according to the timestamp; finally upload the data to the cloud storage service.
5. The vehicle head-up display method for hearing-impaired drivers according to claim 2, characterized in that: The specific method of S13 is as follows: S131, preprocessing the sound signal; A noise reduction algorithm is used to remove the ambient noise in the S11 and S12 audio files; the driving scene sound signal is divided into short time frames, and each frame signal is Fourier transformed to obtain a spectrum; the noise spectrum is subtracted from the spectrum of the noisy signal, and finally the processed spectrum is converted back to the time domain signal; then the noise-reduced sound data is standardized or normalized, and the waveform data is scaled to between [-1,1]; S132, constructing a sound generation model; The sound generation model includes a generator and a discriminator, the generator is constructed by one-dimensional transposed convolution, and the discriminator is constructed by one-dimensional convolution; the input of the generator is a random noise vector, and the output is the generated sound waveform data; the generator includes multiple fully connected layers or convolutional layers; first, the random noise data is mapped to an intermediate dimension through a fully connected layer through the noise vector; the deconvolution layer is used to gradually upsample the data to the dimension of the target waveform data; the output layer outputs the generated sound waveform data, and the dimension is the same as the real data; the input of the discriminator is the real driving sound signal data or the false data generated by the generator, and the output is a scalar, which represents the probability that the input data is the real data; the binary cross entropy loss function is used to alternately train the generator and the discriminator, so that the discriminator can correctly distinguish between real data and false data, and the false data generated by the generator can deceive the discriminator into thinking that the false data is real, and finally generate data identical to the real driving sound waveform data; S133, performing reliability analysis on the sound signal data generated in S132; Post-process the generated sound signal; use the signal-to-noise ratio and log spectrum distance objective indicators to evaluate the quality of the generated signal. The higher the signal-to-noise ratio, the better the signal quality; the smaller the log spectrum distance, the closer the generated signal is to the target signal; and evaluate the quality of the generated signal through listening tests to ensure that the timbre, rhythm and dynamic range of the driving sound signal meet the auditory characteristics of the real scene; S134, integrating sound signal data for complex driving scenarios; Integrate the reliability driving sound data generated by S11, S12 and S132 L , L =( L 1, L 2,..., L i ,..., L N )total N Each piece of driving sound signal data has a sound waveform signal data L i The data length is different.
6. The vehicle head-up display method for hearing-impaired drivers according to claim 5, characterized in that: The specific method of S2 is as follows: S21, constructing a driving scene sound signal dataset; S22, constructing a driving sound recognition model integrating time-frequency domain features; S23. Perform performance test and evaluation on the driving sound recognition model.
7. The vehicle head-up display method for hearing-impaired drivers according to claim 6, characterized in that: The specific method of S21 is as follows: S211, extracting data features; Performing feature extraction on the equal-length sound waveform data integrated by S134, including time domain feature extraction and frequency domain feature extraction; the time domain features include mean, variance and peak value; using fast Fourier transform to convert the time domain sound signal into a frequency domain sound signal, expressed as a two-dimensional array of frequency and time, calculating the power spectrum from the complex matrix of Fourier transform, performing matrix multiplication of the power spectrum with the Mel filter bank, and obtaining a Mel spectrum diagram, wherein the horizontal axis is time and the vertical axis is Mel frequency; S212, labeling data; The collected driving sound data is labeled; the driving sound data is classified and sorted according to scenes and sound categories, and one or more category labels are assigned to the sound in each time period; S213, dividing into training set, test set and validation set; The driving sound data is divided into equal-length sound wavelength data according to the sliding window method. The sliding window is w , the sliding step length is s ; The calculation formula for the number of samples of each sound data is shown in (1): (1) In the formula, is the length of the i-th sound data sample; N The calculation formula (2) for the total number of driving sound signal data samples is: (2) Finally, the samples are divided into training set, test set and validation set according to the proportion.
8. The vehicle head-up display method for hearing-impaired drivers according to claim 6, characterized in that: The specific method of S22 is as follows: S221, Temporal feature representation learning based on attention mechanism; Extract time series features from time domain waveform signals; Sound waveform signal data with time stamp X ={ x 1, x 2,…, x T },in, x t For time t The waveform signal, T is the time step; the original sound waveform signal , the waveform signal after position encoding is expressed as formula (3): (3) The encoder extracts features, which consists of a multi-layer multi-head self-attention mechanism and a feedforward neural network, and outputs features to represent the results. H t , ; (4) S222, Frequency domain feature representation learning based on CNN; right X Perform short-time Fourier transform to obtain the spectrum diagram ,in, F is the frequency dimension; is the number of time frames; the intermediate feature representation result is obtained through multi-layer 2D convolution and pooling operations P , the calculation formula (5) is: (5) Finally, the frequency domain feature representation result is obtained through dimension transformation and flattening operation. H f , ; S223, fusion of features based on cross-attention mechanism; The temporal feature representation results of S221 obtained by the cross attention mechanism H t , and S222 frequency domain feature representation results H f Fusion is performed to obtain the time-frequency domain fusion representation result H fused , the calculation formula (6) is: (6) The calculation formula (7) of CrossAttention is: (7) Through the learnable weight matrix W Q ,right H t Perform a linear transformation to generate the corresponding Q vector, through the weight matrix W K and W V right H f Perform a linear transformation to generate the corresponding K , V vector; S224, generating a traffic signal recognition text; Through the transformer decoder, word probability distribution is generated through linear projection and Softmax; the decoder consists of 1 layer, including self-attention, cross-attention and feedforward network; for the lth layer of the decoder, the input is the output of the previous layer, and the initial input is the embedding of the target sequence Y =[ y 1, y 2,…, y (t-1) ], plus positional encoding; For the processing of the self-attention sub-layer, the decoding result Zs is obtained through residual connection and layer normalization. The calculation formula (8) of Zs is: (8) in , , , W Qs , W Ks and W Vs is a learnable parameter; For the processing of the cross attention sub-layer, the decoding result Zc is obtained, and the calculation formula (9) of Zc is: (9) in , , , W Qc , W Kc and W Vc is a learnable parameter; Then, through the feedforward network sublayer, the calculation formula (10) is: (10) After residual connection and layer normalization, we get l The layer representation result is calculated by formula (11): (11) The word probability distribution is generated by linear projection and Softmax, and the calculation formula (12) is: (12) Finally, the sound signal text output is realized.
9. The vehicle head-up display method for hearing-impaired drivers according to claim 6, characterized in that: The specific method of S23 is as follows: S231, performing offline testing and evaluation on the model; The mean absolute error, root mean square error and mean absolute percentage error are used as evaluation indicators of the driving sound recognition model to evaluate the difference between the predicted value and the true value of the sound signal recognition model; the calculation formulas (13), (14) and (15) of MAE, RMSE and MAPE are as follows: (13) (14) (15) In the formula, and They are respectively the real sound signal text and the predicted sound signal text; S232, conducting in-loop testing and evaluation on the model; Deploy the trained model to the vehicle hardware; ensure that the model can process sound data in real time and provide timely feedback when running on the hardware; and verify the model's real-time recognition capability for key sounds in a real driving environment.
10. The vehicle head-up display method for hearing-impaired drivers according to claim 8, characterized in that: The specific method of S3 is as follows: S31. Define the vehicle head-up display system architecture; S32, obtaining vehicle head-up display information; S33. Design a vehicle head-up display strategy for hearing-impaired drivers.
11. The vehicle head-up display method for hearing-impaired drivers according to claim 10, characterized in that: The specific method of S31 is as follows: S311. Design system indicators and parameters; Define the indicators of the head-up display system, including the resolution and contrast of the displayed information, horizontal and vertical field of view, brightness and color uniformity; S312, forming a head-up display system; The head-up display system includes a MEMS scanning mirror, a laser, a laser diode driver and a projection screen; S313, generating MEMS laser projection imaging; Combining micro-electromechanical systems and laser scanning technology, red, green and blue primary color lasers are used as light sources, and optical elements are used to collimate, focus and modulate the lasers. After the laser beam passes through the MEMS micro-vibration mirror, it is scanned according to the pixel information of the image data, and finally a complete image is formed on the screen.
12. The vehicle head-up display method for hearing-impaired drivers according to claim 10, characterized in that: The specific method of S32 is as follows: S321, converting the sound signal into text information; The driving sound recognition model built by S22 that integrates time-frequency domain features is deployed on the vehicle hardware. The driving sound signal data obtained by the on-board sound sensor is input into the vehicle controller VCU through the CAN bus, and the corresponding text information is output after processing. S322, obtaining vehicle driving status information; The acceleration and angular velocity are obtained through the inertial navigation system, and the speed and position information are obtained through integration processing; the battery voltage, current and temperature are monitored through the battery management system to obtain the battery capacity and driving range; S323, generating vehicle collision warning information; The original data of the target object obtained by LiDAR, millimeter-wave radar, ultrasonic radar and camera is processed and fused, combined with the vehicle's own speed and acceleration, to calculate the relative distance to the target object. d and relative speed v rel , and predict the collision time ttc , ttc = d / v rel , generate vehicle collision warning information; S324, integrating vehicle head-up display information; Integrate the sound signal text information, driving status information and vehicle collision warning information output by S321, S322 and S323; VCU sends the information to the head-up display system controller in a fixed data format and message format. The head-up display system controller receives the signal, converts the received signal into a visual graphic, and projects it on the screen through the optical projection system.
13. The vehicle head-up display method for hearing-impaired drivers according to claim 10, characterized in that: The specific method of S33 is as follows: S331, design the head-up display interface; According to different priorities, the display information is projected to different positions of the display screen; the driving sound signal is taken as the first priority and projected to the center of the screen where the driver is looking, the display information of the second priority is displayed below the center of the screen, and the display information of the fourth priority and the third priority are displayed on the left and right sides respectively; S332, determining the vehicle head-up display information priority; Prioritize the sound signal information that hearing-impaired drivers cannot obtain to help drivers understand the driving environment; the second priority display information includes collision warning information, vehicle battery low power prompt information, lane departure information, and fault information; The third priority displays navigation information, including lane information, lane line information, traffic light information and traffic sign information; the fourth priority displays external environment interaction information, including caller ID information, weather information, and in-car temperature information.
14. A vehicle head-up display system for hearing-impaired drivers, used to implement a vehicle head-up display method for hearing-impaired drivers as described in any one of claims 1 to 13, characterized in that: include: An acquisition module, used to acquire sound signals for complex driving environments; The first building block is used to build a voice recognition model for hearing-impaired drivers; The second building block is used to build a vehicle head-up display system for hearing-impaired drivers.
Citation Information
Patent Citations
Liquid crystal display device and HUD (head up display) system
CN110531551A
Head-up display control system and head-up display method
CN119065139A
A car new line display device for anerythrochloropsia and anomalous trichromatism crowd
CN206773941U
Method and apparatus for speech translation, device and computer readable storage medium
CN108766414A
Control voice retelling consistency verification method based on multi-modal fusion
CN113053366A