A vehicle automatic sound production control method based on infrared recognition

By using infrared recognition and machine learning technology, the horn sound pressure level and frequency are dynamically adjusted, solving the problem that the vehicle acoustic control unit cannot adjust the horn sound pressure level in real time in urban environments, thus improving vehicle driving safety and reducing traffic accidents.

CN119821273BActive Publication Date: 2026-01-06LANGSHA TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411912910.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2026-01-06
Estimated Expiration
2044-12-24

AI Technical Summary

Technical Problem

In urban environments, vehicle acoustic control units struggle to adjust horn sound pressure levels in real time based on the movement patterns of animals and people, resulting in suboptimal avoidance effects. Furthermore, excessively high or low sound pressure levels may startle or fail to attract attention, increasing the risk of traffic accidents.

Method used

The system acquires images of the vehicle's front using infrared recognition, identifies heat sources and tracks their movement, generates predictive information using trajectory clustering and prediction algorithms, dynamically adjusts the horn's sound pressure level and frequency using decision neural networks and reinforcement learning models, generates targeted acoustic warning strategies, and optimizes trajectory prediction through machine learning.

Benefits of technology

It enables dynamic adjustment of horn sound pressure level and frequency based on the movement trajectory of animals or people, thereby improving vehicle driving safety and reducing the incidence of traffic accidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119821273B_ABST
    Figure CN119821273B_ABST
Patent Text Reader

Abstract

The application provides a vehicle automatic sound control method based on infrared recognition, comprising: acquiring an infrared image in a predetermined range in front of the vehicle, identifying a heat source target in the image, judging whether the heat source target is an animal or a person according to the shape, size and temperature distribution characteristics of the heat source target; predicting the moving track of the animal or the person in a future period of time according to the typical moving track mode of the animal and the person to obtain predicted moving position and speed information, and delivering the prediction information to an acoustic control unit; the acoustic control unit dynamically adjusts the sound pressure level and frequency of the loudspeaker according to the received predicted moving information and the target category to generate an acoustic warning mode matched with the predicted moving position, speed and target type; while the acoustic signal is emitted, the actual moving track of the animal or the person is continuously tracked, the track prediction method is optimized by comparing the deviation between the actual track and the predicted track, and the best avoidance effect is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information technology, and in particular to a method for automatic vehicle sound control based on infrared recognition. Background Technology

[0002] In urban environments, when animals or people may collide with vehicles, the acoustic control unit needs to adjust the horn sound pressure level in real time based on the movement patterns of the animals and people to achieve the best avoidance effect. However, urban residents are generally not fond of sudden, loud horn sounds and are easily startled. When the system identifies the heat source characteristics of animals and people, the acoustic control unit needs to adjust the horn sound pressure level in real time based on the movement patterns of the animals and people to achieve the best avoidance effect. This requires the acoustic control unit to accurately predict the movement trajectories of animals and people and adjust the horn sound pressure level in advance based on the prediction results, so that the change in sound pressure level matches the movement trajectory of animals and people. However, the movement trajectories of animals and people are often highly uncertain and random, and different species of animals and people have different sensitivities to sound, which poses a great challenge to the prediction and adjustment of the acoustic control unit. At the same time, excessively high sound pressure levels may have a startling effect on animals and people, causing them to make unpredictable behaviors or even causing traffic accidents; while excessively low sound pressure levels may fail to attract the attention of animals and people, failing to achieve the expected avoidance effect. Therefore, how to dynamically adjust the sound pressure level based on the movement patterns of animals and humans, and ultimately achieve the best avoidance effect, is a key technical problem that urgently needs to be solved. Summary of the Invention

[0003] This invention provides a method for automatic vehicle sound control based on infrared recognition, mainly comprising:

[0004] Acquire infrared images of a predetermined area in front of the vehicle, identify heat source targets in the images, and determine whether the heat source target is an animal or a human based on its shape, size, and temperature distribution characteristics.

[0005] For identified animal or human heat source targets, their movement trajectory is tracked in real time to obtain movement trajectory data within a predetermined time period. The trajectory data is analyzed through trajectory clustering algorithm to obtain the movement trajectory pattern of the animal or human.

[0006] Based on the movement trajectory pattern of an animal or human, its movement trajectory within the predetermined time period is predicted to obtain the predicted movement position and speed information, and the prediction information is transmitted to the acoustic control unit.

[0007] The acoustic control unit dynamically adjusts the sound pressure level and frequency of the speaker based on the received predicted movement information and target category, in order to generate an acoustic warning mode that matches the predicted movement location, speed and target type.

[0008] If the target identification result is a pedestrian, the acoustic control unit inputs the target type, predicted trajectory, and speed parameters into the pre-trained decision neural network. Based on historical data, it learns and generates a warning strategy for pedestrians, obtaining the real-time frequency and volume adjustment curve of the prompt tone. The acoustic control unit then generates a dynamic prompt tone signal that changes continuously as the target approaches.

[0009] If the target identification result is an animal or a vehicle, the acoustic control unit inputs the target parameters into the pre-trained reinforcement learning model to obtain the optimal warning time, volume step adjustment amplitude and frequency. Based on this, the acoustic control unit generates a horn signal in real time and adaptively adjusts it according to the target's reaction until the target is safely avoided.

[0010] While emitting acoustic signals, the system continuously tracks the actual movement trajectory of animals or humans. By comparing the deviation between the actual trajectory and the predicted trajectory, the system optimizes the trajectory prediction method to achieve the best avoidance effect.

[0011] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects:

[0012] This invention discloses an automatic vehicle sound warning control method based on infrared recognition. The method acquires images of the vehicle's front using an infrared thermal imager, employs image segmentation and target recognition algorithms to determine whether the heat source target is an animal or a person, and tracks its movement trajectory in real time. Through trajectory clustering and prediction algorithms, the future position and speed of the target can be predicted. Based on the predicted information, the acoustic control unit dynamically adjusts the horn's sound pressure level and frequency to generate a targeted acoustic warning strategy. For pedestrians, a decision neural network is used to generate the optimal warning strategy; for animals or vehicles, a reinforcement learning model is used to determine the best warning scheme. It can also adaptively adjust based on the target's actual reaction and optimize the trajectory prediction model in real time using machine learning algorithms. This intelligent acoustic warning method can effectively improve vehicle driving safety and reduce the incidence of traffic accidents. Attached Figure Description

[0013] Figure 1 This is a flowchart of an automatic vehicle sound control method based on infrared recognition according to the present invention. Detailed Implementation

[0014] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.

[0015] like Figure 1 This embodiment of an automatic vehicle voice control method based on infrared recognition may specifically include:

[0016] Step S101: Acquire an infrared image of a predetermined range in front of the vehicle, identify the heat source target in the image, and determine whether the heat source target is an animal or a human based on the shape, size, and temperature distribution characteristics of the heat source target.

[0017] Infrared images of the vehicle's front are acquired using an infrared thermal imager. A temperature correction curve is established by calibrating the temperature of a blackbody radiation source and the actual temperature. Brightness and temperature corrections are then applied to the infrared image to obtain a first infrared temperature map. The image quality level is assessed based on the signal-to-noise ratio and image clarity metrics of the first infrared temperature map. A second infrared temperature map is obtained by processing the image using a wavelet transform denoising algorithm. The contrast of the second infrared temperature map is enhanced using a Laplacian operator. Region growing and edge detection algorithms are used to segment regions with temperatures higher than a preset threshold value for the background temperature, resulting in a heat source candidate region map. The area size, contour shape, edge curvature, and temperature distribution feature parameters of each candidate region are extracted from the heat source candidate region map. Based on animal and human feature parameters in a pre-established heat source feature database, an infrared target recognition result map with target type markings and temperature distribution feature parameters is generated.

[0018] Specifically, an infrared thermal imager is used to acquire an infrared image of the area in front of the vehicle. A temperature correction curve is established by calibrating the temperature of the blackbody radiation source and the actual temperature. The infrared image is then subjected to brightness and temperature correction to generate a first infrared temperature map. The first infrared temperature map is then evaluated for image quality. The image quality level is assessed using signal-to-noise ratio calculation and image sharpness quantification indicators. When the quality level reaches a preset quality threshold, a wavelet transform denoising algorithm is used to obtain a second infrared temperature map. The contrast of the second infrared temperature map is enhanced using the Laplacian operator. A region growing algorithm and an edge detection algorithm are used to segment regions where the temperature is higher than a preset threshold of the background temperature, resulting in a candidate heat source region map. The area size, contour shape, edge curvature, and temperature distribution feature parameters of each candidate region are extracted from the candidate heat source region map. Based on animal and human feature parameters in a pre-established heat source feature database, a support vector machine classification algorithm is used to determine the candidate region category. For each candidate region category, if it is determined to be an animal, it is marked as an animal target region; if it is determined to be a human, it is marked as a human target region, generating a target type labeling map. Based on the temperature distribution gradient and edge features of the animal or human target area in the target type marker image, a skeleton extraction algorithm is used to identify target posture feature parameters and establish a target temperature distribution feature vector. This target temperature distribution feature vector is then superimposed on the target type marker image to generate an infrared target recognition result image containing target location, type, posture information, and temperature distribution feature parameters. Infrared thermal imagers acquire target temperature distribution information by detecting the infrared energy radiated by the target. For homeothermic organisms like animals and humans, their surface temperature is typically higher than the ambient temperature, appearing as a distinct heat source area in the infrared image. The raw image signal acquired by the infrared thermal imager needs temperature correction. A correction curve is established using the radiation intensity measured at different temperature points by a calibrated blackbody and the actual temperature value. The relationship between blackbody radiation intensity and temperature conforms to the Stefan-Boltzmann law. Within the temperature range of 283K to 313K, a correction point is recorded every 5K, and the temperature correction curve is obtained by least squares fitting. Infrared image quality assessment uses signal-to-noise ratio (SNR) and sharpness indicators. The SNR calculation formula is...

[0019] SNR = 20lg(As / An), where As is the image signal amplitude and An is the noise amplitude. A signal-to-noise ratio greater than 35dB indicates good image quality. Image sharpness is evaluated by calculating the image variance and gradient. A larger variance indicates higher image contrast, and a larger gradient indicates sharper edges. The image variance threshold is set to 225, and the gradient threshold is set to 0.45. Heat source region segmentation uses an improved region growing algorithm, selecting points with temperatures more than 8K higher than the ambient temperature as seed points, expanding outwards in an 8-neighborhood manner. If the temperature difference between adjacent pixels is less than 1.5K, they are merged into the current region. Edge detection uses the Laplacian operator for image enhancement, highlighting the target edge contour features. Target feature extraction includes shape features such as area, perimeter, and roundness, as well as temperature distribution features. The roundness C is calculated using the formula C = 4πA / L², where A is the region area and L is the perimeter. The roundness of human targets is typically between 0.6 and 0.8, while animal targets have lower roundness due to the presence of limbs, approximately between 0.4 and 0.6. Temperature distribution characteristics are characterized by calculating the mean, standard deviation, and kurtosis of the temperature within a region. The temperature distribution in the human torso is relatively uniform, with a standard deviation less than 2K, while in animals, the temperature distribution is more uneven due to fur coverage. Target pose recognition employs a skeleton extraction algorithm. A thinning algorithm extracts the target region into single-pixel-width skeleton lines, calculating the main direction and branch points of these lines. When a human is standing, a distinct vertical main skeleton line is observed, while in animals, a horizontal main skeleton line with multiple branch points is displayed. The temperature distribution feature vector includes parameters such as the region's average temperature, maximum temperature, temperature standard deviation, and temperature kurtosis. These parameters, combined with the target's location coordinates and pose feature parameters, form a target feature description.

[0020] Step S102: For the identified heat source target of the animal or human, track its movement trajectory in real time, obtain the movement trajectory data within a predetermined time period, and analyze the trajectory data through the trajectory clustering algorithm to obtain the movement trajectory pattern of the animal or human.

[0021] A Kalman filter operator is used to track the target area in the heat source target identification result image in real time. A first motion trajectory coordinate sequence is generated based on the centroid coordinates of the target area. Based on the first motion trajectory coordinate sequence, cubic spline interpolation is used to supplement points exceeding the preset coordinate range threshold to obtain a second motion trajectory coordinate sequence. For the second motion trajectory coordinate sequence, a Gaussian filter is used to eliminate coordinate jitter error, and the displacement distance and direction angle between trajectory points are calculated to obtain a third motion trajectory coordinate sequence. Based on the third motion trajectory coordinate sequence, density clustering algorithm is used to extract inflection point coordinates and trajectory segment points, and trajectory feature vectors are calculated to obtain the target motion trajectory feature pattern recognition result.

[0022] Specifically, based on the target type marking area in the heat source target identification result image, a Kalman filter operator is used to obtain the centroid coordinates of the human or animal target region in real time. The target centroid position coordinates are recorded at 200-millisecond sampling intervals to generate a first motion trajectory coordinate sequence. Outlier detection is performed on the first motion trajectory coordinate sequence, removing points exceeding a preset coordinate range threshold. Missing trajectory points are supplemented using cubic spline interpolation to generate a second motion trajectory coordinate sequence. For the second motion trajectory coordinate sequence, a Gaussian filter is used to remove coordinate jitter errors, and the displacement distance and direction angle between adjacent trajectory points are calculated to generate a third motion trajectory coordinate sequence. Based on the third motion trajectory coordinate sequence, the trajectory curvature value and angle change are calculated. The trajectory points are clustered using a density clustering algorithm, and inflection point coordinates and trajectory segmentation points are extracted to generate trajectory feature segmentation data. For the trajectory feature segmentation data, the average velocity, acceleration, rate of curvature change, and rate of change of motion direction for each trajectory segment are calculated to construct a trajectory feature vector. Based on the trajectory feature vector, a support vector machine classifier is used to train the trajectory feature data to establish animal and human motion trajectory models. The newly acquired trajectory is matched using the animal and human motion trajectory model to obtain the target motion trajectory feature pattern recognition result. In moving target tracking, the Kalman filter operator estimates the target position through two stages: prediction and update. The prediction stage predicts the position at the next moment based on the target motion state equation, and the update stage corrects the prediction result by incorporating actual observations. The state equation includes position coordinates and velocity components, and the observation equation reflects the relationship between sensor measurements and the actual state. In practical applications, for uniformly moving targets, the state transition matrix is ​​set as follows:

[0023] The data is defined as [[1,0,dt,0],[0,1,0,dt],[0,0,1,0],[0,0,0,1]], where dt is the sampling time interval of 200 milliseconds. In trajectory data preprocessing, outlier detection uses the 3σ criterion, calculating the mean and standard deviation of trajectory point coordinates. If a point's coordinates deviate from the mean by more than three times the standard deviation, it is considered an outlier. For missing trajectory points, cubic spline interpolation is used to supplement them, with the interpolation function satisfying the requirements of position and first derivative continuity at the nodes. Gaussian filtering uses a 5×5 filter kernel with a standard deviation of 1.5 to eliminate trajectory jitter. During trajectory feature extraction, the displacement distance between adjacent trajectory points is calculated using Euclidean distance, and the direction angle is calculated using the arctangent function. Curvature calculation employs a three-point method. For three consecutive points P1(x1,y1), P2(x2,y2), and P3(x3,y3) in a trajectory point sequence, the area S of the triangle formed by the three points is calculated using a determinant, with curvature K = 4S / (d1²×d2³×d3¹), where dij represents the distance between point i and point j. When grouping trajectory points using a density clustering algorithm, a neighborhood radius of 50 pixels and a minimum point threshold of 8 are set. The density value of each trajectory point is calculated, and points with density values ​​greater than the threshold are designated as cluster centers. Inflection point extraction is based on the change in direction angle. When the change in direction angle of three consecutive trajectory points exceeds 45 degrees, the intermediate point is identified as an inflection point. The trajectory feature vector contains five dimensions: average velocity, maximum acceleration, mean curvature, mean rate of change of direction angle, and number of trajectory segments. Human motion trajectories typically exhibit relatively stable speed and gentle curvature changes, with typical values ​​of: average speed 1.2 m / s, maximum acceleration 0.8 m / s², mean curvature 0.015, mean rate of change of direction angle 12 degrees / s, and 2 to 3 trajectory segments. Animal motion trajectories, on the other hand, show greater speed variations and sharp turns, with typical values ​​of: average speed 2.5 m / s, maximum acceleration 2.2 m / s², mean curvature 0.045, mean rate of change of direction angle 35 degrees / s, and 4 to 6 trajectory segments. The support vector machine uses a radial basis function kernel with a penalty factor of 10, and the training sample size is no less than 1000 sets.

[0024] Step S103: Based on the movement trajectory pattern of the animal or human, predict its movement trajectory within a predetermined time period to obtain the predicted movement position and speed information, and transmit the prediction information to the acoustic control unit.

[0025] A long short-term memory neural network is used to train animal and human motion trajectory feature data. The input features include historical trajectory point coordinate sequences and velocity sequences to obtain trajectory prediction training data. The trajectory prediction training data is processed through a sliding window to calculate the difference value of the position coordinate point sequence within the window, thereby obtaining the target motion acceleration sequence. The target motion acceleration sequence is then calculated using a numerical integration method to obtain the position prediction coordinate sequence. The position prediction coordinate sequence is smoothed, and the target motion direction angle value and motion velocity value are calculated. Finally, the target motion trajectory prediction curve is obtained through a neural network prediction algorithm.

[0026] Specifically, a Long Short-Term Memory (LSTM) neural network is used to train the animal and human motion trajectory feature data. The input features include historical trajectory point coordinate sequences and velocity sequences. The network structure includes an input layer, a hidden layer, and an output layer. The number of neurons in the hidden layer is twice the length of the input sequence, generating trajectory prediction training data. The trajectory prediction training data is processed using a sliding window. The window length is the forward five seconds of trajectory data, and the sliding step is one second. The difference value of the position coordinate point sequence within the window is calculated to obtain the target motion acceleration sequence. Based on the target motion acceleration sequence, the target motion trajectory within the next three seconds is calculated using numerical integration, obtaining a position prediction coordinate sequence. The position prediction coordinate sequence is smoothed, and the target motion direction angle and motion velocity values ​​are calculated. A neural network prediction algorithm is used to obtain the target motion trajectory prediction curve. Based on the target motion trajectory prediction curve, the target position coordinates, motion velocity, and motion direction angle values ​​are extracted to generate a target prediction motion data packet. The target prediction motion data packet is encoded using a serial communication protocol and sent to the acoustic control unit via a serial data interface. The Long Short-Term Memory (LSTM) neural network uses a gating structure to model temporal information, comprising three components: a forget gate, an input gate, and an output gate. The input feature vector contains the target's position coordinate sequence (x, y) and velocity sequence (vx, vy) over a 5-second historical period, sampled every 200 milliseconds, for a total of 25 time points. The network's hidden layer uses 50 neurons, with the hyperbolic tangent activation function, and is trained using stochastic gradient descent with a learning rate of 0.01. In the sliding window processing, the window length is chosen to be 5 seconds, corresponding to the target's attention duration, and the window slides forward by 1 second each time to ensure prediction continuity. Differential operations are performed on the position coordinate sequence within the window, and the acceleration between adjacent points is calculated using the formula a = (v2 - v1) / dt, where v2 and v1 are the velocities at adjacent times, and dt is the sampling time interval of 0.2 seconds. For human motion, typical acceleration values ​​range from -0.8 to 0.8 m / s², while animal motion acceleration can reach -2.2 to 2.2 m / s². Numerical integration employs a modified Euler method with an integration step size of 0.2 seconds to predict the trajectory over the next 3 seconds. The integration formulas are: v(t+dt)=v(t)+a(t)×dt, x(t+dt)=x(t)+v(t)×dt+0.5×a(t)×dt², where v(t) and x(t) are the velocity and position at time t, respectively, a(t) is the acceleration at time t, dt is the time step, and x(t+dt) is the new position at time t+dt, i.e., the position at the next time step. For uniformly accelerated motion, the predicted trajectory exhibits a parabolic shape, while for variablely accelerated motion, it displays more complex curvilinear characteristics. The trajectory prediction curve is smoothed using cubic spline interpolation to ensure the continuity of position, velocity, and acceleration at the nodes.The motion direction angle is calculated using the position difference between adjacent predicted points, θ = arctan((y2-y1) / (x2-x1)). For human motion, the rate of change of direction angle is typically less than 30 degrees / second, while for animal motion, it can reach 90 degrees / second. The motion velocity value is calculated using displacement and time, v = √((x2-x1)² + (y2-y1)²) / dt. The target prediction motion data packet is encapsulated in a fixed format, containing a 32-byte header and a 96-byte data segment. The header includes a data packet identifier, timestamp, and checksum. The data segment contains the position coordinates (x, y), velocity (vx, vy), and direction angle θ of 15 predicted points. Data is transmitted at a baud rate of 115200 via a serial communication protocol, with each data packet taking approximately 10 milliseconds to transmit, meeting real-time requirements. The sampling period of the prediction data is matched to the transmission delay to ensure data continuity and timeliness.

[0027] In step S104, the acoustic control unit dynamically adjusts the sound pressure level and frequency of the speaker based on the received predicted movement information and target category, in order to generate an acoustic warning mode that matches the predicted movement position, speed and target type.

[0028] An adaptive neural network is used to calculate the distance from the speaker to the predicted target location. The sound pressure attenuation coefficient is obtained based on the predicted distance and the sound wave propagation attenuation formula. The sound pressure modulation amplitude is calculated based on the sound pressure attenuation coefficient and the target velocity value. An envelope extraction algorithm is used to generate a sound pressure level parameter sequence. A reference sound waveform corresponding to the target type is extracted from a sound feature database. The spectral distribution characteristics are extracted from the reference sound waveform to obtain sound spectrum parameters. Phase modulation is performed using a sinusoidal carrier wave based on the sound spectrum parameters and the sound pressure level parameter sequence. A modulated sound signal is generated using a digital frequency synthesizer. The modulation depth is proportional to the predicted target velocity.

[0029] Specifically, based on the target type marker and the target location coordinates in the predicted moving data packet, an adaptive neural network is used to calculate the distance from the horn to the predicted target location. The sound pressure attenuation coefficient is calculated using the sound wave propagation attenuation formula, generating a first sound pressure level parameter sequence. This first sound pressure level parameter sequence is normalized, and the sound pressure modulation amplitude is calculated based on the target velocity and displacement distance. An envelope extraction algorithm is then used to generate a second sound pressure level parameter sequence. A reference sound waveform corresponding to the target type is extracted from a sound feature database, which includes animal sound waveforms and human warning sound waveforms. The spectral distribution characteristics of the reference waveform are extracted to obtain sound spectrum parameters. Based on the sound spectrum parameters and the second sound pressure level parameter sequence, a bandpass filter bank is used to adjust the gain of each frequency band, and a Fourier synthesis algorithm is used to generate the target sound waveform. A sinusoidal carrier wave is used to phase-modulate the target sound waveform, with the modulation depth proportional to the predicted target velocity. A digital frequency synthesizer is used to generate a modulated sound signal. The modulated sound signal is then converted from digital to analog, and the gain control quantity is calculated based on the rate of change of signal amplitude within adjacent control cycles. This signal is then used to drive the horn speaker unit to generate an acoustic signal through a power amplifier. Sound waves attenuate with distance as they propagate through the air; the sound pressure level decreases with increasing distance. The attenuation is calculated using the formula ΔL = 20lg(r2 / r1), where r2 is the predicted distance and r1 is the reference distance of 1 meter. When the target is 20 meters from the speaker, the sound pressure attenuation is 26 dB, and at 50 meters, it reaches 34 dB. An adaptive neural network calculates the sound pressure compensation value in real time to keep the target's received sound pressure level within a preset range. The sound pressure modulation amplitude is related to the target's speed. An envelope extraction algorithm is used to calculate the modulation depth, calculated using the formula m = k × v, where k is a proportionality coefficient of 0.2 s / m and v is the target speed in meters per second. For a walking person, the speed is approximately 1.2 m / s, and the modulation depth is 0.24; for a running animal, the speed can reach 2.5 m / s, and the modulation depth is 0.5. The sound pressure level parameter sequence is normalized, mapping the amplitude to between 0 and 1. The sound feature database stores different types of reference sound waveforms with a sampling rate of 44,100 Hz and a quantization precision of 16 bits. Human warning waveforms primarily exhibit a spectrum concentrated in the 500-2000 Hz range, possessing a strong mid-frequency component; animal sound waveforms cover a wider frequency band, reaching 200-8000 Hz. Fourier transform is used to extract the waveform spectral features and calculate the energy distribution of each frequency band. The bandpass filter bank contains eight sub-band filters with center frequencies of 250, 500, 1000, 2000, 4000, and 8000 Hz, respectively, and a bandwidth 0.2 times the center frequency. The gain of each filter is adjusted according to the spectral parameters, with a gain range of -12 to 12 dB. Fourier synthesis uses inverse Fourier transform to convert the adjusted spectrum into a time-domain waveform. Phase modulation employs a sinusoidal carrier with a carrier frequency of 20 kHz to avoid the range of human hearing.The phase offset of the modulated signal φ(t) = β × sin(2πfmt), where β is the modulation index, related to the target speed, and fm is the modulation frequency, taking a value of 2 Hz. The digital frequency synthesizer uses a direct digital synthesis method, generating the modulated carrier through a phase accumulator. The power amplifier uses a Class D amplifier circuit, operating at a frequency of 100 kHz, with a gain dynamic range of 40 dB. The adjacent control period is 20 milliseconds, dynamically adjusting the gain according to the rate of change of the signal amplitude. When the rate of change exceeds a preset threshold, a soft-start method is used for smooth transition to avoid sound distortion. The speaker unit uses a dual voice coil structure, with a frequency response range of 20 to 20000 Hz, a rated power of 50 watts, and a sound pressure sensitivity of 92 dB.

[0030] In step S105, if the target identification result is a pedestrian, the acoustic control unit inputs the target type, predicted trajectory, and speed parameters into the pre-trained decision neural network. Based on historical data, it learns and generates a warning strategy for pedestrians, obtaining the real-time frequency and volume adjustment curve of the prompt tone. The acoustic control unit then generates a dynamic prompt tone signal that changes continuously as the target approaches to effectively remind pedestrians.

[0031] Based on the target recognition results, historical trajectory parameters are extracted from a preset pedestrian trajectory database. These parameters are then processed using a deep neural network to obtain a pedestrian trajectory feature vector. For this feature vector, a recurrent neural network is used to train the historical trajectory parameters, and a time-series unfolding method is employed to generate trajectory prediction data. For the target distance, speed, and direction parameters in the trajectory prediction data, an adaptive filter is used to eliminate trajectory jitter, and a Kalman filter algorithm is used to obtain a smooth pedestrian trajectory curve. For this smooth pedestrian trajectory curve, a reinforcement learning algorithm is used to train warning data in a historical warning database. An acoustic warning decision-maker generates sound frequency modulation parameters and volume modulation parameters, which are then used by a digital power amplifier to drive a speaker array to generate an acoustic signal.

[0032] Specifically, based on the pedestrian target recognition results and the target type label, historical trajectory parameters are extracted from a preset pedestrian trajectory database. A deep neural network is used to process the predicted trajectory data, generating a pedestrian trajectory feature vector. The pedestrian trajectory feature vector is standardized, and the trajectory features are temporally expanded using a stacked decoder. A recurrent neural network is used to train the historical trajectory data, generating pedestrian trajectory prediction data. Based on the target distance, speed, and direction parameters in the pedestrian trajectory prediction data, an adaptive filter is used to eliminate trajectory jitter, and a Kalman filter is used to calculate a smoothing curve for the pedestrian trajectory. For the smoothing curve, a reinforcement learning algorithm is used to train historical warning data to construct an acoustic warning decision-maker, generating sound frequency modulation parameters. Based on the sound frequency modulation parameters, a reference prompt audio is synthesized using a digital waveform generator, and an envelope detector is used to extract the audio energy curve, generating volume modulation parameters. Distance compensation is applied to the volume modulation parameters, and a multi-segment bandpass filter bank is used to adjust the spectral distribution. A digital power amplifier is used to drive a speaker array to generate a dynamically changing acoustic signal. The pedestrian trajectory feature vector contains temporal data in four dimensions: position, velocity, acceleration, and direction, each represented by a 32-bit floating-point number. The deep neural network employs an encoder-decoder structure. The encoder contains three convolutional layers with kernel sizes of 3×3, 5×5, and 7×7, followed by a ReLU activation function and a batch normalization layer. The decoder uses a deconvolutional structure, outputting a 128-dimensional feature vector. The stacked decoder performs temporal unrolling of the trajectory features, with a time window of 5 seconds and a stride of 0.2 seconds. The recurrent neural network uses a bidirectional long short-term memory (LSTM) network structure with 256 hidden neurons. The initial bias values ​​for the input and forget gates are set to 1.0 to avoid gradient vanishing. The training data includes 1000 sets of pedestrian trajectory data from different scenarios, each set containing feature data at 25 time points. The adaptive filter uses the minimum mean square error criterion, with a filter order of 8 and a stride factor μ = 0.01. The Kalman filter state vector contains position and velocity components, and the state transition matrix F = [[1,dt],[0,1]], where dt is the sampling interval of 0.2 seconds. The diagonal elements of the measurement noise covariance matrix R are set to 0.01, and the diagonal elements of the process noise covariance matrix Q are set to 0.001. The reinforcement learning algorithm uses a deep Q-network. The state space includes three dimensions: target distance, velocity, and approach angle, and the action space includes two dimensions: sound frequency and volume. The reward function is designed to consider both warning effect and sound interference. The warning effect is inversely proportional to the target distance, and the interference is directly proportional to the volume. The network contains three fully connected layers with 128, 64, and 32 neurons respectively, and the learning rate is set to 0.001. The digital waveform generator uses a direct digital synthesis method with a clock frequency of 100MHz, a phase accumulator bit width of 32 bits, and a frequency resolution of 0.023Hz.The reference audio signal is synthesized using a sine wave superposition method, with the fundamental frequency set between 1kHz and 2kHz. It contains three harmonic components with an amplitude ratio of 1:0.5:0.25. Envelope detection uses peak detection with a detection time constant of 20ms. The bandpass filter bank contains six sub-bands with center frequencies of 500Hz, 1kHz, 2kHz, 4kHz, and 8kHz, and a bandwidth 0.3 times the center frequency. The filter uses a Butterworth structure with an order of 4. Distance compensation uses the inverse square law, with a compensation amount ΔL = 20lg(r / r0), where r is the target distance and r0 is the reference distance of 1 meter. The power amplifier uses a Class D structure with a switching frequency of 200kHz, an output power of 50W, and a total harmonic distortion of less than 0.1%. The speaker array is linearly arranged with a unit spacing of 0.1 meters, and beam directivity is controlled by adjusting the phase of each unit.

[0033] If the predicted pedestrian movement trajectory intersects with the vehicle's direction of travel, and the distance between the intersection points is less than the safety threshold given by the network, the frequency and volume of the alert audio will be increased to effectively remind the pedestrian to avoid the vehicle. If the predicted pedestrian movement trajectory is parallel to the vehicle's direction of travel, and the distance is less than the safety threshold given by the network, an alert audio with moderate volume and a moderately increased frequency will be generated to remind the pedestrian that the vehicle is approaching.

[0034] Based on the vehicle's driving direction vector and the predicted pedestrian trajectory vector, spatial analytical geometry is used to calculate the coordinates of the intersection point of the trajectory extension lines, obtaining a first trajectory risk parameter and a second trajectory risk parameter. For the first and second trajectory risk parameters, historical safety threshold data is processed using a deep neural network to obtain a safety risk level determination result. If the safety risk level determination result shows that the distance between the trajectory intersection points is less than the safety threshold, a sound frequency growth curve is calculated to generate sound modulation parameters. For the sound modulation parameters, the spectral distribution is adjusted using a Butterworth bandpass filter bank, and a digital power amplifier is used to drive a speaker array to generate an acoustic prompt signal.

[0035] Specifically, based on the vehicle's driving direction vector and the predicted pedestrian trajectory vector, spatial analytical geometry is used to calculate the coordinates of the intersection point of the trajectory extension lines, generating a first trajectory risk parameter. Numerical analysis is performed on the first trajectory risk parameter to calculate the angle between the vehicle's driving direction and the pedestrian's movement direction, generating a second trajectory risk parameter. Based on the first and second trajectory risk parameters, historical safety threshold data is processed using a deep neural network to generate a safety risk level determination result. If the safety risk level determination result shows that the distance between the trajectory intersection points is less than the safety threshold, an exponential function y = 2^x is used to calculate the sound frequency growth curve and volume amplification factor, generating a first sound modulation parameter. If the safety risk level determination result shows that the trajectory direction angle is less than the angle threshold, a linear function y = kx + b is used to calculate the sound frequency growth curve and volume amplification factor, generating a second sound modulation parameter. Based on the first or second sound modulation parameter, a digital synthesizer generates a reference prompt audio, and a sine wave modulator dynamically adjusts the waveform frequency to generate a modulated audio signal. For the modulated audio signal, a Butterworth bandpass filter bank is used to adjust the spectral distribution, and a digital power amplifier drives a speaker array to generate an acoustic prompt signal. When calculating the intersection point of trajectories using spatial analytic geometry, the vehicle's direction vector is represented as v1(x1,y1), and the pedestrian's trajectory vector is represented as v2(x2,y2). The equations of the two trajectories are l1: y = k1x + b1 and l2: y = k2x + b2, where the slopes are k1 = y1 / x1 and k2 = y2 / x2. Solving the system of equations yields the intersection point coordinates (x,y): x = (b2 - b1) / (k1 - k2) and y = k1x + b1. When k1 = k2, it indicates that the two trajectories are parallel. In the calculation of trajectory risk parameters, the direction angle is also considered.

[0036] θ = arccos((v1·v2) / (|v1|·|v2|)). When θ < 30 degrees, it is considered parallel; when θ > 60 degrees, it is considered intersecting. For intersecting cases, the distance d from the intersection point to the vehicle's current position is calculated. The safety threshold is dynamically adjusted with speed, with a base value of 20 meters. For every 10 km / h increase in speed, the threshold increases by 5 meters. The deep neural network uses a feedforward structure. The input layer contains four parameters: trajectory intersection distance, direction angle, vehicle speed, and pedestrian speed. The hidden layer uses a three-layer structure with 64, 32, and 16 neurons respectively, and the activation function is ReLU. The output layer outputs the risk level probability distribution through the Softmax function, divided into low, medium, and high levels. When a high-risk intersection is determined, the frequency growth curve uses the exponential function y = 2^x, where x is the normalized distance parameter. The distance increases from the safety threshold to 0, corresponding to x increasing from 0 to 1, and the frequency increases from 1000 Hz to 4000 Hz. The volume increases exponentially, from 60 dB to 90 dB. For parallel proximity, a linear growth function y = kx + b, k = -30, b = 90, is used, with the volume increasing linearly from 60 dB to 75 dB as the distance increases from the safety threshold to 5 meters. The digital synthesizer uses a direct digital synthesis method, generating a reference frequency signal through a phase accumulator. The clock frequency is 100 MHz, and the frequency resolution is 0.023 Hz. Sine wave modulation uses frequency modulation, with the instantaneous frequency ω(t) = ω0 + kf × m(t), where ω0 is the carrier frequency, kf is the frequency modulation sensitivity, and m(t) is the modulation signal. The carrier frequency is chosen to be 2000 Hz, and the frequency modulation range is ±1000 Hz. The Butterworth filter bank contains three sub-bands with center frequencies of 1000, 2000, and 4000 Hz, respectively, a bandwidth of 0.5 times the center frequency, and a filter order of 4. Gain control employs a piecewise linear function, with values ​​of 0-6 dB in the low-frequency range, 3-9 dB in the mid-frequency range, and 6-12 dB in the high-frequency range. The power amplifier utilizes a Class D architecture, a switching frequency of 200 kHz, a maximum output power of 50 watts, and a total harmonic distortion (THD) of less than 0.1%. The speaker array uses a 4-unit linear arrangement with a unit spacing of 0.2 meters, and beam directivity is controlled by adjusting the phase of each unit.

[0037] In step S106, if the target identification result is an animal or a vehicle, the acoustic control unit inputs the target parameters into the pre-trained reinforcement learning model to obtain the optimal warning time point, volume step adjustment amplitude and frequency. Based on this, the acoustic control unit generates a horn signal in real time and adaptively adjusts it according to the target's reaction until the target is safely avoided.

[0038] A reinforcement learning neural network is trained based on historical avoidance data. This neural network generates first target feature data from three dimensions: target distance, velocity, and direction of motion. A sliding time window is used to process the target trajectory data, generating second target feature data by calculating target position, velocity, and acceleration parameters. A reward function is established based on the first target feature data, and this reward function is processed by a Bayesian decision-maker to obtain the optimal warning timing. Acoustic control parameters are calculated for the optimal warning timing, and these parameters are used to generate dynamic alert audio via a digital waveform generator. The dynamic alert audio is updated in real-time based on the target response feature vector.

[0039] Specifically, based on the target type labeling results and target predicted trajectory data, a reinforcement learning neural network is used to train historical avoidance data. The state space includes three dimensions: target distance, velocity, and direction of motion, generating first target feature data. The target trajectory data is processed through a five-second sliding time window with a one-second step size, calculating the target's current position, velocity, and acceleration parameters to generate second target feature data. The second target feature data is standardized, and a reward function is established based on the first target feature data. A Bayesian decision-maker is used to generate the optimal warning timing. Based on the optimal warning timing, it is compared with a preset set of baseline acoustic parameters. A piecewise function is used to calculate the volume and frequency adjustment amounts to generate acoustic control parameters. For the acoustic control parameters, a baseline prompt audio is generated using a digital waveform generator. A step function is used to modulate the audio signal to generate a dynamic prompt audio. Based on real-time detection of the target avoidance status by an infrared image feature recognizer, the target displacement increment and direction change are calculated to generate a target response feature vector. A dynamic programming algorithm is used to process the target response feature vector, updating the acoustic control parameters in real time and adaptively adjusting the dynamic prompt audio to generate the final warning signal. The reinforcement learning neural network employs a deep Q-learning structure. The state space includes three dimensions: target distance, velocity, and direction of motion. The distance ranges from 0 to 50 meters, the velocity from 0 to 20 meters per second, and the direction angle from 0 to 360 degrees, discretely represented as 50 × 20 × 36 states. The action space includes three dimensions: warning timing, volume, and frequency, each discretely represented as 10 values. The reward function design considers both avoidance effectiveness and sound interference. Avoidance effectiveness is related to the angle between the displacement direction and sound interference, while sound interference is related to the sound pressure level. Specifically, the function is defined as R = k1 × cos(θ) - k2 × (L - L0), where θ is the angle between the displacement direction and the desired direction, L is the sound pressure level, L0 is the baseline sound pressure level, R is the final reward value or score, and k1 and k2 are weighting coefficients used to adjust the relative importance of avoidance effectiveness and sound interference in the total reward. A 5-second time window is used to analyze the trajectory data, updated every 1 second to maintain data continuity and real-time performance. The target motion parameters are calculated within the window. The position coordinates are fitted using the least squares method to obtain a smooth trajectory curve. Velocity is calculated through position difference, and acceleration is calculated through velocity difference. For vehicle targets, the typical acceleration range is -3 to 3 m / s², while for animal targets, the acceleration range is larger, reaching -5 to 5 m / s². The Bayesian decision-maker calculates the optimal warning timing based on conditional probability: P(t|s) = P(s|t) × P(t) / P(s), where t is the warning timing and s is the current state. The prior probability P(t) is obtained from historical data statistics, and the likelihood function P(s|t) is modeled using a Gaussian distribution. P(s) represents the total probability of the current state s occurring, and is the weighted average of the probabilities of s occurring at all possible warning timings t.For vehicle targets, the optimal warning distance is directly proportional to speed, using the empirical formula d = v × t0 + d0, where t0 is the baseline reaction time of 2 seconds and d0 is the baseline safe distance of 10 meters. The preset baseline acoustic parameter set includes three levels, corresponding to different levels of urgency. Level 1: 70 dB volume, 1000 Hz frequency; Level 2: 80 dB volume, 2000 Hz frequency; Level 3: 90 dB volume, 4000 Hz frequency. Volume modulation uses a piecewise linear function, and frequency modulation uses an exponential function. For animal targets, the sound frequency range is selected between 2000 and 6000 Hz to avoid their auditory sensitive areas. Target avoidance status recognition uses infrared image processing to extract the target contour feature point sequence and calculate the displacement vector of feature points between adjacent frames. The displacement increment is obtained by vector superposition, and the directional change is calculated by the vector angle. When the displacement increment is greater than 0.5 meters and the directional angle is greater than 30 degrees, the target avoidance is determined to begin. The dynamic programming algorithm adjusts acoustic parameters in real time based on the avoidance status, with the adjustment step size proportional to the degree of avoidance to prevent parameter oscillations. The adaptive adjustment process continues until the target displacement exceeds the safe distance or the direction deviates from the danger zone.

[0040] If the predicted animal / vehicle movement trajectory intersects with the vehicle's direction of travel, and the predicted collision time is less than the safety threshold, a horn signal with maximum volume is generated to prompt the animal / vehicle to take emergency avoidance measures. If the predicted animal / vehicle movement trajectory is parallel to the vehicle's direction of travel, and the distance is less than the safety threshold, the horn signal is gradually increased to guide the animal / vehicle to avoid the collision until a safe distance is created between them.

[0041] Vector analysis is performed based on the vehicle's driving direction vector and the target's predicted trajectory vector to obtain first trajectory feature data, including the coordinates of the trajectory intersection point and the intersection angle. For this first trajectory feature data, historical collision data is processed using a feedforward neural network. The input layer of the feedforward neural network includes three parameters: trajectory intersection distance, intersection angle, and relative velocity, to obtain the collision prediction time. If the collision prediction time is less than a safe time threshold and the trajectory intersection angle is greater than an angle threshold, maximum sound pressure level control parameters are generated. If the collision prediction time is greater than a safe time threshold and the trajectory intersection angle is less than an angle threshold, progressive sound pressure level control parameters are generated. An infrared image processor is used to detect the target's position coordinate sequence, and avoidance feature data is obtained by calculating the target displacement vector and direction angle changes. Based on this avoidance feature data, the sound pressure level control parameters are updated in real time using an adaptive filter.

[0042] Specifically, based on the vehicle's driving direction vector and the target's predicted trajectory vector, vector analysis methods are used to calculate the coordinates of the trajectory intersection points and the intersection angle, generating first trajectory feature data. For this first trajectory feature data, a feedforward neural network processes historical collision data. The input layer includes three parameters: trajectory intersection distance, intersection angle, and relative speed, generating a collision prediction time. Based on the comparison between the collision prediction time and a safe time threshold, if the prediction time is less than the safe time threshold and the trajectory intersection angle is greater than the angle threshold, a step function is used to generate maximum sound pressure level control parameters. If the collision prediction time is greater than the safe time threshold and the trajectory intersection angle is less than the angle threshold, a piecewise linear function is used to generate progressive sound pressure level control parameters. Based on the maximum sound pressure level control parameters or progressive sound pressure level control parameters, a digital waveform generator generates a reference acoustic signal, and a volume modulator generates a dynamic acoustic signal. An infrared image processor is used to detect the target position coordinate sequence in real time, calculate the target displacement vector and direction angle changes, and generate avoidance feature data. Based on the avoidance feature data, an adaptive filter is used to calculate the avoidance degree parameter, and the sound pressure level control parameter is updated in real time to optimize and adjust the dynamic acoustic signal. In the vector analysis calculation, the intersection angle between the vehicle's driving direction vector v1(x1,y1) and the target trajectory vector v2(x2,y2) is calculated using the vector angle formula: θ=arccos((v1·v2) / (|v1|·|v2|)), where v1·v2 is the dot product of v1 and v2, and |v1| and |v2| are the magnitudes of vectors v1 and v2. The coordinates of the trajectory intersection point are obtained by solving a system of linear equations. When the equations of the two trajectory lines are y=k1x+b1 and y=k2x+b2 respectively, the x-coordinate of the intersection point is x=(b2-b1) / (k1-k2). For a vehicle speed of 20 m / s and a target speed of 10 m / s, when the intersection angle is 90 degrees, a collision is expected to occur after 5 seconds. The feedforward neural network employs a three-layer structure. The input layer has three neurons corresponding to the distance to the intersection point, the intersection angle, and the relative velocity. The hidden layer has 64 neurons, and the output layer has one neuron that outputs the collision prediction time. ReLU is used as the activation function to avoid the vanishing gradient problem. The training data includes 1000 historical collision records, trained using stochastic gradient descent with a learning rate of 0.01, resulting in a training error of less than 0.1 seconds. The safe time threshold is dynamically set according to the target type: 3 seconds for vehicle targets and 2 seconds for animal targets. The angle threshold is set to 30 degrees; values ​​greater than this value are considered intersecting trajectories, while values ​​less are considered parallel trajectories. The step function directly increases the sound pressure level to a maximum of 90 dB when a dangerous situation is detected, while the piecewise linear function gradually increases the sound pressure level from 60 dB to 75 dB according to the distance. The digital waveform generator uses a direct digital synthesis method, generating a reference frequency signal through a phase accumulator with a clock frequency of 100 MHz and a frequency resolution of 0.023 Hz.The acoustic signal is synthesized using dual frequencies: 1000 Hz and 2000 Hz for vehicle targets, and 2000 Hz and 4000 Hz for animal targets. Volume modulation uses an exponential envelope, with modulation depth proportional to the urgency of the avoidance. Infrared image processing employs a target tracking algorithm to extract the target centroid coordinate sequence and calculate the displacement vector between adjacent frames. The angle between the displacement direction and the desired avoidance direction is used as an evaluation index for avoidance effectiveness; when the angle is less than 45 degrees and the displacement velocity is greater than 1 m / s, effective target avoidance is considered to have begun. An adaptive filter dynamically adjusts the filter coefficients based on the avoidance severity; the filter order is 8, and the step factor μ = 0.01. Avoidance feature data includes three dimensions: displacement velocity, direction angle, and acceleration, and measurement noise is eliminated using Kalman filtering. The avoidance severity parameter is calculated using fuzzy rules; the input variables are the normalized values ​​of the three feature dimensions, and the output variable is the sound pressure level adjustment. When the avoidance severity reaches the expected target, the sound pressure level gradually decreases until it returns to the ambient background sound pressure level.

[0043] Step S107: While emitting acoustic signals, continuously track the actual movement trajectory of the animal or human, and optimize the trajectory prediction method by comparing the deviation between the actual trajectory and the predicted trajectory to achieve the best avoidance effect.

[0044] A target tracker is used to acquire the centroid coordinate sequence of an animal or human target. The centroid coordinate sequence is smoothed using a Kalman filter to obtain the first measured trajectory data. Based on the first measured trajectory data, the displacement increment and velocity component within the trajectory segment are calculated. The trajectory curve is fitted using the recursive least squares method to obtain the second measured trajectory data. The Euclidean distance between the measured trajectory points and the predicted trajectory points is calculated based on the second measured trajectory data. The trajectory prediction error vector is calculated using the weighted average method to obtain the trajectory correction parameters. Based on the trajectory correction parameters, the current position and velocity of the target are predicted and calculated. By calculating the rate of change of the angle between the target's movement trajectory and the vehicle's driving direction, a fuzzy controller is used to adjust the horn sound.

[0045] Specifically, a target tracker acquires the centroid coordinate sequence of an animal or human target in real time. The trajectory is divided into segments according to a fixed time window, and the target coordinate sequence is smoothed using a Kalman filter to generate the first measured trajectory data. Feature extraction is performed on the first measured trajectory data, and the displacement increment and velocity components within the trajectory segments are calculated. The trajectory curve is fitted using the recursive least squares method to generate the second measured trajectory data. Based on the second measured trajectory data, the Euclidean distance between the measured trajectory points and the predicted trajectory points is calculated, and the trajectory prediction error vector is calculated using a weighted average method to generate trajectory correction parameters. For the trajectory correction parameters, a deep neural network is used for online learning. The input layer includes three parameters: position error, velocity error, and direction error. The network weights are updated using a backpropagation algorithm. Based on the updated neural network, the current position and velocity of the target are predicted, generating optimized predicted trajectory data. The optimized predicted trajectory data is evaluated in real time, and the rate of change of the angle between the target's movement trajectory and the vehicle's driving direction is calculated to generate trajectory evaluation parameters. Based on the trajectory evaluation parameters, a fuzzy controller is used to generate sound pressure level adjustment and frequency modulation parameters to optimize and adjust the horn sound in real time. The target tracker uses infrared image processing to obtain the target centroid coordinates, with a sampling frequency of 50 Hz and a fixed time window length of 5 seconds, with adjacent windows overlapping by 2.5 seconds. The Kalman filter state vector contains position and velocity components, and the state transition matrix F = [[1,dt],[0,1]], where dt is the sampling interval of 0.02 seconds. The diagonal elements of the measurement noise covariance matrix R are set to 0.01, and the diagonal elements of the process noise covariance matrix Q are set to 0.001. The trajectory prediction error is calculated using the weighted Euclidean distance d, where d = √(w1Δx). 2 +w2Δy2+w3Δvx 2The error vector is calculated using an exponentially weighted average (w1 = w2 = 1.0) and a smoothing coefficient α = 0.7. The deep neural network employs a three-layer structure: 3 neurons in the input layer, 64 neurons in the hidden layer, and 2 neurons in the output layer corresponding to the predicted position increment. The activation function is ReLU, the loss function is mean squared error, and the initial learning rate is 0.01, decreasing with each training iteration. Backpropagation uses stochastic gradient descent with a batch size of 32. Trajectory evaluation calculates the angle θ between the vehicle's direction vector and the target's movement direction vector as θ = arccos((v1·v2) / (|v1|·|v2|)), and the rate of change of the angle, dθ / dt, is calculated using difference. When the rate of change exceeds a preset threshold of 0.5 radians / second, it indicates that the target begins to steer to avoid it. The fuzzy controller's input variables are the included angle θ and the rate of change dθ / dt, and its output variables are the sound pressure level adjustment ΔL and the frequency modulation depth m. The fuzzy rules employ Mamdani inference, using a triangular membership function, and the rule base contains 25 rules. The sound pressure level adjustment range is -12 to 12 dB, and the frequency modulation depth ranges from 0 to 1. Defuzzification uses the centroid method, with a control period of 20 milliseconds. For cases with good obstacle avoidance, the sound pressure level gradually decreases, and the modulation depth decreases; for cases with insufficient obstacle avoidance, the sound pressure level remains constant or increases, and the modulation depth increases. In practical applications, if the target is a deer moving on a road, the tracker collects the deer's position coordinates every 0.02 seconds. To reduce the impact of measurement noise, a Kalman filter is used to smooth the original coordinate sequence. For example, when the deer moves from one side of the road to the other, measurement errors may cause the trajectory to appear jagged; filtering yields a smoother trajectory curve. Feature extraction is performed on the smoothed trajectory data, calculating the displacement and velocity between adjacent sampling points. Taking deer movement as an example, if the horizontal displacement between two sampling points is 0.5 meters and the vertical displacement is 0.3 meters, the deer's average speed during this time can be calculated. By fitting the trajectory curve using recursive least squares, a more accurate description of the movement trajectory can be obtained. The measured trajectory is compared with the predicted trajectory, and the distance deviation between trajectory points is calculated. Suppose the predicted trajectory shows the deer will continue moving in a straight line, but the actual trajectory shows the deer starting to turn; this deviation is used to update the prediction model to optimize prediction accuracy. When a large prediction error is detected, the neural network automatically adjusts the weight parameters. The trajectory evaluation phase focuses on the relative motion relationship between the target and the vehicle. For example, when the angle between the deer's direction of movement and the vehicle's direction of travel begins to increase, it indicates that the deer is avoiding an obstacle. The rate of change of this angle can then be calculated to determine the effectiveness of the avoidance behavior. If the rate of change of this angle reaches 0.5 radians per second, it indicates that the deer's avoidance action is significant. Based on the trajectory evaluation results, the fuzzy controller adjusts the parameters of the warning sound accordingly.When the deer begins to effectively avoid the noise, the system gradually reduces the sound pressure level and decreases the frequency modulation depth, making the sound softer. Conversely, if the deer does not avoid the noise sufficiently, the system maintains or increases the sound pressure level and increases the modulation depth, making the sound more alarming. This dynamic adjustment ensures the warning effect while avoiding excessive disturbance.

[0046] The above are only some preferred embodiments of the present invention, but the present invention is not limited thereto, and many improvements and modifications can be made. Any improvements and modifications made based on the basic principles of the present invention should be considered to fall within the protection scope of the present invention.

Claims

1. A method for automatic sound generation control of a vehicle based on infrared recognition, characterized in that, The method comprises: acquiring an infrared image in a predetermined range in front of a vehicle, identifying a heat source target in the image, judging whether the heat source target is an animal or a person according to the shape, size and temperature distribution characteristics of the heat source target; tracking the moving track of the identified animal or human heat source target in real time, acquiring moving track data within a predetermined time, analyzing the track data through a track clustering algorithm, and obtaining the moving track mode of the animal or person; according to the moving track mode of the animal or person, predicting the moving track of the animal or person within the predetermined time to obtain predicted moving position and speed information, and transmitting the prediction information to an acoustic control unit; the acoustic control unit dynamically adjusts the sound pressure level and frequency of the loudspeaker according to the received prediction information and target category to generate an acoustic warning mode matched with the predicted moving position, speed and target type; if the target identification result is a pedestrian, the acoustic control unit inputs the target type, predicted track and speed parameters into a pre-trained decision neural network to generate a warning strategy for the pedestrian according to historical data learning, obtains a real-time frequency and volume adjustment curve of the prompt sound, and the acoustic control unit generates a dynamic prompt sound signal according to the real-time frequency and volume adjustment curve of the prompt sound, which continuously changes as the target approaches; if the target identification result is an animal or a vehicle, the acoustic control unit inputs the target parameters into a pre-trained reinforcement learning model to obtain the best warning time point, volume step adjustment amplitude and frequency, and the acoustic control unit generates a loudspeaker signal in real time according to the best warning time point, volume step adjustment amplitude and frequency, and adaptively adjusts according to the target reaction until the safety avoidance target is achieved.

2. The method of claim 1, wherein, The method comprises: acquiring an infrared image in a predetermined range in front of a vehicle, identifying a heat source target in the image, judging whether the heat source target is an animal or a person according to the shape, size and temperature distribution characteristics of the heat source target; acquiring an infrared image in a predetermined range in front of a vehicle, identifying a heat source target in the image, judging whether the heat source target is an animal or a person according to the shape, size and temperature distribution characteristics of the heat source target; according to the moving track mode of the animal or person, predicting the moving track of the animal or person within the predetermined time to obtain predicted moving position and speed information, and transmitting the prediction information to an acoustic control unit; the acoustic control unit dynamically adjusts the sound pressure level and frequency of the loudspeaker according to the received prediction information and target category to generate an acoustic warning mode matched with the predicted moving position, speed and target type; 3. The method of claim 1, wherein, if the target identification result is a pedestrian, the acoustic control unit inputs the target type, predicted track and speed parameters into a pre-trained decision neural network to generate a warning strategy for the pedestrian according to historical data learning, obtains a real-time frequency and volume adjustment curve of the prompt sound, and the acoustic control unit generates a dynamic prompt sound signal according to the real-time frequency and volume adjustment curve of the prompt sound, which continuously changes as the target approaches; if the target identification result is an animal or a vehicle, the acoustic control unit inputs the target parameters into a pre-trained reinforcement learning model to obtain the best warning time point, volume step adjustment amplitude and frequency, and the acoustic control unit generates a loudspeaker signal in real time according to the best warning time point, volume step adjustment amplitude and frequency, and adaptively adjusts according to the target reaction until the safety avoidance target is achieved. The method comprises: acquiring an infrared image in a predetermined range in front of a vehicle, identifying a heat source target in the image, judging whether the heat source target is an animal or a person according to the shape, size and temperature distribution characteristics of the heat source target; acquiring an infrared image in a predetermined range in front of a vehicle, identifying a heat source target in the image, judging whether the heat source target is an animal or a person according to the shape, size and temperature distribution characteristics of the heat source target; according to the moving track mode of the animal or person, predicting the moving track of the animal or person within the predetermined time to obtain predicted moving position and speed information, and transmitting the prediction information to an acoustic control unit; the acoustic control unit dynamically adjusts the sound pressure level and frequency of the loudspeaker according to the received prediction information and target category to generate an acoustic warning mode matched with the predicted moving position, speed and target type; if the target identification result is a pedestrian, the acoustic control unit inputs the target type, predicted track and speed parameters into a pre-trained decision neural network to generate a warning strategy for the pedestrian according to historical data learning, obtains a real-time frequency and volume adjustment curve of the prompt sound, and the acoustic control unit generates a dynamic prompt sound signal according to the real-time frequency and volume adjustment curve of the prompt sound, which continuously changes as the target approaches; if the target identification result is an animal or a vehicle, the acoustic control unit inputs the target parameters into a pre-trained reinforcement learning model to obtain the best warning time point, volume step adjustment amplitude and frequency, and the acoustic control unit generates a loudspeaker signal in real time according to the best warning time point, volume step adjustment amplitude and frequency, and adaptively adjusts according to the target reaction until the safety avoidance target is achieved. The Kalman filter operator is used to track the target region in the heat source target recognition result image in real time, and a first motion trajectory coordinate sequence is generated according to the target region centroid coordinates; According to the first motion trajectory coordinate sequence, a second motion trajectory coordinate sequence is obtained by supplementing the points exceeding the preset coordinate range threshold through cubic spline interpolation; For the second motion trajectory coordinate sequence, a third motion trajectory coordinate sequence is obtained by using a Gaussian filter to eliminate coordinate jitter errors, calculating the displacement distance and direction angle between trajectory points, and calculating the trajectory feature vector. According to the third motion trajectory coordinate sequence, a target motion trajectory feature pattern recognition result is obtained by extracting the inflection point coordinates and trajectory segmentation points through a density clustering algorithm.

4. The method of claim 1, wherein, According to the animal or human movement trajectory pattern, the movement trajectory of the animal or human in the predetermined time is predicted to obtain predicted movement position and speed information, and the predicted information is transmitted to the acoustic control unit, including: The long short-term memory neural network is used to train the animal and human motion trajectory feature data, wherein the input features include historical trajectory point coordinate sequences and speed sequences, and trajectory prediction training data is obtained; According to the trajectory prediction training data, the target motion acceleration sequence is obtained by processing through a sliding window and calculating the difference value of the position coordinate point sequence in the window; The numerical integral method is used to calculate the target motion acceleration sequence to obtain the position prediction coordinate sequence; The position prediction coordinate sequence is smoothed, the target motion direction angle value and the motion speed value are calculated, and the target motion trajectory prediction curve is obtained through the neural network prediction algorithm.

5. The method of claim 1, wherein, The acoustic control unit dynamically adjusts the sound pressure level and frequency of the loudspeaker according to the received prediction information and target category to generate an acoustic warning mode matching the predicted movement position, speed and target type, including: The adaptive neural network is used to calculate the distance from the loudspeaker to the predicted target position, and the sound pressure attenuation coefficient is obtained according to the predicted position distance and the sound wave propagation attenuation formula; The sound pressure modulation amplitude is calculated according to the sound pressure attenuation coefficient and the target speed value, and the envelope extraction algorithm is used to generate a sound pressure level parameter sequence; The target type corresponding reference sound waveform is extracted from the sound feature database, and the sound spectrum parameters are obtained by extracting the frequency spectrum features of the reference sound waveform; The sound spectrum parameters and the sound pressure level parameter sequence are phase-modulated by a sinusoidal carrier, and a modulated sound signal is generated by a digital frequency synthesizer, and the depth of modulation is proportional to the target prediction speed.

6. The method of claim 1, wherein, If the target recognition result is a pedestrian, the acoustic control unit inputs the target type, prediction trajectory and speed parameters into the pre-trained decision neural network, learns from historical data to generate a warning strategy for pedestrians, obtains the real-time frequency and volume adjustment curve of the prompt tone, and the acoustic control unit generates a dynamic prompt tone signal according to the above, which changes continuously as the target approaches, including: According to the target recognition result, the historical trajectory parameters are extracted from the preset pedestrian trajectory database, and the pedestrian trajectory feature vector is obtained by processing the historical trajectory parameters through a deep neural network. The historical trajectory parameters are trained by a recurrent neural network for the pedestrian trajectory feature vector, and trajectory prediction data is generated in a time series expansion manner; The target distance, speed and direction parameters in the trajectory prediction data are eliminated by an adaptive filter to eliminate trajectory jitter, and a Kalman filter algorithm is used to obtain a pedestrian trajectory smooth curve; For the pedestrian trajectory smooth curve, a reinforcement learning algorithm is used to train warning data in a historical warning database, and sound frequency modulation parameters and volume modulation parameters are generated by an acoustic warning decision maker to drive a horn array to generate an acoustic signal by a digital power amplifier; It also includes: if the predicted pedestrian movement trajectory intersects with the vehicle driving direction, and the intersection distance is less than the safety threshold given by the network, the prompt sound frequency and volume are increased to effectively remind the pedestrian to avoid; if the predicted pedestrian movement trajectory is parallel to the vehicle driving direction, and the distance is less than the safety threshold given by the network, a prompt sound with moderate volume and moderately increased frequency is generated to remind the pedestrian that the vehicle is approaching.

7. The method of claim 6, wherein, The if the predicted pedestrian movement trajectory intersects with the vehicle driving direction, and the intersection distance is less than the safety threshold given by the network, the prompt sound frequency and volume are increased to effectively remind the pedestrian to avoid; if the predicted pedestrian movement trajectory is parallel to the vehicle driving direction, and the distance is less than the safety threshold given by the network, a prompt sound with moderate volume and moderately increased frequency is generated to remind the pedestrian that the vehicle is approaching, includes: According to the vehicle driving direction vector and the predicted pedestrian trajectory vector, the intersection coordinates of the trajectory extension line are calculated by using spatial analytic geometry method to obtain the first trajectory risk parameter and the second trajectory risk parameter; for the first trajectory risk parameter and the second trajectory risk parameter, the historical safety threshold data is processed by a deep neural network to obtain a safety risk level determination result; if the safety risk level determination result shows that the trajectory intersection distance is less than the safety threshold, a sound frequency growth curve is calculated to generate a sound modulation parameter; for the sound modulation parameter, the spectral distribution is adjusted by a Butterworth band-pass filter set, and a digital power amplifier is used to drive a horn array to generate an acoustic prompt signal.

8. The method of claim 1, wherein, The if the target recognition result is an animal or a vehicle, the target parameter is input into the pre-trained reinforcement learning model to obtain the best warning time point, the volume step adjustment amplitude and the frequency, and the acoustic control unit generates a horn signal according to the best warning time point, the volume step adjustment amplitude and the frequency in real time, and adjusts adaptively according to the target reaction until the safety avoidance target is achieved, includes: A reinforcement learning neural network is trained according to historical avoidance data, and the reinforcement learning neural network generates first target feature data from three dimensions of target distance, speed and motion direction; The target trajectory data is processed by using a sliding time window, and the sliding time window generates second target feature data by calculating target position, speed and acceleration parameters; According to the first target feature data, a reward function is established, and the reward function processes the second target feature data by a Bayesian decision maker to obtain the optimal warning opportunity; An acoustic control parameter is calculated for the optimal warning timing, and the acoustic control parameter generates a dynamic prompt audio through a digital waveform generator, and the dynamic prompt audio is updated in real time according to a target response feature vector; The method further comprises: if the predicted moving track of the animal / vehicle intersects with the driving direction of the vehicle and the predicted collision time is less than a safety threshold, a horn signal with a maximum volume is generated to prompt the animal / vehicle to take emergency avoidance measures; and if the predicted moving track of the animal / vehicle is parallel to the driving direction of the vehicle and the distance is less than the safety threshold, a gradually increasing volume is used to guide the animal / vehicle to take avoidance measures in advance until the safety distance is ensured.

9. The method of claim 8, wherein, The method further comprises: if the predicted moving track of the animal / vehicle intersects with the driving direction of the vehicle and the predicted collision time is less than a safety threshold, a horn signal with a maximum volume is generated to prompt the animal / vehicle to take emergency avoidance measures; and if the predicted moving track of the animal / vehicle is parallel to the driving direction of the vehicle and the distance is less than the safety threshold, a gradually increasing volume is used to guide the animal / vehicle to take avoidance measures in advance until the safety distance is ensured. First track feature data of a track intersection coordinate and an intersection angle are obtained through vector analysis according to a vehicle driving direction vector and a target predicted track vector; A collision prediction time is obtained by processing historical collision data through a feedforward neural network for the first track feature data, and the input layer of the feedforward neural network includes three parameters of a track intersection point distance, an intersection angle and a relative speed; If the collision prediction time is less than a safety time threshold and the track intersection angle is greater than an angle threshold, a maximum sound pressure level control parameter is generated; If the collision prediction time is greater than the safety time threshold and the track intersection angle is less than the angle threshold, a gradual sound pressure level control parameter is generated; A target position coordinate sequence is detected by using an infrared image processor, avoidance feature data is obtained by calculating a target displacement vector and a direction angle change, and a sound pressure level control parameter is updated in real time by using an adaptive filter according to the avoidance feature data.

10. The method of claim 1, wherein, The method further comprises: A target tracker is used to obtain an animal or human target centroid coordinate sequence, and first measured track data are obtained by smoothing the target centroid coordinate sequence through a Kalman filter; A trajectory curve is fitted through a recursive least square method according to a displacement increment and a speed component in a track segment calculated based on the first measured track data, and second measured track data are obtained. A trajectory correction parameter is obtained by calculating an Euclidean distance between a measured track point and a predicted track point for the second measured track data, and calculating a trajectory prediction error vector by using a weighted average method. A target current position and speed are predicted and calculated according to the trajectory correction parameter, a horn sound is adjusted by calculating a change rate of an angle between a target moving track and a vehicle driving direction, and using a fuzzy controller.

Citation Information

Patent Citations

  • Dynamic advancing track processing system and method

    CN117555333A

  • Traffic conflict identification and risk level prediction method and system based on deep learning

    CN118736824A