Traffic safety early warning method for pedestrian crosswalk and related device
Through multi-dimensional Fourier transform analysis, radar echo signal and occlusion analysis and positioning faces, combined with bidirectional flow analysis and pedestrian pass time prediction, the shortcomings of vehicle driving information error and flow analysis in the existing technology are solved, and more reliable traffic safety warning is achieved, and high-quality voice warning is generated through voice token and rhythm analysis.
Patent Information
- Application Number
- CN202510668120.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-23
AI Technical Summary
In the existing technology, in the traffic safety warning of crosswalks, video analysis is prone to video frame loss, resulting in large errors in vehicle driving information, human flow analysis fails to effectively consider face occlusion, insufficient accuracy and fluency of voice conversion, and cannot provide reliable traffic warning.
Multidimensional Fourier transform is used to analyze radar echo signals to obtain vehicle driving information, and faces are positioned through occlusion analysis and multi-layer filtering, two-way traffic analysis is carried out, and vehicle recommended driving information is analyzed based on pedestrian pass time prediction data. Use voice tokens and rhythm analysis to generate target voice information and provide a pass warning.
It improves the accuracy of vehicle driving information and the reliability of traffic data, enhances the intelligence of traffic safety warnings and the accuracy and fluency of voice generation, and provides a more reliable traffic warning.
Smart Images

Figure CN120183169A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent transportation, and particularly to a method and related device for warning of the passing safety of a crosswalk. Background Art
[0002] When a vehicle passes through a crosswalk without traffic lights, due to the driver's failure to take timely deceleration measures or pedestrians' failure to be informed in time that a vehicle is coming, the safety problems between vehicles and pedestrians are increasing. In response to this, how to effectively warn of the passing safety of a crosswalk has become a research focus. For the warning of the passing safety of a crosswalk, the analysis of vehicle driving information is an important link. At present, it is usually realized through video analysis, but video analysis is prone to video frame loss, resulting in large errors in the analyzed vehicle driving information and unable to guarantee the reliability of the passing safety warning. In order to improve the intelligence of the passing safety warning, the recommended driving information of the vehicle is usually analyzed. The statistical analysis of the pedestrian flow of the crosswalk is a key step in the analysis of the recommended driving information. At present, the analysis of the pedestrian flow mainly counts the number of faces appearing in the video frame images, but currently, less consideration is given to face positioning when the face is partially blocked, resulting in incorrect face positioning and large deviations in the obtained pedestrian flow. After obtaining the recommended driving information, it needs to be converted into voice for passing warning. At present, it is mainly through an autoregressive model for voice conversion, but the accuracy of the voice conversion by this method is low, and the fluency is also insufficient, and reliable passing warning cannot be provided. Summary of the Invention
[0003] The purpose of the present invention is to overcome the deficiencies of the prior art. The present invention provides a method and related device for warning of the passing safety of a crosswalk, which can accurately analyze the recommended driving information of a vehicle, so as to provide a more reliable passing warning for vehicles and pedestrians.
[0004] To solve the above technical problems, the present invention provides a method for warning of the passing safety of a crosswalk, which is applied to a radar warning device and a pedestrian warning device, and the radar warning device is communicatively connected with the pedestrian warning device; the method includes: Performing vehicle driving analysis on the radar echo signals collected by the radar warning device in the radar detection area based on multi-dimensional Fourier transform to obtain vehicle driving information; Performing face positioning and multi-layer filtering processing on a plurality of frame target images collected by the pedestrian warning device for the crosswalk based on occlusion analysis to obtain a target face set; Performing two-way pedestrian flow analysis based on the target face set to obtain two-way pedestrian flow data; Predicting the pedestrian passing time of the crosswalk based on the two-way pedestrian flow data and a plurality of frame target images to obtain pedestrian passing time prediction data; Analyze the recommended driving information of the vehicle in the warning area based on the predicted pedestrian passage time data and vehicle driving information; Based on the recommended driving information, use voice tokens and prosody analysis for voice generation to obtain target voice information, and the radar warning device and pedestrian warning device perform passage warnings based on the target voice information.
[0005] Optionally, the vehicle driving analysis is performed on the radar echo signals collected by the radar warning device in the radar detection area by using multi-dimensional Fourier transform to obtain vehicle driving information, including: Perform filtering and noise reduction processing on the radar echo signals to obtain the radar echo signals after filtering and noise reduction processing; Perform fast Fourier transform on the radar echo signals after filtering and noise reduction processing to obtain the Fourier spectrum, and analyze the first vehicle speed, the first vehicle relative distance, and the first vehicle azimuth based on the Fourier spectrum by using the echo signal matrix; Perform ordered-statistic constant false alarm detection on the radar echo signals after filtering and noise reduction processing to obtain the ordered-statistic constant false alarm detection result; Analyze the second vehicle speed, the second vehicle relative distance, and the second vehicle azimuth based on the ordered-statistic constant false alarm detection result by using peak spectral line analysis, and generate the target vehicle speed, the target vehicle relative distance, and the target vehicle azimuth based on the first vehicle speed, the first vehicle relative distance, the first vehicle azimuth, the second vehicle speed, the second vehicle relative distance, and the second vehicle azimuth; Generate vehicle driving information based on the target vehicle speed, the target vehicle relative distance, and the target vehicle azimuth.
[0006] Optionally, the occlusion analysis uses the pedestrian warning device to perform face localization and multi-layer filtering processing on several frames of target images collected from the crosswalk to obtain a target face set, including: Perform passage direction marking on several frames of target images to obtain several frames of target images after passage direction marking; Perform occlusion marking processing on the first face set obtained by localizing several frames of target images after passage direction marking based on a preset occlusion type table to obtain the first face set after occlusion marking processing, and generate a face feature descriptor based on the first face set after occlusion marking processing; Determine the second face set in several frames of target images based on the face feature descriptor by using a face localization model; Filter the second face set based on a face tracking algorithm to obtain a third face set, and filter the third face set based on a feature clustering algorithm to obtain a target face set corresponding to the passage direction.
[0007] Optionally, performing two-way pedestrian flow analysis based on the target face set to obtain two-way pedestrian flow data, including: Performing two-way pedestrian flow analysis based on the Boolean parameter and the detection frame using the target faces marked with each passing direction in the target face set to obtain two-way pedestrian flow data.
[0008] Optionally, predicting the pedestrian passing time of the crosswalk based on the two-way pedestrian flow data and several frames of target images to obtain pedestrian passing time prediction data, including: Converting the pixel coordinates of each target face located in each frame of the target image marked with the passing direction to obtain the corresponding position coordinates; Analyzing the corresponding pedestrian passing speed based on the time stamps of each frame of the target image using the corresponding position coordinates; Predicting the pedestrian passing time of the crosswalk based on the two-way pedestrian flow data and the corresponding pedestrian passing speed to obtain pedestrian passing time prediction data.
[0009] Optionally, analyzing the recommended driving information of the vehicle in the warning area based on the pedestrian passing time prediction data and the vehicle driving information, including: Analyzing the traffic flow density based on the vehicle information in the warning area and analyzing the vehicle deceleration based on the pedestrian passing time prediction data; Calculating the safety distance based on the vehicle driving information in combination with the driver's reaction time, and analyzing the recommended driving information of the vehicle in the warning area based on the traffic flow density, vehicle deceleration, and safety distance.
[0010] Optionally, performing speech generation based on the recommended driving information using voice tokens and prosody analysis to obtain target voice information, including: Generating warning text information based on the recommended driving information using a preset warning text template, and determining the semantic information corresponding to the warning text information; Determining the voice tokens corresponding to the warning text information based on the semantic information using a convolutional encoder, and determining the initial voice based on the voice tokens using a large language model; Performing linear spectral analysis on several frames corresponding to the warning text information to obtain the target linear spectrum; Analyzing the prosody splicing points of the warning text information, and generating prosody analysis data based on the prosody splicing points in combination with the prosody sequence encoding; Adjusting the initial voice based on the target linear spectrum and prosody analysis data in combination with the voice splicing smoothing algorithm and the volume adjustment factor to obtain the target voice information.
[0011] In addition, the present invention also provides a traffic safety warning device for a crosswalk, which is applied to a radar warning device and a pedestrian warning device, and the radar warning device is communicatively connected to the pedestrian warning device; the device includes: A vehicle driving analysis module: configured to perform vehicle driving analysis on the radar echo signals collected by the radar warning device in the radar detection area based on multi-dimensional Fourier transform to obtain vehicle driving information; A face detection module: configured to perform face positioning and multi-layer filtering processing on a plurality of frames of target images collected by the pedestrian warning device for the crosswalk based on occlusion analysis to obtain a set of target faces; A pedestrian flow analysis module: configured to perform two-way pedestrian flow analysis based on the set of target faces to obtain two-way pedestrian flow data; A passing time prediction module: configured to predict the pedestrian passing time of the crosswalk based on the two-way pedestrian flow data and a plurality of frames of target images to obtain pedestrian passing time prediction data; A driving recommendation analysis module: configured to analyze the recommended driving information of the vehicle in the warning area based on the pedestrian passing time prediction data and the vehicle driving information; A warning module: configured to generate target voice information by performing voice generation based on the recommended driving information using voice tokens and prosody analysis, and the radar warning device and the pedestrian warning device perform traffic warnings based on the target voice information.
[0012] In addition, the present invention also provides a traffic safety warning system for a crosswalk, the system includes a radar warning device and a pedestrian warning device, the radar warning device is communicatively connected to the pedestrian warning device, and the system is configured to execute the above-mentioned traffic safety warning method for a crosswalk.
[0013] In addition, the present invention also provides a computer-readable storage medium, the computer-readable storage medium stores computer instructions, and when the computer instructions run on an electronic device, the electronic device is caused to execute the above-mentioned traffic safety warning method for a crosswalk.
[0014] In the embodiments of the present invention, vehicle driving analysis is performed based on multi-dimensional Fourier transform using radar echo signals collected by a radar warning device in a radar detection area, so that the obtained vehicle driving information is more accurate, and the reliability of traffic safety warning is ensured. Based on occlusion analysis, face positioning and multi-layer filtering processing are performed on several frames of target images collected by a pedestrian warning device for a crosswalk to obtain a set of target faces for two-way pedestrian flow analysis. Considering the situation of face occlusion and introducing filtering of the face set, omission detection or duplicate counting is avoided, and the obtained pedestrian flow data is more accurate. Based on pedestrian passing time prediction data and vehicle driving information, recommended driving information of a vehicle in a warning area is analyzed, which can improve the reliability of recommended driving information analysis and at the same time improve the intelligence of traffic safety warning, enabling a driver to understand the current required driving information. Based on the recommended driving information, voice generation is performed using voice tokens and prosody analysis to obtain target voice information for communication warning, improving the accuracy and fluency of voice generation and being able to provide more reliable traffic warnings for vehicles and pedestrians. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1 It is a schematic flowchart of a traffic safety warning method for a crosswalk in an embodiment of the present invention; Figure 2 It is a schematic flowchart of a traffic safety warning method for a crosswalk in another embodiment of the present invention; Figure 3 It is a schematic diagram of the structural composition of a traffic safety warning system for a crosswalk in an embodiment of the present invention; Figure 4 It is a schematic diagram of the structural composition of a traffic safety warning device for a crosswalk in an embodiment of the present invention; Figure 5 It is an application example diagram of a radar warning device and a pedestrian warning device in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0018] Embodiment 1
[0019] Please refer to Figure 1 , Figure 1 , which is a schematic flowchart of the traffic safety warning method for crosswalks in the embodiments of the present invention. The method is applied to a radar warning device and a pedestrian warning device, and the radar warning device is communicatively connected to the pedestrian warning device; the method includes: S11: Based on multi-dimensional Fourier transform, use the radar echo signal collected by the radar warning device in the radar detection area to perform vehicle driving analysis to obtain vehicle driving information; In the specific implementation process of the present invention, the step of using the radar echo signal collected by the radar warning device in the radar detection area to perform vehicle driving analysis based on multi-dimensional Fourier transform to obtain vehicle driving information includes: performing filtering and noise reduction processing on the radar echo signal to obtain the radar echo signal after filtering and noise reduction processing; performing fast Fourier transform on the radar echo signal after filtering and noise reduction processing to obtain the Fourier spectrum, and analyzing the first vehicle speed, the first vehicle relative distance, and the first vehicle azimuth based on the Fourier spectrum using the echo signal matrix; performing ordered-statistic constant false alarm detection on the radar echo signal after filtering and noise reduction processing to obtain the ordered-statistic constant false alarm detection result; analyzing the second vehicle speed, the second vehicle relative distance, and the second vehicle azimuth based on the ordered-statistic constant false alarm detection result using the peak spectral line, and generating the target vehicle speed, the target vehicle relative distance, and the target vehicle azimuth based on the first vehicle speed, the first vehicle relative distance, the first vehicle azimuth, the second vehicle speed, the second vehicle relative distance, and the second vehicle azimuth. Generate vehicle driving information based on the target vehicle speed, the target vehicle relative distance, and the target vehicle azimuth.
[0020] Specifically, the radar warning device emits pulse signals to the radar detection area. When the pulse signals encounter a moving vehicle, radar echo signals will be generated, and the radar warning device collects the radar echo signals. Filter and reduce the noise of the radar echo signals. Filter the radar echo signals through a high-pass filter, and then reduce the noise of the filtered radar echo signals. The noise reduction process can use algorithms such as wavelet transform to obtain the radar echo signals after filter and noise reduction. Perform a fast Fourier transform on the radar echo signals after filter and noise reduction. By transforming the radar echo signals from the time domain to the frequency domain, the spectral structure of the signals is obtained, and the Fourier spectrum is obtained. Based on the Fourier spectrum, use the echo signal matrix to analyze the speed, relative distance, and azimuth of the first vehicle. Extract the peak value in the Fourier spectrum. The peak position is the position of the moving vehicle. Calculate the relative distance of the first vehicle according to the frequency corresponding to the peak position, that is, the distance between the vehicle and the radar. The echo signal matrix is a signal matrix formed by the radar warning device sending the same electromagnetic wave signals to the detection area at preset time intervals and collecting the corresponding echo signals. Extract the peak value of the Fourier spectrum in the echo signal matrix. Generate the speed of the first vehicle according to the frequency corresponding to the peak value of the Fourier spectrum in the echo signal matrix and the frequency modulation slope. Obtain the relative angle between the first vehicle and the radar according to the wavelength of the radar echo signals after filter and noise reduction and the peak frequency of the Fourier spectrum. The azimuth of the first vehicle is obtained from the relative angle and the vehicle position. Perform an ordered-statistic constant false alarm rate (OS-CFAR) detection based on the radar echo signals after filter and noise reduction. The radar echo signals after filter and noise reduction may still contain non-target signals. After performing corresponding signal processing on the radar echo signals, the target is obtained with a relatively high statistical probability, while noise and other interference signals generate false alarms with a relatively low probability, so as to extract the target from the mixed signals. Its purpose is to extract the target from the input signals according to certain criteria, which is the purpose of the OS-CFAR detection. Obtain the corresponding power sampling values through the OS-CFAR detection. Compare the power sampling values with the preset threshold value to extract the target signals according to the comparison results and obtain the OS-CFAR detection results. Based on the OS-CFAR detection results, use the peak spectral line to analyze the speed, relative distance, and azimuth of the second vehicle. Calculate the relative distance of the second vehicle through the peak spectral line in the spectrogram of the target signals extracted by the OS-CFAR detection results. Calculate the speed of the second vehicle according to the relative distance of the second vehicle, the transmission frequency, and the frequency modulation slope of the radar signals. Calculate the azimuth of the second vehicle according to the radar wavelength and the wave path difference. Generate the target vehicle speed, target vehicle relative distance, and target vehicle azimuth based on the speed, relative distance, azimuth of the first vehicle, the speed, relative distance, and azimuth of the second vehicle. The target vehicle speed is obtained by performing a weighted average on the speed of the first vehicle and the speed of the second vehicle. Similarly, the target vehicle relative distance and the target vehicle azimuth are also obtained by the weighted average method.Generate vehicle driving information based on the target vehicle speed, the relative distance of the target vehicle, and the orientation of the target vehicle, so that the obtained vehicle driving information is more accurate.
[0021] S12: Based on occlusion analysis, perform face localization and multi-layer filtering on a number of target images collected by the pedestrian warning device for the crosswalk to obtain a set of target faces; In the specific implementation process of the present invention, the step of performing face localization and multi-layer filtering on a number of target images collected by the pedestrian warning device for the crosswalk based on occlusion analysis to obtain a set of target faces includes: marking the passing directions of a number of target images to obtain a number of target images with passing direction markings; performing occlusion marking on the first face set obtained by localizing the number of target images with passing direction markings based on a preset occlusion type table to obtain a first face set after occlusion marking processing, and generating a face feature descriptor based on the first face set after occlusion marking processing; determining a second face set in the number of target images based on the face feature descriptor using a face localization model; filtering the second face set based on a face tracking algorithm to obtain a third face set, and filtering the third face set based on a feature clustering algorithm to obtain a set of target faces corresponding to the passing direction.
[0022] Specifically, pedestrian warning devices on both sides of the crosswalk collect images of the traffic conditions of the crosswalk to obtain several frames of target images. Mark the traffic directions for the several frames of target images. The traffic direction of the crosswalk is two-way. The pedestrian warning devices on both sides of the crosswalk mark the corresponding traffic directions for the collected target images, that is, mark which traffic direction the frame of target image belongs to according to the side where the pedestrian warning device is located, and obtain several frames of target images after the traffic direction marking. Perform occlusion marking processing on the first face set obtained by positioning the several frames of target images marked with traffic directions based on a preset occlusion type table. Use a deep convolutional neural network to perform face positioning on the several frames of target images marked with traffic directions to obtain the corresponding first face set. Analyze the occlusion types of the face images in the first face set through the preset occlusion type table. The occlusion types include mouth occlusion, nose occlusion, left and right eye occlusion, etc. Analyze the occlusion types of the face images with occlusions in the first face set, and mark the face images with occlusions in the first face set according to the analyzed occlusion types to obtain the first face set after the occlusion marking processing. Generate a face feature descriptor based on the first face set after the occlusion marking processing. Extract face key points from the first face set after the occlusion marking processing. Construct an octagon neighborhood centered on the face key points. Use the face key points to construct several spatial triangles in the octagon neighborhood. Construct several histograms based on each spatial triangle. Connect each histogram in the form of a vector to form a triangular feature descriptor, which is the face feature descriptor. Both the face images with occlusions and the complete face images have corresponding face feature descriptors. Determine the second face set in the several frames of target images based on the face feature descriptor using a face positioning model. Locate the face set in the several frames of target images according to the face feature descriptor using the face positioning model, match this face set with the first face set, and check whether there are missed detections or misdetections. If there is a missed detection, add the missed face image to the first face set. If there is a misdetection, delete the misdetected face image. Perform the above processing on the face set of each frame of target image to form the second face set.Filter the second face set based on the face tracking algorithm, perform face tracking on the face set of each frame of the target image, delete the repeatedly appearing faces, and only retain the initially appearing faces. For the occluded faces, match their unoccluded face features to obtain the third face set. Then filter the third face set based on the feature clustering algorithm. The filtering of the second face set is only a simple filtering, and there may still be duplicate faces. To ensure that no duplicate faces appear to affect the pedestrian flow analysis, it is necessary to further filter the third face set. Extract the face feature values of each face in the third face set, and determine whether the face feature value belongs to the face feature value set of the clustering filter. If it belongs to the face feature value set in the clustering filter, filter out the face from the third face set. The face feature value set of the clustering filter refers to the face feature values of all the previous target faces saved when judging whether the nth target face is similar to the previous target faces. The preset face feature value is an effective parameter value used to filter out the subsequent target faces with the same face feature value, filter out the duplicate faces, and obtain the final face set. Since the passing direction has been marked at the beginning of the target image, the obtained face set also has the corresponding passing direction mark, that is, obtain the target face set corresponding to the passing direction.
[0023] S13: Perform two-way pedestrian flow analysis based on the target face set to obtain two-way pedestrian flow data; In the specific implementation process of the present invention, the performing two-way pedestrian flow analysis based on the target face set to obtain two-way pedestrian flow data includes: performing two-way pedestrian flow analysis based on the boolean parameter and the detection frame using the target faces with passing direction marks in the target face set to obtain two-way pedestrian flow data.
[0024] Specifically, perform two-way pedestrian flow analysis based on the boolean parameter and the detection frame using the target faces with passing direction marks in the target face set. Frame the detection frames for the target faces with passing direction marks, and assign a boolean parameter to each detection frame, initialized to zero. When the detection frame is counted, its value changes from zero to one, which is used to record whether the detection frame has been counted, so as to avoid double counting. Count the number of detection frames to obtain the pedestrian flow data in different passing directions, that is, obtain two-way pedestrian flow data.
[0025] S14: Predict the pedestrian passing time of the crosswalk based on the two-way pedestrian flow data and several frames of target images to obtain pedestrian passing time prediction data; In the specific implementation process of the present invention, predicting the pedestrian passing time of a crosswalk based on two-way pedestrian flow data and several frames of target images to obtain pedestrian passing time prediction data includes: converting the pixel coordinates of each target face located in each frame of the target image after marking the passing direction to obtain corresponding position coordinates; analyzing the corresponding pedestrian passing speed based on the position coordinates by using the time stamps of each frame of the target image; predicting the pedestrian passing time of the crosswalk based on the two-way pedestrian flow data and the corresponding pedestrian passing speed to obtain pedestrian passing time prediction data.
[0026] Specifically, converting the pixel coordinates of each target face located in each frame of the target image after marking the passing direction, that is, converting the pixel coordinates to world coordinates in the world coordinate system to obtain corresponding position coordinates. Analyzing the corresponding pedestrian passing speed based on the position coordinates by using the time stamps of each frame of the target image. The same face in each frame of the target image has been known in face positioning. According to the change of the position coordinates of the same face in each frame of the image and combining the time stamp of each frame of the target image, the corresponding pedestrian communication speed can be calculated. Predicting the pedestrian passing time of the crosswalk based on the two-way pedestrian flow data and the corresponding pedestrian passing speed, inputting the two-way pedestrian flow data and the corresponding pedestrian passing speed into the passing time prediction model to obtain the predicted time of pedestrian passing, that is, obtaining pedestrian passing time prediction data. Since the pedestrian flow in different passing directions will affect the passing speed of pedestrians, the two-way pedestrian flow data is added to predict the passing time, making the obtained prediction data more in line with the actual situation.
[0027] S15: Analyzing the recommended driving information of the vehicle in the warning area based on the pedestrian passing time prediction data and the vehicle driving information; In the specific implementation process of the present invention, analyzing the recommended driving information of the vehicle in the warning area based on the pedestrian passing time prediction data and the vehicle driving information includes: analyzing the traffic flow density based on the vehicle information in the warning area and analyzing the vehicle deceleration based on the pedestrian passing time prediction data; calculating the safety distance based on the vehicle driving information combined with the driver reaction time, and analyzing the recommended driving information of the vehicle in the warning area based on the traffic flow density, vehicle deceleration and safety distance.
[0028] Specifically, the traffic flow density is analyzed based on the vehicle information in the warning area. The vehicle information in the warning area includes the length of the warning area, the number and speed of vehicles in each lane within the warning area, etc. The traffic flow density of each lane can be analyzed according to the number of vehicles in the warning area and the length of the warning area. The vehicle deceleration is predicted based on the pedestrian passage time prediction data, and the vehicle deceleration is analyzed by combining the pedestrian passage time prediction data and the traffic flow density with the vehicle speed in the vehicle driving information. The safe distance is calculated based on the vehicle driving information in combination with the driver's reaction time, that is, the safe distance from the vehicle in front after entering the warning area is determined according to the vehicle speed and vehicle orientation in the vehicle driving information in combination with the driver's reaction time. The recommended driving information of the vehicle in the warning area is analyzed based on the traffic flow density, vehicle deceleration, and safe distance. The lanes that the vehicle can change and the recommended driving speed in the warning area are analyzed according to the traffic flow density, vehicle deceleration, and safe distance. The recommended driving information is generated from the vehicle deceleration, safe distance, lanes that can be changed, and recommended driving speed.
[0029] S16: Based on the recommended driving information, voice generation is performed using voice tokens and prosody analysis to obtain target voice information, and the radar warning device and pedestrian warning device perform traffic warning based on the target voice information.
[0030] In the specific implementation process of the present invention, the voice generation using voice tokens and prosody analysis based on the recommended driving information to obtain target voice information includes: generating warning text information using a preset warning text template based on the recommended driving information, and determining the semantic information corresponding to the warning text information; determining the voice tokens corresponding to the warning text information based on the convolutional encoder using the semantic information, and determining the initial voice using the voice tokens based on the large language model; performing linear spectral analysis on several frames corresponding to the warning text information to obtain the target linear spectrum; analyzing the prosody splicing points of the warning text information, and generating prosody analysis data based on the prosody splicing points in combination with the prosody sequence coding; adjusting the initial voice based on the target linear spectrum and prosody analysis data in combination with the voice splicing smoothing algorithm and volume adjustment factor to obtain the target voice information.
[0031] Specifically, based on the recommended driving information, warning text information is generated using a preset warning text template. The preset warning text template includes warning text for vehicles and warning text for pedestrians. The warning text for vehicles is like "ahead crosswalk, recommended driving speed is a km / h", and the warning text for pedestrians is like "a vehicle is coming soon, please pay attention to safety for pedestrians". The recommended driving information is filled into the warning text for vehicles, and together with the warning text for pedestrians, warning text information is generated. Determine the semantic information corresponding to the warning text information, and determine the semantic information corresponding to the warning text information through natural language processing technology. Based on a convolutional encoder, determine the speech token corresponding to the warning text information using the semantic information. The speech token includes the encoded information corresponding to the semantic information, and the encoded information is used to identify important features of the semantic information, such as emotional expression and speech rate and tone, etc. Encode the input semantic information through the convolutional encoder, convert it into a hidden representation suitable for subsequent quantization processing, calculate the distance between the hidden representation and each code vector in the codebook, which can be calculated through the Euclidean distance, and take the code vector with the smallest distance as the speech token that can represent the current semantic information. And based on a large language model, determine the initial speech using the speech token. Generate the initial speech through the vocoder in the large language model using the speech token, and obtain the corresponding initial speech. Perform linear spectral analysis on several frames corresponding to the warning text information. The frames of the warning text information can be the frames corresponding to each phoneme of the warning text information. Obtain the vector features and prosodic features of the phoneme corresponding to each frame, and input the vector features and prosodic features into a pre-trained neural network model to obtain the target linear spectrum. Analyze the prosodic splicing points of the warning text information, divide the prosodic structure of the warning text information, and obtain the corresponding target prosodic structure. The prosodic structure includes prosodic words, prosodic phrases, and intonation phrases. Obtain historical warning text and historical warning speech, extract the text prosodic features of the historical warning text and the speech prosodic features of the historical warning speech, perform dimensionality conversion on the speech prosodic features according to the dimension of the text prosodic features, and obtain the speech prosodic features after dimensionality conversion. Obtain the target matrix of the text prosodic features and the speech prosodic features after dimensionality conversion, perform singular value decomposition on the target matrix to obtain the feature mapping matrix, and fuse the text prosodic features and the speech prosodic features after dimensionality conversion according to the feature mapping matrix to obtain the fused prosodic features. Determine the prosodic splicing points of the warning text information according to the fused prosodic features and the target prosodic structure using speech splicing cost analysis. The speech splicing cost analysis is implemented using a preset splicing cost analysis function, which can improve the accuracy of the speech prosodic splicing points. Generate prosodic analysis data based on the prosodic splicing points combined with prosodic sequence encoding. Form the embedding information of prosodic words and prosodic phrases according to the speech prosodic splicing points, generate prosodic sequence encoding using the attention mechanism according to the embedding information of prosodic words and prosodic phrases, and generate prosodic analysis data from the prosodic splicing points and the prosodic sequence encoding.Based on the target linear spectrum and prosody analysis data, the initial speech is adjusted by combining a speech splicing smoothing algorithm and a volume adjustment factor. The phase spectrum of the warning text information is determined according to the target linear spectrum, and the initial speech is preliminarily adjusted according to the phase spectrum. The prosody of the initially adjusted speech is adjusted according to the prosody analysis data. The speech after prosody adjustment is smoothed by the speech splicing smoothing algorithm. The volume of the speech is determined according to the volume adjustment factor, and finally the target speech information is obtained, making the obtained target speech smoother and more fluent, and improving the naturalness and expressiveness of the speech. The radar warning device and the pedestrian warning device perform traffic warnings based on the target speech information, that is, the radar warning device and the pedestrian warning device emit corresponding target speech information for traffic warnings. The applications of the radar warning device and the pedestrian warning device are as follows. Figure 5 As shown, the radar warning device notifies the driving vehicle to pay attention to the crosswalk and informs its recommended driving speed, etc. At the same time, the radar warning device sends a warning signal to the pedestrian warning device. The pedestrian warning device notifies the pedestrians on the crosswalk to cross the crosswalk as soon as possible, as there is a vehicle approaching. It can also display the warning text information and turn on the warning light. At the same time, the pedestrian warning device can trigger and control the spotlight to illuminate the zebra crossing when there are pedestrians at night. The radar warning device can trigger and control the yellow flashing light located on the roadside to warn vehicles, so as to ensure the traffic safety of pedestrians on the crosswalk at night.
[0032] In the embodiment of the present invention, vehicle driving analysis is performed based on multi-dimensional Fourier transform using the radar echo signals collected by the radar warning device in the radar detection area, making the obtained vehicle driving information more accurate and ensuring the reliability of traffic safety warnings. Based on occlusion analysis, face positioning and multi-layer filtering processing are performed on several frames of target images collected by the pedestrian warning device for the crosswalk to obtain a target face set for two-way pedestrian flow analysis. Considering the situation of face occlusion and introducing the filtering of the face set, it is avoided to miss detection or double counting, making the obtained pedestrian flow data more accurate. Based on the pedestrian passing time prediction data and vehicle driving information, the recommended driving information of the vehicle in the warning area is analyzed, which can improve the reliability of the recommended driving information analysis and at the same time improve the intelligence of traffic safety warnings, enabling the driver to understand the current required driving information. Based on the recommended driving information, speech generation is performed using speech tokens and prosody analysis to obtain target speech information for communication warnings, improving the accuracy and fluency of speech generation and being able to provide more reliable traffic warnings for vehicles and pedestrians.
[0033] Embodiment 2
[0034] Please refer to Figure 2 , Figure 2It is a schematic flowchart of a traffic safety warning method for a crosswalk in another embodiment of the present invention. The method is applied to a radar warning device and a pedestrian warning device, and the radar warning device is communicatively connected to the pedestrian warning device. The method includes: S201: Perform vehicle driving analysis on the radar echo signals collected by the radar warning device in the radar detection area based on multi-dimensional Fourier transform to obtain vehicle driving information; S202: Perform face positioning and multi-layer filtering processing on several frames of target images collected by the pedestrian warning device for the crosswalk based on occlusion analysis to obtain a set of target faces; S203: Perform two-way pedestrian flow analysis based on the set of target faces to obtain two-way pedestrian flow data; S204: Predict the pedestrian passing time of the crosswalk based on the two-way pedestrian flow data and several frames of target images to obtain pedestrian passing time prediction data; S205: Analyze the recommended driving information of the vehicle in the warning area based on the pedestrian passing time prediction data and the vehicle driving information; S206: Generate warning text information using a preset warning text template based on the recommended driving information, and determine the semantic information corresponding to the warning text information; S207: Determine the voice tokens corresponding to the warning text information based on the semantic information using a convolutional encoder, and determine the initial voice based on the voice tokens using a large language model; S208: Perform linear spectrum analysis on several frames corresponding to the warning text information to obtain the target linear spectrum; S209: Analyze the prosody splicing points of the warning text information, and generate prosody analysis data based on the prosody splicing points combined with the prosody sequence coding; S210: Adjust the initial voice based on the target linear spectrum and the prosody analysis data in combination with the voice splicing smoothing algorithm and the volume adjustment factor to obtain the target voice information. The radar warning device and the pedestrian warning device perform traffic warnings based on the target voice information.
[0035] In the embodiments of the present invention, vehicle driving analysis is performed based on multi-dimensional Fourier transform using radar echo signals collected by a radar warning device in a radar detection area, making the obtained vehicle driving information more accurate and ensuring the reliability of traffic safety warnings. Occlusion analysis is used to perform face localization and multi-layer filtering on several frames of target images collected by a pedestrian warning device for a crosswalk to obtain a set of target faces for two-way pedestrian flow analysis. Considering the situation of face occlusion and introducing filtering of the face set, false positives or double counting are avoided, making the obtained pedestrian flow data more accurate. Based on pedestrian passing time prediction data and vehicle driving information, recommended driving information of a vehicle in a warning area is analyzed, which can improve the reliability of recommended driving information analysis and the intelligence of traffic safety warnings, enabling a driver to understand the current required driving information. Based on the recommended driving information, voice generation is performed using voice tokens and prosody analysis to obtain target voice information for communication warnings, improving the accuracy and fluency of voice generation and providing more reliable traffic warnings for vehicles and pedestrians.
[0036] Embodiment III
[0037] Please refer to Figure 3 , Figure 3 , which is a schematic structural diagram of a traffic safety warning system for a crosswalk in the embodiments of the present invention. The system includes a radar warning device 31 and a pedestrian warning device 32, and the radar warning device 31 is communicatively connected to the pedestrian warning device 32; the system is configured to execute the traffic safety warning method for the crosswalk described in the above embodiments.
[0038] In the specific implementation process of the present invention, the applications of the radar warning device and the pedestrian warning device can be as Figure 5As shown, the radar warning device 31 is a warning device based on radar sensing. It includes a sensing radar, a horn, a red and blue flashing light, a light-emitting diode (LED) display screen, a unit control board, and an LED control board. The sensing radar senses the data of road vehicles in the radar detection area and transmits it to the unit control board for analysis to form a warning voice and a warning text. The warning voice is transmitted to the horn for playback, and the warning text is transmitted to the LED control board to control the LED display screen for display. The driver can then know the corresponding deceleration or stop operation of the vehicle in the warning area. The red and blue flashing light is used to give a warning light indication. At the same time, when passing at night, the radar warning device can trigger and control the yellow flashing light located on the roadside to warn vehicles. The pedestrian warning device 32 is a warning device based on a sensing camera. It includes a pedestrian sensing camera, a yellow flashing light, an LED display screen, a horn, a unit control board, and an LED control board. The pedestrian sensing camera is used to collect images of the crosswalk. The unit control board is used to analyze the images to generate a warning voice and a warning text. The warning voice is transmitted to the horn for broadcast, and the warning text is transmitted to the LED control board to control the LED display screen for display. The yellow flashing light is used to give a warning light indication. At the same time, the pedestrian warning device can trigger and control the spotlight, which can illuminate the zebra crossing when there are pedestrians at night. When pedestrians pass through the crosswalk at night, their passing safety can be maximally guaranteed. The radar warning device 31 and the pedestrian warning device 32 can also exchange data to achieve the trigger control of the warning voice and text. Figure 3 The system shown does not constitute a limitation on all components and may include more or fewer components than shown, or combine certain components.
[0039] In the specific implementation process of the present invention, the specific implementation manner of the system item can refer to the above embodiments and will not be elaborated here.
[0040] In the embodiments of the present invention, vehicle driving analysis is performed based on multi-dimensional Fourier transform using the radar echo signals collected by the radar warning device in the radar detection area, making the obtained vehicle driving information more accurate and ensuring the reliability of traffic safety warning. Based on occlusion analysis, face positioning and multi-layer filtering processing are performed on several frames of target images collected by the pedestrian warning device for the crosswalk to obtain a target face set for two-way pedestrian flow analysis. Considering the situation of face occlusion and introducing the filtering of the face set, false detections or duplicate counting are avoided, making the obtained pedestrian flow data more accurate. Based on the pedestrian passing time prediction data and vehicle driving information, the recommended driving information of the vehicle in the warning area is analyzed, which can improve the reliability of the recommended driving information analysis and the intelligence of traffic safety warning at the same time, enabling the driver to understand the current required driving information. Based on the recommended driving information, voice generation is performed using voice tokens and prosody analysis to obtain target voice information for communication warning, improving the accuracy and fluency of voice generation and providing more reliable traffic warnings for vehicles and pedestrians.
[0041] Embodiment 4
[0042] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of the traffic safety warning device for the crosswalk in the embodiments of the present invention. The device is applied to the radar warning device and the pedestrian warning device, and the radar warning device is communicatively connected to the pedestrian warning device; the device includes: Vehicle driving analysis module 41: configured to perform vehicle driving analysis based on multi-dimensional Fourier transform using the radar echo signals collected by the radar warning device in the radar detection area to obtain vehicle driving information; Face detection module 42: configured to perform face positioning and multi-layer filtering processing on several frames of target images collected by the pedestrian warning device for the crosswalk based on occlusion analysis to obtain a target face set; Pedestrian flow analysis module 43: configured to perform two-way pedestrian flow analysis based on the target face set to obtain two-way pedestrian flow data; Passing time prediction module 44: configured to predict the pedestrian passing time of the crosswalk based on the two-way pedestrian flow data and several frames of target images to obtain pedestrian passing time prediction data; Driving recommendation analysis module 45: configured to analyze the recommended driving information of the vehicle in the warning area based on the pedestrian passing time prediction data and vehicle driving information; Warning module 46: configured to perform voice generation based on the recommended driving information using voice tokens and prosody analysis to obtain target voice information, and the radar warning device and the pedestrian warning device perform traffic warning based on the target voice information.
[0043] In the specific implementation process of the present invention, the specific implementation manner of the device item can refer to the above embodiments, which will not be elaborated here.
[0044] In the embodiment of the present invention, based on the multi-dimensional Fourier transform, the radar echo signals collected by the radar warning device in the radar detection area are used for vehicle driving analysis, so that the obtained vehicle driving information is more accurate, and the reliability of the traffic safety warning is ensured. Based on the occlusion analysis, several frames of target images collected by the pedestrian warning device for the crosswalk are subjected to face positioning and multi-layer filtering processing to obtain a set of target faces for two-way pedestrian flow analysis. Considering the face occlusion situation, the filtering of the face set is introduced to avoid missed detection or double counting, so that the obtained pedestrian flow data is more accurate. Based on the pedestrian passing time prediction data and the vehicle driving information, the recommended driving information of the vehicle in the warning area is analyzed, which can improve the reliability of the recommended driving information analysis and at the same time improve the intelligence of the traffic safety warning, enabling the driver to understand the current required driving information. Based on the recommended driving information, voice generation is performed using voice tokens and prosody analysis to obtain target voice information for communication warning, improving the accuracy and fluency of voice generation, and being able to provide a more reliable traffic warning for vehicles and pedestrians.
[0045] A computer-readable storage medium provided by an embodiment of the present invention, on which a computer program is stored, and when the program is executed by a processor, it implements the traffic safety warning method for a crosswalk in any one of the above embodiments. Among them, the computer-readable storage medium includes, but is not limited to, any type of disk (including floppy disks, hard disks, optical disks, CD-ROMs, and magneto-optical disks), ROM (Read-Only Memory), RAM (Random Access Memory), EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), flash memory, magnetic cards or optical cards. That is, the storage device includes any medium that can store or transmit information in a readable form by a device (such as a computer, mobile phone), and can be a read-only memory, a magnetic disk or an optical disk, etc.
[0046] In addition, the above has introduced in detail a traffic safety warning method and related devices for a crosswalk provided by the embodiments of the present invention. Specific examples should have been used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for warning of pedestrian crossing safety, characterized in that, Applied to a radar warning device and a pedestrian warning device, the radar warning device is communicatively connected to the pedestrian warning device; the method includes: Based on multi-dimensional Fourier transform, vehicle driving analysis is performed using the radar echo signals collected by the radar warning device in the radar detection area to obtain vehicle driving information; Based on occlusion analysis, face positioning and multi-layer filtering processing are performed on a number of frame target images collected by the pedestrian warning device for the crosswalk to obtain a target face set; Based on the target face set, two-way pedestrian flow analysis is performed to obtain two-way pedestrian flow data; Based on the two-way pedestrian flow data and a number of frame target images, prediction of the pedestrian passing time of the crosswalk is performed to obtain pedestrian passing time prediction data; Based on the pedestrian passing time prediction data and the vehicle driving information, recommended driving information of the vehicle in the warning area is analyzed; Based on the recommended driving information, voice generation is performed using voice tokens and prosody analysis to obtain target voice information, and the radar warning device and the pedestrian warning device perform passing warnings based on the target voice information.
2. The method for warning of pedestrian crossing safety according to claim 1, characterized in that, The vehicle driving analysis using the radar echo signals collected by the radar warning device in the radar detection area based on multi-dimensional Fourier transform to obtain vehicle driving information includes: Performing filtering and noise reduction processing on the radar echo signals to obtain the radar echo signals after filtering and noise reduction processing; Performing a fast Fourier transform on the radar echo signals after filtering and noise reduction processing to obtain a Fourier spectrum, and based on the Fourier spectrum, analyzing the first vehicle speed, the first vehicle relative distance, and the first vehicle azimuth using an echo signal matrix; Performing ordered statistic constant false alarm detection on the radar echo signals after filtering and noise reduction processing to obtain an ordered statistic constant false alarm detection result; Based on the ordered statistic constant false alarm detection result, analyzing the second vehicle speed, the second vehicle relative distance, and the second vehicle azimuth using peak spectral lines, and generating a target vehicle speed, a target vehicle relative distance, and a target vehicle azimuth based on the first vehicle speed, the first vehicle relative distance, the first vehicle azimuth, the second vehicle speed, the second vehicle relative distance, and the second vehicle azimuth; Generating vehicle driving information based on the target vehicle speed, the target vehicle relative distance, and the target vehicle azimuth.
3. The method for warning of pedestrian crossing safety according to claim 1, characterized in that, The occlusion analysis using the pedestrian warning device to perform face positioning and multi-layer filtering processing on a number of frame target images collected by the pedestrian warning device for the crosswalk to obtain a target face set includes: Performing passing direction marking on a number of frame target images to obtain a number of frame target images after passing direction marking; Performing occlusion marking processing on the first face set obtained by positioning the number of frame target images after passing direction marking based on a preset occlusion type table to obtain the first face set after occlusion marking processing, and generating a face feature descriptor based on the first face set after occlusion marking processing; Determining a second face set in the number of frame target images based on the face feature descriptor using a face positioning model; Filtering the second face set based on a face tracking algorithm to obtain a third face set, and filtering the third face set based on a feature clustering algorithm to obtain a target face set corresponding to the passing direction.
4. The method for warning of pedestrian crossing safety according to claim 1, characterized in that, Performing two-way pedestrian flow analysis based on the target face set to obtain two-way pedestrian flow data, including: Performing two-way pedestrian flow analysis based on the boolean parameter and the detection box using the target faces marked with each passing direction in the target face set to obtain two-way pedestrian flow data.
5. The method for warning of pedestrian crossing safety according to claim 1, characterized in that, Predicting the pedestrian passing time of the crosswalk based on the two-way pedestrian flow data and several frames of target images to obtain pedestrian passing time prediction data, including: Converting the pixel coordinates of each target face located in each frame of the target image marked with the passing direction to obtain the corresponding position coordinates; Analyzing the corresponding pedestrian passing speed based on the time stamps of each frame of the target image using the corresponding position coordinates; Predicting the pedestrian passing time of the crosswalk based on the two-way pedestrian flow data and the corresponding pedestrian passing speed to obtain pedestrian passing time prediction data.
6. The method for warning of pedestrian crossing safety according to claim 1, characterized in that, Analyzing the recommended driving information of the vehicle in the warning area based on the pedestrian passing time prediction data and the vehicle driving information, including: Analyzing the traffic flow density based on the vehicle information in the warning area and analyzing the vehicle deceleration based on the pedestrian passing time prediction data; Calculating the safety distance based on the vehicle driving information combined with the driver's reaction time, and analyzing the recommended driving information of the vehicle in the warning area based on the traffic flow density, vehicle deceleration and safety distance.
7. The method for warning of pedestrian crossing safety according to claim 1, characterized in that, Generating target voice information using voice token and prosody analysis based on the recommended driving information, including: Generating warning text information using a preset warning text template based on the recommended driving information, and determining the semantic information corresponding to the warning text information; Determining the voice token corresponding to the warning text information based on the semantic information using a convolutional encoder, and determining the initial voice based on the voice token using a large language model; Performing linear spectrum analysis on several frames corresponding to the warning text information to obtain the target linear spectrum; Analyzing the prosody splicing points of the warning text information, and generating prosody analysis data based on the prosody splicing points combined with the prosody sequence coding; Adjusting the initial voice based on the target linear spectrum and prosody analysis data combined with the voice splicing smoothing algorithm and the volume adjustment factor to obtain the target voice information.
8. A traffic safety warning device for a crosswalk, characterized in that, Applied to a radar warning device and a pedestrian warning device, the radar warning device is communicatively connected to the pedestrian warning device; the device includes: A vehicle driving analysis module: configured to perform vehicle driving analysis based on multi-dimensional Fourier transform using the radar echo signals collected by the radar warning device in the radar detection area to obtain vehicle driving information; A face detection module: configured to perform face positioning and multi-layer filtering processing on several frames of target images collected by the pedestrian warning device for the crosswalk based on occlusion analysis to obtain a target face set; A pedestrian flow analysis module: configured to perform two-way pedestrian flow analysis based on the target face set to obtain two-way pedestrian flow data; A passing time prediction module: configured to predict the pedestrian passing time of the crosswalk based on the two-way pedestrian flow data and several frames of target images to obtain pedestrian passing time prediction data; A driving recommendation analysis module: configured to analyze the recommended driving information of the vehicle in the warning area based on the pedestrian passing time prediction data and the vehicle driving information; Warning module: configured to generate speech based on the recommended driving information using voice tokens and prosody analysis to obtain target speech information, and the radar warning device and pedestrian warning device perform traffic warnings based on the target speech information.
9. A traffic safety warning system for a crosswalk, characterized in that, The system includes a radar warning device and a pedestrian warning device, the radar warning device is communicatively connected to the pedestrian warning device, and the system is configured to execute the traffic safety warning method for a crosswalk according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the computer instructions are run on an electronic device, the electronic device is caused to execute the traffic safety warning method for a crosswalk according to any one of claims 1 to 7.
Citation Information
Patent Citations
Speeding vehicle lane detection method
CN103745601A
Intersection safety early warning system based on machine vision and DSRC and early warning method thereof
CN111210662A
Bidirectional pedestrian volume statistical method based on RGB-D multi-modal data
CN111881749A
Traffic flow detection method and device
CN112233416A
Method and device for converting text into voice, equipment and storage medium
CN119517004A
Cited By
Audio synthesis method and device, medium and equipment
CN121415759A