A traffic safety warning method and related device for a crosswalk
Through the combination of radar and pedestrian early warning equipment, multi-dimensional Fourier transform and occlusion analysis are used to solve the problem of large information error in video analysis, more accurate traffic statistics and voice generation are achieved, and the reliability and intelligence of traffic safety warning are improved.
Patent Information
- Application Number
- CN202510668120.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-05-23
AI Technical Summary
In the prior art, video analysis is prone to video frame loss in crosswalk traffic safety warning, resulting in large errors in vehicle driving information, inaccurate traffic statistics, insufficient accuracy and fluency of voice conversion, and unable to provide reliable traffic warning.
The communication connection between radar early warning equipment and pedestrian early warning equipment is used to analyze radar echo signals through multi-dimensional Fourier transform to obtain vehicle driving information, combine occlusion analysis and multi-layer filtering to process face positioning, perform two-way traffic analysis and pedestrian pass time prediction, and generate target voice information for traffic warning.
It improves the accuracy of vehicle driving information and the reliability of traffic data, enhances the intelligence of traffic safety warnings and the accuracy and fluency of voice generation, and provides a more reliable traffic warning.
Smart Images

Figure CN120183169B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent transportation, and in particular, to a method and related device for warning of the passing safety of a crosswalk. Background Art
[0002] When a vehicle passes through a crosswalk without traffic lights, due to the driver's failure to take timely deceleration measures or pedestrians' failure to timely know that a vehicle is coming, the safety problems between vehicles and pedestrians are increasing. In this regard, how to effectively warn of the passing safety of a crosswalk has become the research focus. For the warning of the passing safety of a crosswalk, the analysis of vehicle driving information is an important link. At present, it is usually realized through video analysis, but video analysis is prone to video frame loss, resulting in large errors in the analyzed vehicle driving information and unable to guarantee the reliability of the passing safety warning. In order to improve the intelligence of the passing safety warning, the recommended driving information of the vehicle is usually analyzed. The statistical analysis of the pedestrian flow of the crosswalk is a key step in the analysis of the recommended driving information. At present, the analysis of the pedestrian flow mainly counts the number of faces appearing in the video frame images, but currently, less consideration is given to face positioning in the case where the face is partially blocked, resulting in incorrect face positioning and large deviations in the obtained pedestrian flow. After obtaining the recommended driving information, it needs to be converted into voice for passing warning. At present, it is mainly through an autoregressive model for voice conversion, but the accuracy of the voice conversion by this method is low, and the fluency is also insufficient, unable to provide a reliable passing warning. Summary of the Invention
[0003] The purpose of the present invention is to overcome the deficiencies of the prior art. The present invention provides a method and related device for warning of the passing safety of a crosswalk, which can accurately analyze the recommended driving information of a vehicle, so as to provide a more reliable passing warning for vehicles and pedestrians.
[0004] To solve the above technical problems, the present invention provides a method for warning of the passing safety of a crosswalk, which is applied to a radar warning device and a pedestrian warning device, and the radar warning device is communicatively connected to the pedestrian warning device; the method includes:
[0005] Performing vehicle driving analysis on the radar echo signal collected by the radar warning device in the radar detection area based on multi-dimensional Fourier transform to obtain vehicle driving information;
[0006] Performing face positioning and multi-layer filtering processing on a plurality of frame target images collected by the pedestrian warning device for the crosswalk based on occlusion analysis to obtain a target face set;
[0007] Performing two-way pedestrian flow analysis based on the target face set to obtain two-way pedestrian flow data;
[0008] Predict the pedestrian passing time of a crosswalk based on two-way pedestrian flow data and several frames of target images, and obtain pedestrian passing time prediction data;
[0009] Analyze the recommended driving information of vehicles in the warning area based on the pedestrian passing time prediction data and vehicle driving information;
[0010] Generate target voice information based on the recommended driving information by using voice tokens and prosody analysis, and the radar warning device and pedestrian warning device perform passing warnings based on the target voice information.
[0011] Optionally, the vehicle driving analysis is performed based on multi-dimensional Fourier transform by using the radar echo signals collected by the radar warning device in the radar detection area to obtain vehicle driving information, including:
[0012] Perform filtering and noise reduction processing on the radar echo signals to obtain the radar echo signals after filtering and noise reduction processing;
[0013] Perform a fast Fourier transform on the radar echo signals after filtering and noise reduction processing to obtain a Fourier spectrum, and analyze the first vehicle speed, the first vehicle relative distance, and the first vehicle azimuth based on the Fourier spectrum by using an echo signal matrix;
[0014] Perform ordered statistic constant false alarm detection based on the radar echo signals after filtering and noise reduction processing to obtain an ordered statistic constant false alarm detection result;
[0015] Analyze the second vehicle speed, the second vehicle relative distance, and the second vehicle azimuth based on the ordered statistic constant false alarm detection result by using peak spectral line analysis, and generate the target vehicle speed, the target vehicle relative distance, and the target vehicle azimuth based on the first vehicle speed, the first vehicle relative distance, the first vehicle azimuth, the second vehicle speed, the second vehicle relative distance, and the second vehicle azimuth;
[0016] Generate vehicle driving information based on the target vehicle speed, the target vehicle relative distance, and the target vehicle azimuth.
[0017] Optionally, the occlusion analysis uses the several frames of target images collected by the pedestrian warning device for the crosswalk to perform face localization and multi-layer filtering processing to obtain a target face set, including:
[0018] Mark the passing directions of several frames of target images to obtain several frames of target images with passing direction marks;
[0019] Perform occlusion marking processing on the first face set obtained by localizing several frames of target images with passing direction marks based on a preset occlusion type table to obtain the first face set after occlusion marking processing, and generate a face feature descriptor based on the first face set after occlusion marking processing;
[0020] Determine a second set of faces in several frames of target images using a face localization model based on the face feature descriptors;
[0021] Filter the second set of faces based on a face tracking algorithm to obtain a third set of faces, and filter the third set of faces based on a feature clustering algorithm to obtain a set of target faces corresponding to the passing direction.
[0022] Optionally, performing two-way pedestrian flow analysis based on the set of target faces to obtain two-way pedestrian flow data, including:
[0023] Perform two-way pedestrian flow analysis using the target faces marked with passing directions in the set of target faces based on boolean parameters and detection frames to obtain two-way pedestrian flow data.
[0024] Optionally, predicting the pedestrian passing time of a crosswalk based on the two-way pedestrian flow data and several frames of target images to obtain pedestrian passing time prediction data, including:
[0025] Convert the pixel coordinates of each target face located in each frame of the target image marked with the passing direction to obtain corresponding position coordinates;
[0026] Analyze the corresponding pedestrian passing speed based on the time stamps of each frame of the target image using the corresponding position coordinates;
[0027] Predict the pedestrian passing time of a crosswalk based on the two-way pedestrian flow data and the corresponding pedestrian passing speed to obtain pedestrian passing time prediction data.
[0028] Optionally, analyzing the recommended driving information of a vehicle in a warning area based on the pedestrian passing time prediction data and vehicle driving information, including:
[0029] Analyze the traffic flow density based on the vehicle information in the warning area, and analyze the vehicle deceleration based on the pedestrian passing time prediction data;
[0030] Calculate a safety distance based on the vehicle driving information combined with the driver's reaction time, and analyze the recommended driving information of the vehicle in the warning area based on the traffic flow density, vehicle deceleration, and safety distance.
[0031] Optionally, generating target voice information using voice tokens and prosody analysis based on the recommended driving information, including:
[0032] Generate warning text information based on the recommended driving information using a preset warning text template, and determine the semantic information corresponding to the warning text information;
[0033] Determine the speech tokens corresponding to the warning text information based on the convolutional encoder using semantic information, and determine the initial speech based on the large language model using the speech tokens;
[0034] Perform linear spectral analysis on several frames corresponding to the warning text information to obtain the target linear spectrum;
[0035] Analyze the prosodic splicing points of the warning text information, and generate prosodic analysis data based on the prosodic splicing points combined with the prosodic sequence encoding;
[0036] Adjust the initial speech based on the target linear spectrum and the prosodic analysis data by combining the speech splicing smoothing algorithm and the volume adjustment factor to obtain the target speech information.
[0037] In addition, the present invention also provides a pedestrian crossing traffic safety warning device, which is applied to a radar warning device and a pedestrian warning device, and the radar warning device is communicatively connected to the pedestrian warning device; the device includes:
[0038] A vehicle driving analysis module: used to perform vehicle driving analysis based on multi-dimensional Fourier transform using the radar echo signals collected by the radar warning device in the radar detection area to obtain vehicle driving information;
[0039] A face detection module: used to perform face positioning and multi-layer filtering processing on several frames of target images collected by the pedestrian warning device for the pedestrian crossing based on occlusion analysis to obtain a set of target faces;
[0040] A pedestrian flow analysis module: used to perform two-way pedestrian flow analysis based on the set of target faces to obtain two-way pedestrian flow data;
[0041] A passing time prediction module: used to predict the pedestrian passing time of the pedestrian crossing based on the two-way pedestrian flow data and several frames of target images to obtain pedestrian passing time prediction data;
[0042] A driving recommendation analysis module: used to analyze the recommended driving information of the vehicle in the warning area based on the pedestrian passing time prediction data and the vehicle driving information;
[0043] A warning module: used to generate speech based on the recommended driving information using speech tokens and prosodic analysis to obtain target speech information, and the radar warning device and the pedestrian warning device perform traffic warnings based on the target speech information.
[0044] In addition, the present invention also provides a pedestrian crossing traffic safety warning system, the system includes a radar warning device and a pedestrian warning device, the radar warning device is communicatively connected to the pedestrian warning device, and the system is configured to execute the above-mentioned pedestrian crossing traffic safety warning method.
[0045] In addition, the present invention also provides a computer-readable storage medium storing computer instructions, which, when running on an electronic device, cause the electronic device to execute the above-mentioned traffic safety warning method for a crosswalk.
[0046] In the embodiments of the present invention, vehicle driving analysis is performed on the radar echo signals collected by the radar warning device in the radar detection area based on multi-dimensional Fourier transform, so that the obtained vehicle driving information is more accurate, ensuring the reliability of traffic safety warning. Based on occlusion analysis, face localization and multi-layer filtering processing are performed on several frames of target images collected by the pedestrian warning device for the crosswalk to obtain a set of target faces for two-way pedestrian flow analysis. Considering the situation of face occlusion and introducing the filtering of the face set, it is avoided to miss detection or double counting, making the obtained pedestrian flow data more accurate. Based on the predicted pedestrian passing time data and vehicle driving information, the recommended driving information of the vehicle in the warning area is analyzed, which can improve the reliability of the recommended driving information analysis and at the same time improve the intelligence of traffic safety warning, enabling the driver to understand the current required driving information. Based on the recommended driving information, voice generation is performed using voice tokens and prosody analysis to obtain target voice information for communication warning, improving the accuracy and fluency of voice generation, and being able to provide more reliable traffic warnings for vehicles and pedestrians. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0048] Figure 1 is a schematic flowchart of the traffic safety warning method for a crosswalk in the embodiments of the present invention;
[0049] Figure 2 is a schematic flowchart of the traffic safety warning method for a crosswalk in another embodiment of the present invention;
[0050] Figure 3 is a schematic diagram of the structural composition of the traffic safety warning system for a crosswalk in the embodiments of the present invention;
[0051] Figure 4 is a schematic diagram of the structural composition of the traffic safety warning device for a crosswalk in the embodiments of the present invention;
[0052] Figure 5 is an application diagram of the radar warning device and the pedestrian warning device in the embodiments of the present invention. Detailed implementation manners
[0053] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0054] Embodiment 1
[0055] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a traffic safety warning method for a crosswalk in an embodiment of the present invention. The method is applied to a radar warning device and a pedestrian warning device, and the radar warning device is communicatively connected to the pedestrian warning device; the method includes:
[0056] S11: Based on multi-dimensional Fourier transform, use the radar echo signal collected by the radar warning device in the radar detection area to perform vehicle driving analysis to obtain vehicle driving information;
[0057] In the specific implementation process of the present invention, the step of using the radar echo signal collected by the radar warning device in the radar detection area to perform vehicle driving analysis based on multi-dimensional Fourier transform to obtain vehicle driving information includes: performing filter denoising processing on the radar echo signal to obtain the radar echo signal after filter denoising processing; performing fast Fourier transform on the radar echo signal after filter denoising processing to obtain a Fourier spectrum, and based on the Fourier spectrum, use the echo signal matrix to analyze the first vehicle speed, the first vehicle relative distance, and the first vehicle azimuth; performing ordered-statistic constant false alarm detection on the radar echo signal after filter denoising processing to obtain an ordered-statistic constant false alarm detection result; based on the ordered-statistic constant false alarm detection result, use the peak spectral line to analyze the second vehicle speed, the second vehicle relative distance, and the second vehicle azimuth, and generate the target vehicle speed, the target vehicle relative distance, and the target vehicle azimuth based on the first vehicle speed, the first vehicle relative distance, the first vehicle azimuth, the second vehicle speed, the second vehicle relative distance, and the second vehicle azimuth. Generate vehicle driving information based on the target vehicle speed, the target vehicle relative distance, and the target vehicle azimuth.
[0058] Specifically, the radar warning device emits pulse signals to the radar detection area. When the pulse signals encounter a moving vehicle, radar echo signals will be generated, and the radar warning device collects the radar echo signals. Filtering and noise reduction processing is performed on the radar echo signals. The radar echo signals are filtered through a high-pass filter, and noise reduction processing is performed on the filtered radar echo signals. The noise reduction processing can adopt algorithms such as wavelet transform to obtain the radar echo signals after filtering and noise reduction processing. Fast Fourier transform is performed on the radar echo signals after filtering and noise reduction processing. By transforming the radar echo signals from the time domain to the frequency domain, the spectral structure of the signals is obtained, the Fourier spectrum is obtained, and based on the Fourier spectrum, the first vehicle speed, the first vehicle relative distance, and the first vehicle azimuth are analyzed using the echo signal matrix. The peak in the Fourier spectrum is extracted, and the peak position is the position of the moving vehicle. The first vehicle relative distance is calculated according to the frequency corresponding to the peak position, that is, the distance between the vehicle and the radar. The echo signal matrix is a signal matrix formed by the radar warning device sending the same electromagnetic wave signals to the detection area at preset time intervals and collecting the corresponding echo signals. The peak of the Fourier spectrum in the echo signal matrix is extracted, and the first vehicle speed is generated according to the frequency corresponding to the peak of the Fourier spectrum in the echo signal matrix combined with the frequency modulation slope. The relative angle between the first vehicle and the radar is obtained according to the wavelength of the radar echo signals after filtering and noise reduction processing and the peak frequency of the Fourier spectrum, and the first vehicle azimuth is obtained from the relative angle and the vehicle position. Ordered statistical constant false alarm detection is performed based on the radar echo signals after filtering and noise reduction processing. The radar echo signals after filtering and noise reduction processing may still contain non-target signals. After corresponding signal processing on the radar echo signals, the target is obtained with a relatively large statistical probability, while noise and other interference signals generate false alarms with a relatively low probability, so as to extract the target from the mixed signals. Its purpose is to extract the target from the input signals according to certain criteria, which is the purpose of ordered statistical constant false alarm detection. The corresponding power sampling values are obtained through ordered statistical constant false alarm detection. By comparing the power sampling values with the preset threshold value, the target signal is extracted according to the comparison result, and the ordered statistical constant false alarm detection result is obtained. Based on the ordered statistical constant false alarm detection result, the second vehicle speed, the second vehicle relative distance, and the second vehicle azimuth are analyzed using peak spectral line analysis. The second vehicle relative distance is calculated through the peak spectral line of the spectrogram in the target signal extracted by the ordered statistical constant false alarm detection result. The second vehicle speed is calculated according to the second vehicle relative distance combined with the transmission frequency and the frequency modulation slope of the radar signal. The second vehicle azimuth is calculated according to the radar wavelength and the wave path difference, and the target vehicle speed, the target vehicle relative distance, and the target vehicle azimuth are generated based on the first vehicle speed, the first vehicle relative distance, the first vehicle azimuth, the second vehicle speed, the second vehicle relative distance, and the second vehicle azimuth. The target vehicle speed is obtained by weighted averaging the first vehicle speed and the second vehicle speed. Similarly, the target vehicle relative distance and the target vehicle azimuth are also obtained by weighted averaging.Generate vehicle driving information based on the target vehicle speed, the relative distance of the target vehicle, and the orientation of the target vehicle, making the obtained vehicle driving information more accurate.
[0059] S12: Based on occlusion analysis, use a pedestrian warning device to perform face localization and multi-layer filtering on several frames of target images collected from a crosswalk to obtain a set of target faces;
[0060] In the specific implementation process of the present invention, the step of using a pedestrian warning device to perform face localization and multi-layer filtering on several frames of target images collected from a crosswalk based on occlusion analysis to obtain a set of target faces includes: marking the passing directions of several frames of target images to obtain several frames of target images with passing direction marks; performing occlusion marking processing on the first face set obtained by localizing several frames of target images with passing direction marks based on a preset occlusion type table to obtain a first face set after occlusion marking processing, and generating a face feature descriptor based on the first face set after occlusion marking processing; determining a second face set in several frames of target images based on the face feature descriptor using a face localization model; filtering the second face set based on a face tracking algorithm to obtain a third face set, and filtering the third face set based on a feature clustering algorithm to obtain a set of target faces corresponding to the passing direction.
[0061] Specifically, pedestrian warning devices on both sides of the crosswalk collect images of the traffic conditions of the crosswalk to obtain a number of target images. Mark the traffic directions for the number of target images. The traffic direction of the crosswalk is two-way. The pedestrian warning devices on both sides of the crosswalk mark the corresponding traffic directions for the collected target images, that is, mark which traffic direction the frame of target image belongs to according to the side where the pedestrian warning device is located, and obtain a number of target images after the traffic direction marking. Perform occlusion marking processing on the first face set obtained by positioning the number of target images marked with traffic directions based on a preset occlusion type table. Use a deep convolutional neural network to perform face positioning on the number of target images marked with traffic directions to obtain the corresponding first face set. Analyze the occlusion types of the face images in the first face set through the preset occlusion type table. The occlusion types include mouth occlusion, nose occlusion, left and right eye occlusion, etc. Analyze the occlusion types of the face images with occlusions in the first face set, and mark the face images with occlusions in the first face set according to the analyzed occlusion types to obtain the first face set after occlusion marking processing. Generate a face feature descriptor based on the first face set after occlusion marking processing. Extract face key points from the first face set after occlusion marking processing. Construct an octagon neighborhood centered on the face key points. Use the face key points to construct a number of spatial triangles in the octagon neighborhood. Construct a number of histograms based on each spatial triangle. Connect each histogram in the form of a vector to form a triangular feature descriptor, that is, the face feature descriptor. Both occluded face images and complete face images have corresponding face feature descriptors. Use the face positioning model to determine the second face set in the number of target images based on the face feature descriptor. Locate the face set in the number of target images according to the face feature descriptor using the face positioning model, match the face set with the first face set, and check whether there are missed detections or misdetections. If there is a missed detection, add the missed face image to the first face set. If there is a misdetection, delete the misdetected face image. Perform the above processing on the face set of each frame of target image to form the second face set.Filter the second face set based on the face tracking algorithm, perform face tracking on the face set of each frame of the target image, delete the repeatedly appearing faces, and only retain the initially appearing faces. For the occluded faces, match their unoccluded face features to obtain the third face set, and filter the third face set based on the feature clustering algorithm. The filtering of the second face set is only a simple filtering, and there may still be duplicate faces. To ensure that no duplicate faces appear to affect the pedestrian flow analysis, it is necessary to further filter the third face set. Extract the face feature values of each face in the third face set, and determine whether the face feature value belongs to the face feature value set of the clustering filter. If it belongs to the face feature value set in the clustering filter, filter out the face from the third face set. The face feature value set of the clustering filter refers to the face feature values of all the previous target faces saved when judging whether the nth target face is similar to the previous target faces. The preset face feature value is an effective parameter value used to filter out the subsequent target faces with the same face feature value, filter out the duplicate faces, and obtain the final face set. Since the target image has been marked with the passing direction at the beginning, the obtained face set also has the corresponding passing direction mark, that is, obtain the target face set corresponding to the passing direction.
[0062] S13: Perform two-way pedestrian flow analysis based on the target face set to obtain two-way pedestrian flow data;
[0063] In the specific implementation process of the present invention, the performing two-way pedestrian flow analysis based on the target face set to obtain two-way pedestrian flow data includes: performing two-way pedestrian flow analysis based on the Boolean parameter and the detection frame using the target faces marked with the passing direction in the target face set to obtain two-way pedestrian flow data.
[0064] Specifically, perform two-way pedestrian flow analysis based on the Boolean parameter and the detection frame using the target faces marked with the passing direction in the target face set. Frame the detection frames for the target faces marked with each communication direction, and assign a Boolean parameter to each detection frame, initialized to zero. When the detection frame is counted, its value changes from zero to one, which is used to record whether the detection frame has been counted, so as to avoid double counting. Count the number of detection frames to obtain the pedestrian flow data in different passing directions, that is, obtain the two-way pedestrian flow data.
[0065] S14: Predict the pedestrian passing time of the crosswalk based on the two-way pedestrian flow data and several frames of target images to obtain pedestrian passing time prediction data;
[0066] In the specific implementation process of the present invention, predicting the pedestrian passing time of a crosswalk based on two-way pedestrian flow data and several frames of target images to obtain pedestrian passing time prediction data includes: converting the pixel coordinates of each target face located in each frame of the target image after marking the passing direction to obtain corresponding position coordinates; analyzing the corresponding pedestrian passing speed based on the time stamps of each frame of the target image using the corresponding position coordinates; predicting the pedestrian passing time of the crosswalk based on the two-way pedestrian flow data and the corresponding pedestrian passing speed to obtain pedestrian passing time prediction data.
[0067] Specifically, converting the pixel coordinates of each target face located in each frame of the target image after marking the passing direction, that is, converting the pixel coordinates to world coordinates in the world coordinate system to obtain corresponding position coordinates. Analyzing the corresponding pedestrian passing speed based on the time stamps of each frame of the target image using the corresponding position coordinates. The same face in each frame of the target image has been known in face localization. According to the change in the position coordinates of the same face in each frame of the image and combining the time stamp of each frame of the target image, the corresponding pedestrian passing speed can be calculated. Predicting the pedestrian passing time of the crosswalk based on the two-way pedestrian flow data and the corresponding pedestrian passing speed, inputting the two-way pedestrian flow data and the corresponding pedestrian passing speed into the passing time prediction model to obtain the predicted time of pedestrian passing, that is, obtaining pedestrian passing time prediction data. Since the pedestrian flow in different passing directions will affect the pedestrian passing speed, the two-way pedestrian flow data is added to predict the passing time, making the obtained prediction data more in line with the actual situation.
[0068] S15: Analyzing the recommended driving information of the vehicle in the warning area based on the pedestrian passing time prediction data and the vehicle driving information;
[0069] In the specific implementation process of the present invention, analyzing the recommended driving information of the vehicle in the warning area based on the pedestrian passing time prediction data and the vehicle driving information includes: analyzing the traffic flow density based on the vehicle information in the warning area and analyzing the vehicle deceleration based on the pedestrian passing time prediction data; calculating the safety distance based on the vehicle driving information combined with the driver reaction time, and analyzing the recommended driving information of the vehicle in the warning area based on the traffic flow density, vehicle deceleration and safety distance.
[0070] Specifically, the traffic flow density is analyzed based on the vehicle information in the warning area. The vehicle information in the warning area includes the length of the warning area, the number and speed of vehicles in each lane within the warning area, etc. The traffic flow density of each lane can be analyzed according to the number of vehicles in the warning area and the length of the warning area. The vehicle deceleration is predicted based on the pedestrian passing time prediction data, and the vehicle deceleration is analyzed by combining the pedestrian passing time prediction data and the traffic flow density with the vehicle speed in the vehicle driving information. The safe distance is calculated based on the vehicle driving information in combination with the driver's reaction time, that is, the safe distance from the vehicle in front after entering the warning area is determined according to the vehicle speed and vehicle orientation in the vehicle driving information in combination with the driver's reaction time. The recommended driving information of the vehicle in the warning area is analyzed based on the traffic flow density, vehicle deceleration, and safe distance. The lanes that the vehicle can change and the recommended driving speed in the warning area are analyzed according to the traffic flow density, vehicle deceleration, and safe distance. The recommended driving information is generated from the vehicle deceleration, safe distance, lanes that can be changed, and recommended driving speed.
[0071] S16: Based on the recommended driving information, voice generation is performed using voice tokens and prosody analysis to obtain target voice information. The radar warning device and the pedestrian warning device perform traffic warning based on the target voice information.
[0072] In the specific implementation process of the present invention, the voice generation using voice tokens and prosody analysis based on the recommended driving information to obtain target voice information includes: generating warning text information using a preset warning text template based on the recommended driving information, and determining the semantic information corresponding to the warning text information; determining the voice tokens corresponding to the warning text information based on the semantic information using a convolutional encoder, and determining the initial voice based on the voice tokens using a large language model; performing linear spectral analysis on several frames corresponding to the warning text information to obtain target linear spectra; analyzing the prosody splicing points of the warning text information, and generating prosody analysis data based on the prosody splicing points in combination with prosody sequence coding; adjusting the initial voice based on the target linear spectra and prosody analysis data in combination with a voice splicing smoothing algorithm and a volume adjustment factor to obtain target voice information.
[0073] Specifically, based on the recommended driving information, warning text information is generated using a preset warning text template. The preset warning text template includes warning text for vehicles and warning text for pedestrians. The warning text for vehicles is like "ahead crosswalk, recommended driving speed is a km / h", and the warning text for pedestrians is like "a vehicle is coming, please pay attention to safety". The recommended driving information is filled into the warning text for vehicles, and together with the warning text for pedestrians, warning text information is generated. The semantic information corresponding to the warning text information is determined, and the semantic information corresponding to the warning text information is determined through natural language processing technology. Based on a convolutional encoder, the speech token corresponding to the warning text information is determined using the semantic information. The speech token includes the encoded information corresponding to the semantic information, and the encoded information is used to identify important features of the semantic information, such as emotional expression and speech rate and pitch. The input semantic information is encoded through the convolutional encoder and converted into a hidden representation suitable for subsequent quantization processing. The distance between the hidden representation and each code vector in the codebook is calculated, which can be calculated through the Euclidean distance. The code vector with the smallest distance is used as the speech token that can represent the current semantic information, and based on a large language model, the initial speech is determined using the speech token. The initial speech is generated using the speech token through the vocoder in the large language model to obtain the corresponding initial speech. Linear spectral analysis is performed on several frames corresponding to the warning text information. The frames of the warning text information can be the frames corresponding to each phoneme of the warning text information. The vector features and prosodic features of each frame corresponding to the phoneme are obtained, and the vector features and prosodic features are input into a pre-trained neural network model to obtain the target linear spectrum. The prosodic splicing points of the warning text information are analyzed, and the prosodic structure of the warning text information is divided to obtain the corresponding target prosodic structure. The prosodic structure includes prosodic words, prosodic phrases, and intonation phrases. The historical warning text and historical warning speech are obtained, and the text prosodic features of the historical warning text and the speech prosodic features of the historical warning speech are extracted. According to the dimension of the text prosodic features, the dimension of the speech prosodic features is transformed to obtain the transformed speech prosodic features. The target matrix of the text prosodic features and the transformed speech prosodic features is obtained, and the target matrix is subjected to singular value decomposition to obtain the feature mapping matrix. According to the feature mapping matrix, the text prosodic features and the transformed speech prosodic features are fused to obtain the fused prosodic features. According to the fused prosodic features and the target prosodic structure, the prosodic splicing points of the warning text information are determined using speech splicing cost analysis. The speech splicing cost analysis is implemented using a preset splicing cost analysis function, which can improve the accuracy of the speech prosodic splicing points. Based on the prosodic splicing points, prosodic analysis data is generated in combination with prosodic sequence encoding. According to the speech prosodic splicing points, the embedding information of prosodic words and prosodic phrases is formed. According to the embedding information of prosodic words and prosodic phrases, the prosodic sequence encoding is generated using the attention mechanism, and the prosodic analysis data is generated from the prosodic splicing points and the prosodic sequence encoding.Based on the target linear spectrum and prosody analysis data, the initial speech is adjusted by combining a speech splicing smoothing algorithm and a volume adjustment factor. The phase spectrum of the warning text information is determined according to the target linear spectrum, and the initial speech is preliminarily adjusted according to the phase spectrum. The prosody of the initially adjusted speech is adjusted according to the prosody analysis data. The initially adjusted speech after prosody adjustment is smoothed by a speech splicing smoothing algorithm. The volume of the speech is determined according to the volume adjustment factor, and finally the target speech information is obtained, making the obtained target speech smoother and more fluent, and improving the naturalness and expressiveness of the speech. The radar warning device and the pedestrian warning device perform traffic warnings based on the target speech information, that is, the radar warning device and the pedestrian warning device issue corresponding target speech information for traffic warnings. The applications of the radar warning device and the pedestrian warning device are as follows. Figure 5 As shown, the radar warning device notifies the driving vehicle to pay attention to the crosswalk and informs it of the recommended driving speed, etc. At the same time, the radar warning device sends a warning signal to the pedestrian warning device. The pedestrian warning device notifies the pedestrians on the crosswalk to cross the crosswalk as soon as possible, as there is a vehicle approaching. It can also display the warning text information and turn on the warning light. At the same time, the pedestrian warning device can trigger the control of the spotlight to illuminate the zebra crossing when there are pedestrians at night. The radar warning device can trigger the control of the yellow flashing light located on the roadside to warn vehicles, so as to ensure the traffic safety of pedestrians on the crosswalk at night.
[0074] In the embodiment of the present invention, vehicle driving analysis is performed based on multi-dimensional Fourier transform using the radar echo signals collected by the radar warning device in the radar detection area, making the obtained vehicle driving information more accurate and ensuring the reliability of traffic safety warnings. Based on occlusion analysis, face positioning and multi-layer filtering processing are performed on several frames of target images collected by the pedestrian warning device for the crosswalk to obtain a target face set for two-way pedestrian flow analysis. Considering the situation of face occlusion and introducing the filtering of the face set, it is possible to avoid missed detections or duplicate counting, making the obtained pedestrian flow data more accurate. Based on the pedestrian passing time prediction data and vehicle driving information, the recommended driving information of the vehicle in the warning area is analyzed, which can improve the reliability of the recommended driving information analysis and at the same time improve the intelligence of traffic safety warnings, enabling the driver to understand the current required driving information. Based on the recommended driving information, speech generation is performed using voice tokens and prosody analysis to obtain target speech information for communication warnings, improving the accuracy and fluency of speech generation and being able to provide more reliable traffic warnings for vehicles and pedestrians.
[0075] Embodiment 2
[0076] Please refer to Figure 2 , Figure 2It is a schematic flowchart of a traffic safety warning method for a crosswalk in another embodiment of the present invention. The method is applied to a radar warning device and a pedestrian warning device, and the radar warning device is communicatively connected to the pedestrian warning device. The method includes:
[0077] S201: Based on multi-dimensional Fourier transform, use the radar echo signals collected by the radar warning device in the radar detection area to perform vehicle driving analysis to obtain vehicle driving information;
[0078] S202: Based on occlusion analysis, use the pedestrian warning device to perform face localization and multi-layer filtering processing on several frames of target images collected for the crosswalk to obtain a set of target faces;
[0079] S203: Perform two-way pedestrian flow analysis based on the set of target faces to obtain two-way pedestrian flow data;
[0080] S204: Predict the pedestrian passing time of the crosswalk based on the two-way pedestrian flow data and several frames of target images to obtain pedestrian passing time prediction data;
[0081] S205: Analyze the recommended driving information of the vehicle in the warning area based on the pedestrian passing time prediction data and the vehicle driving information;
[0082] S206: Generate warning text information using a preset warning text template based on the recommended driving information, and determine the semantic information corresponding to the warning text information;
[0083] S207: Based on a convolutional encoder, determine the speech tokens corresponding to the warning text information using the semantic information, and determine the initial speech using a large language model based on the speech tokens;
[0084] S208: Perform linear spectrum analysis on several frames corresponding to the warning text information to obtain a target linear spectrum;
[0085] S209: Analyze the prosody splicing points of the warning text information, and generate prosody analysis data based on the prosody splicing points combined with a prosody sequence encoding;
[0086] S210: Adjust the initial speech based on the target linear spectrum and the prosody analysis data in combination with a speech splicing smoothing algorithm and a volume adjustment factor to obtain target speech information. The radar warning device and the pedestrian warning device perform traffic warnings based on the target speech information.
[0087] In the embodiment of the present invention, vehicle driving analysis is performed based on multi-dimensional Fourier transform using radar echo signals collected by a radar warning device in a radar detection area, so that the obtained vehicle driving information is more accurate, and the reliability of traffic safety warning is ensured. Based on occlusion analysis, face positioning and multi-layer filtering processing are performed on several frames of target images collected by a pedestrian warning device for a crosswalk to obtain a target face set for two-way pedestrian flow analysis. Considering the situation of face occlusion and introducing filtering of the face set, false detection or double counting is avoided, and the obtained pedestrian flow data is more accurate. Based on the pedestrian passing time prediction data and vehicle driving information, the recommended driving information of the vehicle in the warning area is analyzed, which can improve the reliability of the recommended driving information analysis and the intelligence of traffic safety warning at the same time, so that the driver can understand the current required driving information. Based on the recommended driving information, voice generation is performed using voice tokens and prosody analysis to obtain target voice information for communication warning, which improves the accuracy and fluency of voice generation and can provide more reliable traffic warnings for vehicles and pedestrians.
[0088] Embodiment III
[0089] Please refer to Figure 3 , Figure 3 FIG. is a schematic structural diagram of a traffic safety warning system for a crosswalk in an embodiment of the present invention. The system includes a radar warning device 31 and a pedestrian warning device 32, and the radar warning device 31 is communicatively connected to the pedestrian warning device 32; the system is configured to execute the traffic safety warning method for the crosswalk described in the above embodiment.
[0090] In the specific implementation process of the present invention, the applications of the radar warning device and the pedestrian warning device can be as Figure 5As shown, the radar warning device 31 is a warning device based on radar perception. It includes a perception radar, a horn, a red and blue flashing light, a light-emitting diode (LED) display screen, a unit control board, and an LED control board. The perception radar senses the data of road vehicles in the radar detection area and transmits it to the unit control board for analysis to form a warning voice and a warning text. The warning voice is transmitted to the horn for playback, and the warning text is transmitted to the LED control board to control the LED display screen for display. The driver can then learn about the corresponding deceleration or stop operations of the vehicle in the warning area. The red and blue flashing light is used to give a warning light indication. At the same time, when passing at night, the radar warning device can trigger and control the yellow flashing light located on the roadside to warn vehicles. The pedestrian warning device 32 is a warning device based on a perception camera. It includes a pedestrian perception camera, a yellow flashing light, an LED display screen, a horn, a unit control board, and an LED control board. The pedestrian perception camera is used to collect images of the crosswalk. The unit control board is used to analyze the images, generate a warning voice and a warning text, transmit the warning voice to the horn for broadcasting, and transmit the warning text to the LED control board to control the LED display screen for display. The yellow flashing light is used to give a warning light indication. At the same time, the pedestrian warning device can trigger and control the spotlight to illuminate the zebra crossing when there are pedestrians at night. When pedestrians pass through the crosswalk at night, their passing safety can be maximally guaranteed. The radar warning device 31 and the pedestrian warning device 32 can also exchange data to achieve the trigger control of the warning voice and text. Figure 3 The system shown does not constitute a limitation on all components and may include more or fewer components than shown, or combine certain components.
[0091] In the specific implementation process of the present invention, the specific implementation manners of the system items can refer to the above embodiments and will not be elaborated here.
[0092] In the embodiments of the present invention, vehicle driving analysis is performed based on multi-dimensional Fourier transform using the radar echo signals collected by the radar warning device in the radar detection area, so that the obtained vehicle driving information is more accurate, and the reliability of the traffic safety warning is ensured. Based on occlusion analysis, face positioning and multi-layer filtering processing are performed on a number of target images collected by the pedestrian warning device for the crosswalk to obtain a set of target faces for two-way pedestrian flow analysis. Considering the situation of face occlusion and introducing the filtering of the face set, the occurrence of missed detection or duplicate counting is avoided, and the obtained pedestrian flow data is more accurate. Based on the pedestrian passing time prediction data and vehicle driving information, the recommended driving information of the vehicle in the warning area is analyzed, which can improve the reliability of the recommended driving information analysis and the intelligence of the traffic safety warning at the same time, so that the driver can understand the current required driving information. Based on the recommended driving information, voice generation is performed using voice tokens and prosody analysis to obtain target voice information for communication warning, which improves the accuracy and fluency of voice generation and can provide more reliable traffic warnings for vehicles and pedestrians.
[0093] Embodiment 4
[0094] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of the traffic safety warning device for the crosswalk in the embodiments of the present invention. The device is applied to the radar warning device and the pedestrian warning device, and the radar warning device is communicatively connected to the pedestrian warning device; the device includes:
[0095] Vehicle driving analysis module 41: configured to perform vehicle driving analysis based on multi-dimensional Fourier transform using the radar echo signals collected by the radar warning device in the radar detection area to obtain vehicle driving information;
[0096] Face detection module 42: configured to perform face positioning and multi-layer filtering processing on a number of target images collected by the pedestrian warning device for the crosswalk based on occlusion analysis to obtain a set of target faces;
[0097] Pedestrian flow analysis module 43: configured to perform two-way pedestrian flow analysis based on the set of target faces to obtain two-way pedestrian flow data;
[0098] Passing time prediction module 44: configured to predict the pedestrian passing time of the crosswalk based on the two-way pedestrian flow data and a number of target images to obtain pedestrian passing time prediction data;
[0099] Driving recommendation analysis module 45: configured to analyze the recommended driving information of the vehicle in the warning area based on the pedestrian passing time prediction data and the vehicle driving information;
[0100] Early warning module 46: It is used to generate voice based on the recommended driving information by using voice tokens and prosody analysis to obtain target voice information, and the radar early warning device and the pedestrian early warning device perform traffic warning based on the target voice information.
[0101] In the specific implementation process of the present invention, the specific implementation manners of the device items can refer to the above embodiments and will not be elaborated here.
[0102] In the embodiment of the present invention, vehicle driving analysis is performed based on the multi-dimensional Fourier transform by using the radar echo signals collected by the radar early warning device in the radar detection area, so that the obtained vehicle driving information is more accurate, and the reliability of the traffic safety warning is ensured. Based on occlusion analysis, face positioning and multi-layer filtering processing are performed on several frames of target images collected by the pedestrian early warning device for the crosswalk to obtain a target face set for two-way pedestrian flow analysis. Considering the situation of face occlusion and introducing the filtering of the face set, it is avoided to miss detection or double counting, and the obtained pedestrian flow data is more accurate. Based on the pedestrian passing time prediction data and the vehicle driving information, the recommended driving information of the vehicle in the warning area is analyzed, which can improve the reliability of the recommended driving information analysis and the intelligence of the traffic safety warning at the same time, so that the driver can understand the current required driving information. Based on the recommended driving information, voice generation is performed by using voice tokens and prosody analysis to obtain target voice information for communication warning, which improves the accuracy and fluency of voice generation and can provide more reliable traffic warning for vehicles and pedestrians.
[0103] A computer-readable storage medium provided by an embodiment of the present invention, on which a computer program is stored, and when the program is executed by a processor, it implements the traffic safety warning method for the crosswalk in any one of the above embodiments. Among them, the computer-readable storage medium includes but is not limited to any type of disk (including floppy disks, hard disks, optical disks, CD-ROMs, and magneto-optical disks), ROM (Read-Only Memory), RAM (Random Access Memory), EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), flash memory, magnetic cards or optical cards. That is, the storage device includes any medium that can store or transmit information in a readable form by a device (such as a computer, a mobile phone), and can be a read-only memory, a magnetic disk or an optical disk, etc.
[0104] In addition, the above has introduced in detail a traffic safety warning method and related device for a crosswalk provided by the embodiments of the present invention. Specific examples should have been used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. At the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A traffic safety warning method for a crosswalk, characterized in that, Applied to a radar warning device and a pedestrian warning device, the radar warning device is communicatively connected to the pedestrian warning device; the method includes: Performing vehicle driving analysis on the radar echo signals collected by the radar warning device in the radar detection area by using multi-dimensional Fourier transform to obtain vehicle driving information; Performing face positioning and multi-layer filtering processing on a plurality of frames of target images collected by the pedestrian warning device for the crosswalk based on occlusion analysis to obtain a set of target faces; Performing two-way pedestrian flow analysis based on the set of target faces to obtain two-way pedestrian flow data; Predicting the pedestrian passing time of the crosswalk based on the two-way pedestrian flow data and a plurality of frames of target images to obtain pedestrian passing time prediction data; Analyzing the recommended driving information of the vehicle in the warning area based on the pedestrian passing time prediction data and the vehicle driving information; Generating target voice information by using voice tokens and prosody analysis based on the recommended driving information, and the radar warning device and the pedestrian warning device performing passing warnings based on the target voice information; Among them, the performing face positioning and multi-layer filtering processing on a plurality of frames of target images collected by the pedestrian warning device for the crosswalk based on occlusion analysis to obtain a set of target faces includes: performing passing direction marking on a plurality of frames of target images to obtain a plurality of frames of target images after passing direction marking; performing occlusion marking processing on the first set of faces obtained by positioning the plurality of frames of target images after passing direction marking based on a preset occlusion type table to obtain a first set of faces after occlusion marking processing, and generating a face feature descriptor based on the first set of faces after occlusion marking processing; determining a second set of faces in the plurality of frames of target images based on the face feature descriptor by using a face positioning model; filtering the second set of faces by using a face tracking algorithm to obtain a third set of faces, and filtering the third set of faces by using a feature clustering algorithm to obtain a set of target faces corresponding to the passing direction.
2. The traffic safety warning method for crosswalks according to claim 1, characterized in that The performing vehicle driving analysis on the radar echo signals collected by the radar warning device in the radar detection area by using multi-dimensional Fourier transform to obtain vehicle driving information includes: Performing filtering and noise reduction processing on the radar echo signals to obtain radar echo signals after filtering and noise reduction processing; Performing fast Fourier transform on the radar echo signals after filtering and noise reduction processing to obtain a Fourier spectrum, and analyzing the first vehicle speed, the first vehicle relative distance, and the first vehicle azimuth based on the Fourier spectrum by using an echo signal matrix; Performing ordered-statistic constant false alarm detection on the radar echo signals after filtering and noise reduction processing to obtain an ordered-statistic constant false alarm detection result; Analyzing the second vehicle speed, the second vehicle relative distance, and the second vehicle azimuth based on the ordered-statistic constant false alarm detection result by using peak spectral lines, and generating a target vehicle speed, a target vehicle relative distance, and a target vehicle azimuth based on the first vehicle speed, the first vehicle relative distance, the first vehicle azimuth, the second vehicle speed, the second vehicle relative distance, and the second vehicle azimuth; Generating vehicle driving information based on the target vehicle speed, the target vehicle relative distance, and the target vehicle azimuth.
3. The traffic safety warning method for crosswalks according to claim 1, wherein Performing two-way pedestrian flow analysis based on the target face set to obtain two-way pedestrian flow data, including: Performing two-way pedestrian flow analysis based on the marked target faces in the target face set using boolean parameters and detection frames to obtain two-way pedestrian flow data.
4. The traffic safety warning method for crosswalks according to claim 1, wherein Predicting the pedestrian passing time of the crosswalk based on the two-way pedestrian flow data and several frames of target images to obtain pedestrian passing time prediction data, including: Converting the pixel coordinates of each target face located in each frame of the target image after marking the passing direction to obtain the corresponding position coordinates; Analyzing the corresponding pedestrian passing speed based on the time stamps of each frame of the target image using the corresponding position coordinates; Predicting the pedestrian passing time of the crosswalk based on the two-way pedestrian flow data and the corresponding pedestrian passing speed to obtain pedestrian passing time prediction data.
5. The method for warning of traffic safety at a crosswalk according to claim 1, wherein, Analyzing the recommended driving information of the vehicle in the warning area based on the pedestrian passing time prediction data and vehicle driving information, including: Analyzing the traffic flow density based on the vehicle information in the warning area and analyzing the vehicle deceleration based on the pedestrian passing time prediction data; Calculating the safe distance based on the vehicle driving information combined with the driver's reaction time, and analyzing the recommended driving information of the vehicle in the warning area based on the traffic flow density, vehicle deceleration, and safe distance.
6. The method for warning of pedestrian crossing traffic safety according to claim 1, characterized in that Generating target voice information using voice tokens and prosody analysis based on the recommended driving information, including: Generating warning text information using a preset warning text template based on the recommended driving information, and determining the semantic information corresponding to the warning text information; Determining the voice tokens corresponding to the warning text information based on the semantic information using a convolutional encoder, and determining the initial voice based on the voice tokens using a large language model; Performing linear spectral analysis on several frames corresponding to the warning text information to obtain the target linear spectrum; Analyzing the prosody splicing points of the warning text information, and generating prosody analysis data based on the prosody splicing points combined with the prosody sequence encoding; Adjusting the initial voice based on the target linear spectrum and prosody analysis data combined with a voice splicing smoothing algorithm and a volume adjustment factor to obtain the target voice information.
7. A traffic safety warning device for a crosswalk, characterized in that, Applied to a radar warning device and a pedestrian warning device, the radar warning device is communicatively connected to the pedestrian warning device; the device includes: A vehicle driving analysis module: for performing vehicle driving analysis based on multi-dimensional Fourier transform using the radar echo signals collected by the radar warning device in the radar detection area to obtain vehicle driving information; A face detection module: for performing face positioning and multi-layer filtering processing on several frames of target images collected by the pedestrian warning device for the crosswalk based on occlusion analysis to obtain a target face set; A pedestrian flow analysis module: for performing two-way pedestrian flow analysis based on the target face set to obtain two-way pedestrian flow data; A passing time prediction module: for predicting the pedestrian passing time of the crosswalk based on the two-way pedestrian flow data and several frames of target images to obtain pedestrian passing time prediction data; A driving recommendation analysis module: for analyzing the recommended driving information of the vehicle in the warning area based on the pedestrian passing time prediction data and vehicle driving information; Early warning module: used to generate voice based on the recommended driving information by using voice tokens and prosody analysis to obtain target voice information, and the radar early warning device and pedestrian early warning device perform traffic warning based on the target voice information; Among them, the face positioning and multi-layer filtering processing of several frames of target images collected by the pedestrian warning device for the crosswalk based on occlusion analysis to obtain a target face set, including: marking the traffic direction of several frames of target images to obtain several frames of target images after traffic direction marking; performing occlusion marking processing on the first face set obtained by positioning several frames of target images after traffic direction marking based on a preset occlusion type table to obtain the first face set after occlusion marking processing, and generating a face feature descriptor based on the first face set after occlusion marking processing; determining the second face set in several frames of target images based on the face feature descriptor by using a face positioning model; filtering the second face set based on a face tracking algorithm to obtain a third face set, and filtering the third face set based on a feature clustering algorithm to obtain a target face set corresponding to the traffic direction.
8. A traffic safety warning system for a crosswalk, characterized in that, The system includes a radar early warning device and a pedestrian early warning device, the radar early warning device is communicatively connected to the pedestrian early warning device, and the system is configured to execute the traffic safety warning method for the crosswalk according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the computer instructions run on an electronic device, the electronic device is caused to execute the traffic safety warning method for the crosswalk according to any one of claims 1 to 6.
Citation Information
Patent Citations
Intersection safety early warning system based on machine vision and DSRC and early warning method thereof
CN111210662A
Bidirectional pedestrian volume statistical method based on RGB-D multi-modal data
CN111881749A
Traffic flow detection method and device
CN112233416A
Method and device for converting text into voice, equipment and storage medium
CN119517004A