Voice recognition method, voice recognition model training method and electronic equipment

By performing frequency domain and time domain feature detection, fusion, integration, and segmentation processing on the sound signals around vehicles, the problem of ordinary vehicles having difficulty recognizing the sirens of special vehicles has been solved, achieving more accurate recognition and timely response, thereby improving road safety and user experience.

CN121053969APending Publication Date: 2025-12-02Z-ONE TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511267021.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-05
Publication Date
2025-12-02

AI Technical Summary

Technical Problem

In existing technologies, ordinary vehicles have difficulty accurately identifying the sirens of special vehicles while driving, leading to delayed avoidance, affecting the effectiveness of mission execution, and increasing the risk of traffic accidents.

Method used

By performing feature detection on the target sound signal, frequency domain and time domain feature information is obtained, and then splicing, fusion, integration and segmentation processing are performed, followed by weighted processing, to finally identify whether the target sound signal contains the prompt sound of a special vehicle.

Benefits of technology

It improves the accuracy and timeliness of identifying sirens from special vehicles, enhances emergency response efficiency, reduces the probability of traffic accidents, and improves road safety and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121053969A_ABST
    Figure CN121053969A_ABST
Patent Text Reader

Abstract

The invention discloses a sound recognition method, a sound recognition model training method and electronic equipment, and the method comprises the steps: carrying out the feature detection processing of a target sound signal, and obtaining the frequency domain feature information and time domain feature information of the target sound signal; performing splicing fusion processing on the frequency domain feature information and the time domain feature information to obtain first feature spectrum information corresponding to the target sound signal; performing integration and segmentation processing on the first characteristic spectrum information to obtain second characteristic spectrum information; weighting the first characteristic spectrum information and the second characteristic spectrum information to obtain target characteristic spectrum information; and carrying out identification processing on the target characteristic spectrum information to obtain a target identification result which corresponds to the target sound signal and represents whether the target sound signal comprises a target prompt tone corresponding to the target special vehicle or not. Therefore, the recognition accuracy of the target prompt tone such as siren sound of the target special vehicle is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of sound recognition technology, and in particular to a sound recognition method, a sound recognition model training method, and an electronic device. Background Technology

[0002] While driving, drivers frequently encounter emergency vehicles such as fire trucks, police cars, and ambulances. These vehicles emit warning sounds, such as sirens, to alert other vehicles to give way and minimize obstruction. However, due to factors like the vehicle's own sound insulation, noise interference from inside and outside the vehicle, and driver distraction, ordinary vehicles may not accurately recognize these warning sounds, leading to delayed avoidance and potentially impacting the emergency vehicle's mission performance. Furthermore, emergency vehicles typically travel at high speeds during missions; failure to yield in time or illegal driving by ordinary vehicles could result in a collision and a traffic accident.

[0003] Therefore, more accurate identification of warning sounds such as sirens from special vehicles is crucial for the effectiveness of special vehicle missions and for traffic safety. Summary of the Invention

[0004] This application provides a sound recognition method, a sound recognition model training method, and an electronic device to solve the problem in the prior art that it cannot accurately recognize the warning sounds such as sirens of special vehicles.

[0005] To address the aforementioned technical problems, in a first aspect, this application discloses a sound recognition method. The method includes: performing feature detection processing on a target sound signal to obtain frequency domain feature information and time domain feature information of the target sound signal, wherein the target sound signal is obtained by detecting sounds around a vehicle; performing splicing and fusion processing on the frequency domain feature information and the time domain feature information to obtain first feature map information corresponding to the target sound signal; performing integration and segmentation processing on the first feature map information to obtain second feature map information; performing weighted processing on the first feature map information and the second feature map information to obtain target feature map information; and performing recognition processing on the target feature map information to obtain a target recognition result corresponding to the target sound signal, wherein the target recognition result indicates whether the target sound signal includes a target prompt sound corresponding to a specific target vehicle.

[0006] By employing the above technical solution, frequency domain and time domain feature information of the target sound signal can be obtained based on the target sound signal around the vehicle. Then, the frequency domain and time domain feature information are spliced ​​and fused to obtain the first feature map information corresponding to the target sound signal. Attention information integration and segmentation processing are then performed on the first feature map information to obtain the second feature map information. The first and second feature map information are then weighted to obtain the target feature map information. Finally, the target feature map information is processed for recognition to obtain the target recognition result indicating whether the target sound signal includes the target prompt sound corresponding to the target vehicle. Therefore, by acquiring and processing the time domain and frequency domain feature information of the target sound signal, the accuracy of the spliced ​​first feature map information can be effectively increased. Furthermore, the first and second feature map information are weighted, so that the resulting target feature map information can pay more attention to the target prompts of special vehicles, such as sirens, in the target sound signal. This makes the target recognition results more accurate, increases the accuracy and timeliness of target prompts such as sirens of special vehicles, improves the emergency response efficiency and rescue effect of special vehicles, improves road safety, reduces the probability of traffic accidents, and enhances the user experience.

[0007] According to another specific implementation of this application, the implementation of this application discloses a sound recognition method that performs feature detection processing on a target sound signal to obtain frequency domain feature information and time domain feature information of the target sound signal, including: performing short-time Fourier transform processing on the target sound signal to obtain frequency domain feature information corresponding to the target sound signal; and performing time domain modeling processing on the target sound signal to obtain time domain feature information corresponding to the target sound signal.

[0008] By adopting the above technical solution, by acquiring the frequency domain feature information and time domain feature information of the target sound signal, it is possible to obtain the spatial characteristics of the target sound signal and the dynamic changes of the target sound signal. This effectively increases the accuracy and timeliness of target warning sound recognition, such as the siren sound of a special vehicle, improves the emergency response efficiency and rescue effect of the target special vehicle, enhances road safety, and reduces the probability of traffic accidents.

[0009] According to another specific implementation of this application, the voice recognition method disclosed in this implementation integrates and segments first feature map information to obtain second feature map information, including: integrating the first feature map information to obtain spatial feature map information corresponding to the first feature map information; and segmenting the spatial feature map information to obtain second feature map information.

[0010] By adopting the above technical solution, through integration and segmentation processing, the second feature map information can pay more attention to the target prompt sound of special vehicles under noisy conditions, thereby increasing the accuracy of target recognition results.

[0011] According to another specific implementation of this application, a sound recognition method disclosed in this implementation integrates and processes first feature map information to obtain spatial feature map information corresponding to the first feature map information, including: encoding the first feature map information along the X direction corresponding to the first feature map information to obtain spatial feature map information in the X direction; encoding the first feature map information along the Y direction corresponding to the first feature map information to obtain spatial feature map information along the Y direction; and integrating the spatial feature map information along the X direction and the spatial feature map information along the Y direction to obtain spatial feature map information corresponding to the first feature map information.

[0012] By adopting the above technical solution, the first feature map information is first encoded to obtain spatial feature map information in the X direction and spatial feature map information in the Y direction. Then, the spatial feature map information along the X direction and the spatial feature map information along the Y direction are integrated to obtain the spatial feature map information corresponding to the first feature map information. This allows the spatial feature map information to pay more attention to the target prompt sound of the target special vehicle under noisy conditions, thereby increasing the accuracy of the target recognition result.

[0013] According to another specific implementation of this application, the sound recognition method disclosed in this implementation method involves segmenting spatial feature map information to obtain second feature map information, including: segmenting spatial feature map information to obtain first sub-feature map information and second sub-feature map information as the second feature map information, wherein the first sub-feature map information and the second sub-feature map information are feature map information including different spatial information.

[0014] By employing the above technical solution, the spatial feature map information is segmented to obtain a first sub-feature map information and a second sub-feature map information, which include different spatial information. This second feature map information can better focus on the target prompt sound of special vehicles under noisy conditions, thereby increasing the accuracy of target recognition results.

[0015] According to another specific implementation of this application, the voice recognition method disclosed in this application performs weighted processing on first feature map information and second feature map information to obtain target feature map information, including: performing weighted processing on first feature map information, first sub-feature map information and second sub-feature map information to obtain target feature map information.

[0016] By adopting the above technical solution and through weighted processing, the accuracy of the obtained target feature map information is effectively increased.

[0017] According to another specific implementation of this application, the implementation of this application discloses a sound recognition method that performs recognition processing on target feature map information to obtain a target recognition result, including: performing recognition processing on target feature map information to obtain a prompt sound recognition score as the target recognition result, wherein if the prompt sound recognition score is greater than or equal to a score threshold, it indicates that the target sound signal includes the target prompt sound corresponding to the target special vehicle, and if the prompt sound recognition score is less than the score threshold, it indicates that the target sound signal does not include the target prompt sound corresponding to the target special vehicle.

[0018] By employing the above technical solution, target feature map information is processed to obtain a prompt sound recognition score as the target recognition result. Based on this score, it is determined whether the target sound signal includes the prompt sound corresponding to the target vehicle. This makes the target recognition result more intuitive and effectively increases the accuracy and timeliness of recognizing target prompt sounds such as sirens of target vehicles.

[0019] Secondly, this application also discloses a sound recognition method applied to a target sound recognition model. The target sound recognition model includes a sound detection module, a splicing and fusion module, a coordinate attention module, a weighting module, and a classification module. The method includes: the sound detection module performing feature detection processing on the input target sound signal to obtain frequency domain feature information and time domain feature information of the target sound signal; inputting the frequency domain feature information and time domain feature information into the splicing and fusion module; the target sound signal is obtained by detecting the sound around the target vehicle; the splicing and fusion module performing splicing and fusion processing on the input frequency domain feature information and time domain feature information to obtain the first corresponding to the target sound signal. The system firstly inputs a feature map into a coordinate attention module and a weighting module. The coordinate attention module integrates and segments the first feature map to obtain a second feature map, which is then input into the weighting module. The weighting module weights the first and second feature maps to obtain a target feature map, which is then input into a classification module. The classification module performs recognition processing on the target feature map to obtain a target recognition result corresponding to the target sound signal. This target recognition result indicates whether the target sound signal includes the target prompt sound corresponding to the target special vehicle.

[0020] Thirdly, the implementation of this application also discloses a sound recognition model training method, which includes: determining a first dataset, the first dataset including multiple training samples, the multiple training samples being obtained by collecting sound samples from real road scenes; training an initial sound recognition model based on the first dataset until the loss function corresponding to the initial sound recognition model satisfies the target condition, thereby obtaining a target sound recognition model, which is used to implement any of the sound recognition methods described in the first aspect above, or to implement the sound recognition method described in the second aspect above.

[0021] Fourthly, this application also discloses a sound recognition device, comprising: a first processing module for performing feature detection processing on a target sound signal to obtain frequency domain feature information and time domain feature information of the target sound signal, wherein the target sound signal is obtained by detecting sounds around the vehicle; a second processing module for splicing and fusing the frequency domain feature information and the time domain feature information to obtain first feature map information corresponding to the target sound signal; a third processing module for integrating and segmenting the first feature map information to obtain second feature map information; a fourth processing module for weighting the first feature map information and the second feature map information to obtain target feature map information; and a fifth processing module for recognizing the target feature map information to obtain a target recognition result corresponding to the target sound signal, wherein the target recognition result indicates whether the target sound signal includes a target prompt sound corresponding to a specific target vehicle.

[0022] Fifthly, this application also discloses an electronic device, including: a processor and a memory communicatively connected to the processor; the memory stores a computer program; the processor executes the computer program stored in the memory to enable the electronic device to implement any of the voice recognition methods described in the first aspect above, or to implement the voice recognition method described in the second aspect above, or to implement the voice recognition model training method described in the third aspect above.

[0023] Sixthly, the implementation of this application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, is used to implement any of the sound recognition methods described in the first aspect above, or to implement the sound recognition method described in the second aspect above, or to implement the sound recognition model training method described in the third aspect above.

[0024] In a seventh aspect, the present application provides a computer program product, including a computer program that, when executed by a processor, implements any of the sound recognition methods described in the first aspect above, or implements the sound recognition method described in the second aspect above, or implements the sound recognition model training method described in the third aspect above.

[0025] It is understood that the beneficial effects of the second to seventh aspects mentioned above can also be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0026] Figure 1 This is a schematic flowchart of a voice recognition method provided in an embodiment of this application;

[0027] Figure 2 This is a schematic diagram of the frequency domain-time domain spectrum fusion process provided in an embodiment of this application;

[0028] Figure 3 This is a schematic diagram of a process for obtaining second feature map information provided in an embodiment of this application;

[0029] Figure 4 This is a schematic diagram of a process for obtaining spatial feature map information provided in an embodiment of this application;

[0030] Figure 5 This is a schematic diagram of a process for obtaining target feature map information provided in an embodiment of this application;

[0031] Figure 6 This is a schematic diagram of the structure of the target sound recognition model provided in the embodiments of this application;

[0032] Figure 7 This is another schematic flowchart of the voice recognition method provided in the embodiments of this application;

[0033] Figure 8 This is a schematic flowchart of a sound recognition model training method provided in an embodiment of this application;

[0034] Figure 9 This is another schematic flowchart of the voice recognition method provided in the embodiments of this application;

[0035] Figure 10 This is another schematic flowchart of the voice recognition method provided in the embodiments of this application;

[0036] Figure 11 This is a schematic diagram illustrating the detection performance of the target sound recognition model under different input features provided in this application embodiment on a test set;

[0037] Figure 12 This is a schematic diagram of the output of different sound recognition models provided in the embodiments of this application on representative siren sounds and noise samples;

[0038] Figure 13 This is a schematic diagram of the output of the sound recognition model provided in this application embodiment on samples of siren sounds in a stationary state and at high speed.

[0039] Figure 14 This is a schematic diagram of a voice recognition device provided in an embodiment of this application.

[0040] Figure 15 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0041] Specific implementation method

[0042] As mentioned earlier, in the existing technology, ordinary vehicles may not accurately recognize warning sounds such as sirens from emergency vehicles outside the vehicle due to factors such as the vehicle's own physical sound insulation, interference from noise inside or outside the vehicle, and the driver's own distraction. This can lead to delayed avoidance of emergency vehicles, affecting the speed and effectiveness of emergency vehicle missions. Moreover, emergency vehicles generally travel at relatively high speeds during missions, and if ordinary vehicles fail to avoid them in time or violate traffic regulations, they are highly likely to collide with emergency vehicles, causing traffic accidents.

[0043] A siren is a special acoustic signal emitted by vehicles, typically emergency vehicles such as fire trucks, police cars, and ambulances, to alert pedestrians and other vehicles on the road. Siren detection helps drivers better understand road conditions, effectively improving driving safety. To increase the driver's ability to recognize external siren sounds, current technology uses directional light or electromagnetic waves as warning signals on special vehicles. However, most ordinary vehicles do not have electromagnetic wave receivers, and light waves transmitted from behind cannot be received in time when the driver's attention is diverted.

[0044] Furthermore, while siren sounds possess stable frequency components, at high speeds, the frequency distribution of the collected sound signal may change over time due to the Doppler effect and environmental noise. Spectral analysis has limited ability to process such non-stationary signals. Traditional siren detection primarily relies on spectral analysis methods, analyzing frequency components in the spectrum that change with specific periodic patterns, resulting in unsatisfactory detection performance. In recent years, due to the rise of deep learning, more and more researchers are leveraging the powerful nonlinear modeling capabilities of deep neural networks to improve the performance of models in real-world siren detection scenarios. Most methods are based on Convolutional Neural Networks (CNNs) to model sound signals in the frequency domain. CNN models process frequency domain features such as the log-Mel spectrum or Mel-frequency cepstral coefficients (MFCCs) of the signal and utilize spatial correlation for classification to determine the presence or absence of a siren. However, this method of modeling signals only in the frequency domain has certain limitations. It is difficult to effectively detect target sounds in noisy environments with vehicles traveling at high speeds. Relying solely on frequency domain data ignores the dynamic changes of the sound signal in the time domain, which are crucial for identifying sirens and other similar sounds. Furthermore, most existing models assume the microphone is stationary during sample recording, neglecting the effects of wind noise, tire noise, and other environmental interference during high-speed vehicle movement, as well as the Doppler effect between the vehicle and the vehicle containing the sirens.

[0045] Based on this, this application provides a sound recognition method that, by acquiring and processing the time-domain and frequency-domain feature information of the target sound signal, can effectively increase the accuracy of the spliced ​​first feature map information. Furthermore, by weighting the first and second feature map information, the resulting target feature map information can better focus on target alert sounds of special vehicles, such as sirens, within the target sound signal. This makes the target recognition result more accurate, increases the accuracy and timeliness of target alert sound recognition (e.g., sirens), improves emergency response efficiency and rescue effectiveness for special vehicles, enhances road safety, and reduces the probability of traffic accidents.

[0046] Next, with reference to the accompanying drawings, the steps and advantages of the voice recognition method provided in this application will be described in detail.

[0047] In one implementation of this application, such as Figure 1 As shown, the voice recognition method provided in this application includes the following steps.

[0048] S100: Perform feature detection processing on the target sound signal to obtain the frequency domain feature information and time domain feature information of the target sound signal. The target sound signal is obtained by detecting the sound around the vehicle. The vehicle can be an ordinary vehicle driven by the user, or a special vehicle that is responding to an emergency or not (such as a police car, fire truck, ambulance, etc.).

[0049] In one implementation of this application, the target sound signal can be obtained by detecting sound within a preset range around the vehicle.

[0050] S200: The frequency domain feature information and time domain feature information are spliced ​​and fused to obtain the first feature map information corresponding to the target sound signal.

[0051] S300: Integrate and segment the first feature map information to obtain the second feature map information.

[0052] S400: Weight the first feature map information and the second feature map information to obtain the target feature map information.

[0053] S500: The target feature map information is processed to obtain the target recognition result corresponding to the target sound signal. The target recognition result is used to indicate whether the target sound signal includes the target prompt sound corresponding to the target special vehicle.

[0054] Target prompts could be, for example, the sirens of special vehicles.

[0055] The sound recognition method provided in this application can obtain the frequency domain feature information and time domain feature information of the target sound signal based on the target sound signal around the vehicle. Then, the frequency domain feature information and time domain feature information are spliced ​​and fused to obtain the first feature spectrum information corresponding to the target sound signal. The first feature spectrum information is then integrated and segmented to obtain the second feature spectrum information. The first and second feature spectrum information are then weighted to obtain the target feature spectrum information. Finally, the target feature spectrum information is processed to obtain the target recognition result indicating whether the target sound signal includes the target prompt sound corresponding to the target vehicle. Therefore, by acquiring and processing the time domain feature information and frequency domain feature information of the target sound signal, the accuracy of the spliced ​​first feature spectrum information can be effectively increased. Furthermore, the first and second feature map information are weighted, so that the resulting target feature map information can pay more attention to the target prompts of special vehicles, such as sirens, in the target sound signal. This makes the target recognition results more accurate, increases the accuracy and timeliness of target prompts such as sirens of special vehicles, improves the emergency response efficiency and rescue effect of special vehicles, improves road safety, reduces the probability of traffic accidents, and enhances the user experience.

[0056] In another implementation of this application, feature detection processing is performed on the target sound signal to obtain frequency domain feature information and time domain feature information of the target sound signal, including: performing short-time Fourier transform processing on the target sound signal to obtain frequency domain feature information corresponding to the target sound signal; and performing time domain modeling processing on the target sound signal to obtain time domain feature information corresponding to the target sound signal.

[0057] In another implementation of this application, such as Figure 2 As shown, for the input siren sound sample (i.e., the target sound signal), on the one hand, a Short-time Fourier Transform (STFT) is performed on the sound signal to obtain its frequency domain representation (i.e., frequency domain feature information, frequency domain spectrogram); on the other hand, a Convolutional Neural Network (CNN) is used to model the sound signal in the time domain to obtain its time domain representation (i.e., time domain feature information, time domain spectrogram). Then, the time domain spectrogram and the frequency domain spectrogram are concatenated to obtain the fused spectrogram of the sound signal (i.e., the first feature map information), which serves as the subsequent input.

[0058] In another implementation of this application, such as Figure 3 As shown, the second feature map information is obtained by integrating and segmenting the first feature map information, including the following steps:

[0059] S310: Integrate and process the first feature map information to obtain the spatial feature map information corresponding to the first feature map information.

[0060] S320: Segment the spatial feature map information to obtain the second feature map information.

[0061] In another implementation of this application, such as Figure 4 As shown, the spatial feature map information corresponding to the first feature map information is integrated and processed, including the following steps:

[0062] S311: Encode the first feature map information along the X direction corresponding to the first feature map information to obtain spatial feature map information along the X direction; and encode the first feature map information along the Y direction corresponding to the first feature map information to obtain spatial feature map information along the Y direction.

[0063] S312: Integrate the spatial feature map information along the X direction and the spatial feature map information along the Y direction to obtain the spatial feature map information corresponding to the first feature map information.

[0064] In another implementation of this application, the spatial feature map information is segmented to obtain the second feature map information, including: segmenting the spatial feature map information to obtain the first sub-feature map information and the second sub-feature map information, which serve as the second feature map information. The first sub-feature map information and the second sub-feature map information are feature map information that includes different spatial information.

[0065] In another implementation of this application, the first feature map information and the second feature map information are weighted to obtain the target feature map information, including: weighting the first feature map information, the first sub-feature map information and the second sub-feature map information to obtain the target feature map information.

[0066] In another implementation of this application, the target feature map information is processed to obtain a target recognition result, including: processing the target feature map information to obtain a prompt sound recognition score as the target recognition result; and if the prompt sound recognition score is greater than or equal to a score threshold, it means that the target sound signal includes the target prompt sound corresponding to the target special vehicle; if the prompt sound recognition score is less than the score threshold, it means that the target sound signal does not include the target prompt sound corresponding to the target special vehicle.

[0067] For example, a classifier can be built using CNN networks and fully connected networks. The classifier calculates a score (i.e., a prompt sound recognition score) for each sample (i.e., target feature map information) to determine whether a siren sound (i.e., target prompt sound) is present in the sample. The output results are then used with fully connected networks and SoftMax.

[0068] In another implementation of this application, a coordinate attention module (CAM) is introduced, such as... Figure 5 As shown, the method includes the following steps:

[0069] The fused spectrum (i.e., the first feature map information) is obtained, and the fused spectrum is weighted with attention information through the coordinate attention module CAM.

[0070] Specifically, CAM uses two parallel average pooling kernels (X-direction average pooling kernel and Y-direction average pooling kernel) to encode the features along the X and Y directions, obtaining spatial feature map information along the X and Y directions to integrate the spatial information related to the features. Then, a two-dimensional CNN network is used to interact with the information to obtain spatial feature map information. Furthermore, a two-dimensional CNN network is used again to segment the interacting information, splitting it into two different position-aware feature maps (e.g., the aforementioned first and second sub-feature map information). These two different position-aware feature maps contain different spatial information. Applying the position-aware feature maps back to the input features (i.e., the first feature map information) provides attention weighting to the original input (i.e., the target sound signal), enabling greater attention to siren sounds under noisy conditions.

[0071] In another implementation of this application, a sound recognition method is provided, which is applied to a target sound recognition model. For example... Figure 6 As shown, the target sound recognition model includes a sound detection module, a splicing and fusion module, a coordinate attention module, a weighting module, and a classification module.

[0072] like Figure 7 As shown, the method includes the following steps:

[0073] S10: The sound detection module performs feature detection processing on the input target sound signal to obtain the frequency domain feature information and time domain feature information of the target sound signal, and inputs the frequency domain feature information and time domain feature information to the splicing and fusion module. The target sound signal is obtained by detecting the sound around the target vehicle.

[0074] S20: The splicing and fusion module splices and fuses the input frequency domain feature information and time domain feature information to obtain the first feature map information corresponding to the target sound signal, and inputs the first feature map information to the coordinate attention module.

[0075] S30: The coordinate attention module integrates and segments the input first feature map information to obtain the second feature map information, and then inputs the second feature map information into the weighting module.

[0076] S40: The weighting module performs weighted processing on the input first feature map information and second feature map information to obtain the target feature map information, and inputs the target feature map information into the classification module.

[0077] S50: The classification module performs recognition processing on the input target feature map information to obtain the target recognition result corresponding to the target sound signal.

[0078] The target recognition result is used to indicate whether the target sound signal includes the target prompt sound corresponding to the target special vehicle.

[0079] In another implementation of this application, a method for training a sound recognition model is provided, characterized in that, as Figure 8 As shown, the method includes the following steps:

[0080] S001: Determine the first dataset. The first dataset includes multiple training samples, which are obtained by collecting sound samples from real road scenes.

[0081] S002: Train the initial sound recognition model based on the first dataset until the loss function corresponding to the initial sound recognition model satisfies the target condition, and obtain the target sound recognition model.

[0082] First, we import a sample of siren sounds collected from a real highway scene. The sound sample is usually a superposition of the target siren sound and various other noises. Then, we set a loss function for the model, and the training objective of the model is to minimize the loss function. After the model is trained, it will be able to perform the siren sound detection task.

[0083] The sound recognition method provided in this application can also be described as a special vehicle sound detection algorithm that integrates frequency domain and time domain features. For example, such as... Figure 9 and 10As shown, firstly, the target detection sound signal (i.e., the target sound signal, input signal, or target detection sound source) is imported, and the loss function for the model training process is set. Secondly, algorithmic processing is performed to obtain the frequency domain encoded spectrogram (i.e., frequency domain representation) and the time domain encoded spectrogram (i.e., time domain representation) of the target detection sound signal. These two spectrograms are then concatenated and fused to form a fused feature spectrogram (i.e., the first feature spectrogram information). Then, the fused feature spectrogram is processed by a coordinate attention module to integrate information, resulting in an integrated feature spectrogram (i.e., spatial feature spectrogram information). A CNN network is used to perform information interaction and segmentation processing on the integrated feature spectrogram, resulting in a segmented feature spectrogram (i.e., the second feature spectrogram information). The fused feature spectrogram and the segmented feature spectrogram are then reweighted. Finally, the weighted feature spectrogram (i.e., target feature spectrogram information) is output. Based on the CNN network, the weighted feature spectrogram is classified, and the detection result (target recognition result) is output. The sound recognition method provided in this application can, after inputting a target detection sound signal, output a judgment result (whether there is a siren sound emitted by a special vehicle in the input signal) after processing by the model trained by the algorithm.

[0084] The sound recognition method provided in this application is adaptable to different noise environments and has strong adaptability. It solves the problem of low accuracy in detecting alarm sounds in existing pure time-domain detection methods, and also addresses the insufficient detection accuracy of existing methods in noisy environments at high speeds. This improves road safety, enhances emergency response efficiency, reduces driver stress and anxiety caused by not hearing the siren, and ultimately enhances the driving experience.

[0085] This application addresses the shortcomings of existing siren detection methods, such as difficulty in effectively detecting special vehicle sounds at high speeds and insufficient detection accuracy due to neglecting temporal variations in the detection signal. It proposes a special vehicle sound detection algorithm that integrates frequency and temporal features. This reduces the probability of traffic accidents caused by the failure to hear siren, thereby minimizing the resulting economic losses and lowering the likelihood of traffic accidents.

[0086] Next, based on specific experimental data, we will demonstrate the effectiveness of the voice recognition method provided in this application.

[0087] like Figure 11The results shown are experimental verification results of the sound recognition method provided in this application on the LSAD-EVSRN dataset. Compared with using temporal features alone as input, using both features improves the area under the ROC curve (AUC) by 14.88% and the area under the pROC curve (pAUC) by 11.06%. Compared with using frequency domain features alone as input, using the fusion of the two features as model input improves AUC by 2.52% and pAUC by 13.31%. This indicates that using both features as input allows the model to better learn the correlation between the two features, which is beneficial to improving the model's detection performance. On the other hand, compared with using temporal features, using frequency domain features improves AUC by 12.36% while decreasing pAUC by 2.15%, indicating that the information provided by the two features is different. After using both features as input, both AUC and pAUC are effectively improved, indicating that there is complementarity between the two features when performing siren sound detection tasks.

[0088] like Figure 12 As shown, to demonstrate the model's performance on road sound samples with stronger noise interference under different input features, four noisy siren sound samples (Ambulance820.wav, Ambulance821.wav, Ambulance822.wav, Ambulance823.wav) and four corresponding noise samples (Road96.wav, Road97.wav, Road98.wav, Road99.wav) were selected from the dataset. In the figure, the true label for sound samples containing siren sounds is 1, and the true label for sound samples without siren sounds is 0. Model 1 uses frequency domain features as input, Model 2 uses only time domain features as input, and Model 3 uses both types of features as input. The model detection threshold is set to 0.5; that is, if the model outputs a detection score greater than or equal to 0.5, it considers a siren sound to be present; otherwise, the task indicates no siren sound was detected. Due to the uncertainty and interference in real road environments, some models may mistakenly identify horn sounds as siren sounds, leading to misjudgments. On siren sound samples numbered 820, 822, and 823, both Model 1 and Model 2 exhibited false detections. This is due to interference from human voices and other sound sources in these samples, coupled with the fact that the siren sound was far from the microphone and its intensity was low. Model 3, using time-domain and frequency-domain features, also detected accurately in these samples with strong interference signals and weak siren sounds, demonstrating the superior performance of the proposed method. In the demonstrated road noise samples, all models detected correctly, indicating that the models can accurately model noise signals.

[0089] Furthermore, such as Figure 13 As shown, to demonstrate the effectiveness of the sound recognition method provided in this application, a verification experiment was conducted using a real dataset, collecting sound samples from four different real-world scenarios: siren samples in a stationary state, noise samples in a stationary state, siren samples when a vehicle is traveling at 80 km / h, and noise samples when a vehicle is traveling at 80 km / h. Figure 13 The sample of the siren sound in a stationary state is shown in the sample, where both the data collection vehicle and the simulated emergency vehicle are stationary. Samples 1 and 2 of the siren sound at high speed represent the sound samples of the simulated emergency vehicle approaching and moving away from the data collection vehicle, respectively. These samples introduce noise interference such as wind noise and tire noise, so the driver may not be able to detect the siren sound when the two vehicles are far apart. Figure 13 As can be seen, the model can accurately detect the presence of sirens in both stationary and high-speed driving states.

[0090] In another implementation of this application, such as Figure 14 As shown, a sound recognition device is provided, comprising: a first processing module for performing feature detection processing on a target sound signal to obtain frequency domain feature information and time domain feature information of the target sound signal, wherein the target sound signal is obtained by detecting sounds around a vehicle; a second processing module for splicing and fusing the frequency domain feature information and the time domain feature information to obtain a first feature spectrum information corresponding to the target sound signal; a third processing module for integrating and segmenting the first feature spectrum information to obtain a second feature spectrum information; a fourth processing module for weighting the first feature spectrum information and the second feature spectrum information to obtain a target feature spectrum information; and a fifth processing module for recognizing the target feature spectrum information to obtain a target recognition result corresponding to the target sound signal, wherein the target recognition result indicates whether the target sound signal includes a target prompt sound corresponding to a specific target vehicle.

[0091] The first to fifth processing modules may, for example, include a sound detection module, a splicing and fusion module, a coordinate attention module, a weighting module, and a classification module, respectively, as part of the aforementioned target sound recognition model.

[0092] Please see Figure 15 , Figure 15 The diagram shown is a structural schematic of an electronic device provided in an embodiment of this application. Figure 15 As shown, the electronic device may include: transceiver 121, processor 122, and memory 123.

[0093] The processor 122 executes computer execution instructions stored in the memory, causing the processor 122 to perform part of the technical solutions of the voice recognition method or voice recognition model training method in the above embodiments. The processor 122 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital data processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0094] The memory 123 is connected to the processor 122 via the system bus and completes communication between them. The memory 123 is used to store computer program instructions.

[0095] For example, and not as a limitation, memory 123 may include a hard disk drive (HDD), a floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 123 may include removable or non-removable (or fixed) media. Where appropriate, memory 123 may be internal or external to the integrated gateway device. In a particular embodiment, memory 123 is non-volatile solid-state memory. In a particular embodiment, memory 123 includes read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or flash memory, or a combination of two or more of these.

[0096] Transceiver 121 can be used to obtain the task to be run and its configuration information.

[0097] The system bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The system bus can be divided into address bus, data bus, control bus, etc. For ease of representation, only one thick line is used in the diagram, but this does not indicate that there is only one bus or one type of bus. Transceivers are used to enable communication between database access devices and other computers (e.g., clients, read-write libraries, and read-only libraries). Memory may include random access memory (RAM) and may also include non-volatile memory.

[0098] This application also provides a chip for executing instructions, which is used to execute the technical solutions of the sound recognition method or sound recognition model training method in the above embodiments.

[0099] This application also provides a computer-readable storage medium storing computer instructions. When the computer instructions are executed on the processor of an electronic device, the processor of the electronic device executes the technical solution of the sound recognition method or sound recognition model training method described in the above embodiments.

[0100] In some possible implementations, various aspects of the methods provided in this application can also be implemented as a program product, comprising program code. When the program product is run on the processor of an electronic device, the program code causes the processor of the electronic device to execute the steps of the methods according to the various exemplary implementations of this application described above. For example, the electronic device can execute the voice recognition method or voice recognition model training method described in the embodiments of this application. The program product can take the form of any combination of one or more readable media. The readable media can be a readable data medium or a readable storage medium.

[0101] This application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. At least one processor can read the computer program from the computer-readable storage medium. When the at least one processor executes the computer program, it can implement the technical solutions of the sound recognition method or the sound recognition model training method in the above embodiments.

[0102] This application is described with reference to flowchart illustrations and / or block diagrams of the methods, apparatus, and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable information processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable information processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0103] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable information processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0104] These computer program instructions may also be loaded onto a computer or other programmable information processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0105] It should be noted that, in addition to the specific embodiments described above, those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. Although the description of this application is presented in conjunction with preferred embodiments, this does not mean that the features of this invention are limited to this implementation. On the contrary, the purpose of describing the invention in conjunction with the implementation is to cover other options or modifications that may be derived based on the claims of this application. To provide a thorough understanding of this application, many specific details are included in the above description, and this application may also be implemented without using these details. Furthermore, to avoid confusion or obscuring the focus of this application, some specific details will be omitted in the description. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other.

[0106] It should be noted that in this specification, similar reference numerals and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0107] Although this application has been illustrated and described with reference to certain implementations thereof, those skilled in the art should understand that the above description is a further detailed explanation of the application in conjunction with specific implementations, and should not be construed as limiting the specific implementation of the application to these descriptions. Those skilled in the art can make various changes in form and detail, including some simple deductions or substitutions, without departing from the spirit and scope of this application.

Claims

1. A voice recognition method, characterized in that, The method includes: The target sound signal is processed by feature detection to obtain its frequency domain feature information and time domain feature information. The target sound signal is obtained by detecting the sound around the vehicle. The frequency domain feature information and the time domain feature information are spliced ​​and fused to obtain the first feature map information corresponding to the target sound signal; The first feature map information is integrated and segmented to obtain the second feature map information; The first feature map information and the second feature map information are weighted to obtain the target feature map information; The target feature map information is processed to obtain the target recognition result corresponding to the target sound signal. The target recognition result is used to indicate whether the target sound signal includes the target prompt sound corresponding to the target special vehicle.

2. The voice recognition method as described in claim 1, characterized in that, The target sound signal is subjected to feature detection processing to obtain its frequency domain feature information and time domain feature information, including: Perform a short-time Fourier transform on the target sound signal to obtain the frequency domain feature information corresponding to the target sound signal; and The target sound signal is subjected to time-domain modeling processing to obtain the time-domain feature information corresponding to the target sound signal.

3. The voice recognition method as described in claim 1 or 2, characterized in that, The first feature map information is integrated and segmented to obtain the second feature map information, including: The first feature map information is integrated and processed to obtain the spatial feature map information corresponding to the first feature map information. The spatial feature map information is segmented to obtain the second feature map information.

4. The voice recognition method as described in claim 3, characterized in that, The first feature map information is integrated and processed to obtain the spatial feature map information corresponding to the first feature map information, including: The first feature map information is encoded along the X direction corresponding to the first feature map information to obtain spatial feature map information along the X direction; and The first feature map information is encoded along the Y direction corresponding to the first feature map information to obtain spatial feature map information along the Y direction; The spatial feature map information along the X direction and the spatial feature map information along the Y direction are integrated and processed to obtain the spatial feature map information corresponding to the first feature map information.

5. The voice recognition method as described in claim 4, characterized in that, The spatial feature map information is segmented to obtain the second feature map information, including: The spatial feature map information is segmented to obtain a first sub-feature map information and a second sub-feature map information, which are used as the second feature map information. The first sub-feature map information and the second sub-feature map information are feature map information that include different spatial information.

6. The voice recognition method as described in claim 5, characterized in that, The first feature map information and the second feature map information are weighted to obtain the target feature map information, including: The first feature map information, the first sub-feature map information, and the second sub-feature map information are weighted to obtain the target feature map information.

7. The voice recognition method according to any one of claims 1-6, characterized in that, The target feature map information is processed to obtain the target recognition result, including: The target feature map information is processed to obtain a prompt sound recognition score as the target recognition result. If the prompt sound recognition score is greater than or equal to a score threshold, it means that the target sound signal includes the target prompt sound corresponding to the target special vehicle. If the prompt sound recognition score is less than the score threshold, it means that the target sound signal does not include the target prompt sound corresponding to the target special vehicle.

8. A voice recognition method, characterized in that, The method is applied to a target sound recognition model, which includes a sound detection module, a splicing and fusion module, a coordinate attention module, a weighting module, and a classification module. The sound detection module performs feature detection processing on the input target sound signal to obtain the frequency domain feature information and time domain feature information of the target sound signal. The frequency domain feature information and the time domain feature information are then input to the splicing and fusion module. The target sound signal is obtained by detecting the sound around the target vehicle. The splicing and fusion module performs splicing and fusion processing on the input frequency domain feature information and the time domain feature information to obtain the first feature map information corresponding to the target sound signal, and inputs the first feature map information to the coordinate attention module and the weighting module; The coordinate attention module integrates and segments the input first feature map information to obtain second feature map information, and inputs the second feature map information into the weighting module; The weighting module performs weighted processing on the input first feature map information and second feature map information to obtain target feature map information, and inputs the target feature map information into the classification module; The classification module performs recognition processing on the input target feature map information to obtain the target recognition result corresponding to the target sound signal. The target recognition result is used to indicate whether the target sound signal includes the target prompt sound corresponding to the target special vehicle.

9. A method for training a voice recognition model, characterized in that, The method includes: A first dataset is determined, which includes multiple training samples obtained by collecting sound samples from real road scenes; The initial sound recognition model is trained based on the first dataset until the loss function corresponding to the initial sound recognition model satisfies the target condition, thereby obtaining the target sound recognition model. The target sound recognition model is used to implement the sound recognition method as described in any one of claims 1-8, or to implement the sound recognition model training method as described in claim 9.

10. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer programs; The processor executes the computer program stored in the memory to enable the electronic device to implement the sound recognition method as described in any one of claims 1-8, or to implement the sound recognition model training method as described in claim 9.

Citation Information

Patent Citations

  • Alarm sound recognition method and device and training method of alarm sound recognition model

    CN117746890A

  • Abnormal sound detection method and device, storage medium and electronic equipment

    CN118280386A

  • Emergency vehicle detection method based on acoustic spectrum-time domain information fusion

    CN119724230A

  • Emergency Response Vehicle Detection for Autonomous Driving Applications

    US20220157165A1