Audio Data Processing Method, Device, and Medium for Identifying Ambient Sounds in the Wild

Through pre-trained identification model screening and data enhancement processing of audio samples, the problem of high difficulty in identifying sounds of power facilities in outdoor environments is solved, more accurate identification results are achieved, and safety hazards of power facilities are reduced.

CN114387991BActive Publication Date: 2025-07-08JINAN XINTONG ELECTRIC TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111416357.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-25
Publication Date
2025-07-08
Estimated Expiration
2041-11-25

AI Technical Summary

Technical Problem

In the wild environment, the recognition of sound collection samples of power facilities is difficult and the recognition accuracy is low, resulting in the failure to identify and manage safety hazards in a timely manner.

Method used

A pre-trained recognition model is used to filter environmental audio samples, process background audio samples through data augmentation, and integrate them with the environmental audio samples to form training samples, which are used to supervise the training of the field ambient sound recognition model.

Benefits of technology

It improves the accuracy of field ambient sound recognition, reduces the safety risks of power facilities, and ensures the normal operation of the power system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114387991B_ABST
    Figure CN114387991B_ABST
Patent Text Reader

Abstract

The present application discloses an audio data processing method, device and medium for identifying field environmental sounds. The method includes: obtaining an audio sample collected by a sound collection device; identifying the audio sample through a pre-trained identification model to obtain an identification result, and selecting the audio sample that meets the first specified type according to the identification result as the environmental audio sample. Select the audio sample that meets the second specified type according to the identification result, perform data augmentation processing on the audio sample that meets the second specified type to obtain a background audio sample, and integrate it as a training sample for the field environmental sound identification model to be trained. The number of the sample set is expanded, greatly enriching the diversity of the training samples. At the same time, through the recognition accuracy of the recognition model before and after data augmentation, this solution can obtain a more accurate recognition effect, ensuring that the recognition model can accurately recognize the environmental audio of power facilities and further reducing the safety hazards of power facilities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing technologies, and particularly to an audio data processing method, device, and medium for identifying ambient sounds in the wild. Background Art

[0002] With the development of the power system, more and more power facilities are being built. However, power facilities such as transmission lines and towers in outdoor scenarios often have various potential safety hazards, such as lightning strikes, icing, and bird damage. If the above potential safety hazards are not promptly noticed and effectively addressed, it will endanger the normal operation of the power system.

[0003] In the prior art, in order to accurately identify various potential safety hazards, workers often use sound recognition technology to identify and give early warnings of the above potential safety hazards. However, since power facilities such as transmission lines and towers are usually built in outdoor scenarios with complex background sounds, the collected sound samples have problems such as high recognition difficulty and low recognition accuracy. Summary of the Invention

[0004] To solve the above problems, that is, to solve the problems of high recognition difficulty and low recognition accuracy of the collected sound samples in the identification of ambient sounds in the wild, this application proposes an audio data processing method, device, and medium for identifying ambient sounds in the wild, including:

[0005] In a first aspect, this application proposes an audio data processing method for identifying ambient sounds in the wild, including: obtaining an audio sample collected by a sound collection device; identifying the audio sample through a pre-trained recognition model to obtain a recognition result, and selecting the audio sample that meets the first specified type as the ambient audio sample according to the recognition result, where the first specified type includes at least one of the following: thunderstorm sound, animal calls; selecting the audio sample that meets the second specified type according to the recognition result, where the second specified type includes at least one of the following: silence, wind sound, noisy sound; performing data augmentation processing on the audio sample that meets the second specified type to obtain a background audio sample, where the data augmentation processing method includes at least one of the following: waveform displacement, high-pitch transformation; integrating the ambient audio sample and the background audio sample as a training sample for a wild ambient sound recognition model to be trained.

[0006] In one example, the audio sample is recognized by a pre-trained recognition model to obtain a recognition result, and an audio sample that conforms to the first specified type is selected according to the recognition result as the environmental audio sample. Specifically, it includes: inputting the audio sample into the pre-trained recognition model to obtain a recognition result, where the recognition result includes the sound types corresponding to each segment of audio in the audio sample; according to the recognition result, for any one of the segments of audio in the audio sample, determining that the segment of audio contains more than three sound types and deleting the segment of audio; according to the recognition result, taking the remaining audio samples whose corresponding sound types conform to the first specified sound type as the environmental audio sample, and the first specified sound type includes at least one of the following: thunderstorm sound, animal calls.

[0007] In one example, after taking the remaining audio samples whose corresponding sound types conform to the first specified type as the environmental audio sample according to the recognition result, the method further includes: obtaining the position coordinates and operating time period of the sound collection device, and determining through query that there is a thunderstorm weather at the position coordinates during the operating time period; obtaining the occurrence time period of the thunderstorm weather and the audio samples during the occurrence time period; performing a decibel intensity detection on the audio samples during the occurrence time period to obtain a detection result, and according to the detection result, selecting those with a decibel intensity greater than the first preset threshold as the enhancement samples; adding the enhancement samples to the environmental audio samples.

[0008] In one example, before performing data enhancement processing on the audio samples that conform to the second specified type to obtain background audio samples, the method further includes: classifying the audio samples that conform to the second specified type according to the sound type according to the second specified type to obtain multiple groups of classified audio samples; for each group of the multiple groups of classified audio samples, obtaining multiple waveform signals corresponding to the classified audio samples in the group, and performing a mean processing on the multiple waveform signals corresponding to the classified audio samples to obtain processed audio samples, and taking the processed audio samples as the audio templates corresponding to the classified audio samples in the group; obtaining multiple groups of audio templates corresponding to the multiple groups of classified audio samples respectively, and replacing the audio samples that conform to the second specified type with the multiple groups of audio templates.

[0009] In one example, data augmentation processing is performed on the audio sample that conforms to the second specified type to obtain a background audio sample, which specifically includes: collecting the waveform signal of the audio sample that conforms to the second specified type, and within a preset time domain range, randomly rolling the waveform signal of the audio sample that conforms to the second specified type along the time axis corresponding to the time domain range to obtain a first processed audio sample; collecting the pitch value of the first processed audio sample, and performing a stretching process on the pitch value according to a second preset threshold to obtain a second processed audio sample; using the second processed audio sample as the background audio sample.

[0010] In one example, after data augmentation processing is performed on the audio sample that conforms to the second specified type to obtain a background audio sample, the method further includes: collecting a network audio sample that conforms to the second specified type according to the second specified type, where the network audio sample is an audio sample of the same type as the second specified type collected from the network; optimizing the background audio sample through the network audio sample of the second specified type, and the optimization methods at least include one of the following: audio splicing, audio superposition.

[0011] In one example, after integrating the environmental audio sample and the background audio sample as the training sample of the wild environmental sound recognition model to be trained, the method further includes: performing supervised training on the wild environmental sound recognition model to be trained through the training sample to obtain a first wild environmental sound recognition model; performing supervised training on the wild environmental sound recognition model to be trained through the audio sample to obtain a second wild environmental sound recognition model; obtaining accident case information of wild power facilities, and determining accident cause information and audio information before the accident according to the accident case information; inputting the audio information before the accident into the first wild environmental sound recognition model and the second wild environmental sound recognition model respectively to obtain a first recognition result and a second recognition result; comparing the first recognition result and the second recognition result with the accident cause information respectively to obtain a comparison result, and obtaining the recognition accuracy of the first wild environmental sound recognition model and the recognition accuracy of the second wild environmental sound recognition model according to the comparison result.

[0012] In one example, after comparing the first recognition result and the second recognition result with the accident cause information respectively to obtain a comparison result, and obtaining the recognition accuracy of the first field ambient sound recognition model and the recognition accuracy of the second field ambient sound recognition model according to the comparison result, the method further includes: determining whether the recognition accuracy of the first field ambient sound recognition model is greater than a preset multiple of the recognition accuracy of the second field ambient sound recognition model; if so, determining that the training effect corresponding to the training sample is qualified; if not, re-collecting the audio sample through the sound collection device, and expanding the training sample with the audio sample to obtain an expanded training sample; re-training the first field ambient sound recognition model with the expanded training sample.

[0013] On the other hand, the present application also provides an audio data processing device for recognizing field ambient sounds, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the following instructions: obtaining an audio sample collected by a sound collection device; recognizing the audio sample through a pre-trained recognition model to obtain a recognition result, and selecting an audio sample that conforms to a first specified type as an environmental audio sample according to the recognition result, the first specified type includes at least one of the following: thunderstorm sound, animal sound; selecting an audio sample that conforms to a second specified type according to the recognition result, the second specified type includes at least one of the following: silence, wind sound, noise; performing data enhancement processing on the audio sample that conforms to the second specified type to obtain a background audio sample, and the data enhancement processing method includes at least one of the following: waveform displacement, high pitch transformation; integrating the environmental audio sample and the background audio sample as a training sample for a to-be-trained field ambient sound recognition model.

[0014] On the other hand, the present application also provides a non-volatile computer storage medium storing computer-executable instructions, and the computer-executable instructions are configured to: obtain an audio sample collected by a sound acquisition device; identify the audio sample through a pre-trained recognition model to obtain a recognition result, and select, according to the recognition result, an audio sample that conforms to a first specified type as an environmental audio sample, where the first specified type includes at least one of the following: thunderstorm sound, animal calls; select, according to the recognition result, an audio sample that conforms to a second specified type, where the second specified type includes at least one of the following: silence, wind sound, noise; perform data augmentation processing on the audio sample that conforms to the second specified type to obtain a background audio sample, and the data augmentation processing method includes at least one of the following: waveform displacement, high-pitch transformation; integrate the environmental audio sample and the background audio sample as a training sample for training a wild environmental sound recognition model.

[0015] The audio data processing method, device, and medium for identifying wild environmental sounds proposed by the present application can bring the following beneficial effects: The wild environmental sound extraction and data augmentation technologies are adopted, and at the same time, an audio fusion method is proposed. The collected audio samples are directly processed, the number of the sample set is expanded, the diversity of the training samples is greatly enriched, and at the same time, through the recognition accuracy of the recognition model before and after data augmentation, this solution can obtain a more accurate recognition effect, ensure that the recognition model can accurately identify the environmental audio of power facilities, and can further reduce the safety hazards of power facilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0017] Figure 1 It is a flowchart of the audio data processing method for identifying wild environmental sounds in an embodiment of the present application;

[0018] Figure 2 It is a schematic diagram of the audio data processing device for identifying wild environmental sounds in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0020] First of all, it should be noted that an audio data processing method for identifying wild environmental sounds described in this application is in the form of a program and stored in a corresponding system or server according to the process designed in this application. The terminal or server where the system is located should be built-in with corresponding hardware devices, including but not limited to processors, memories, communication modules, etc., to implement the various technical solutions in this application. In the embodiments of this application, the system is taken as an example for explanation. The system is set in a corresponding terminal, and the terminal includes but not limited to: mobile phones, tablets, computers, or other terminal devices with corresponding computing power and functions. The system can determine its interaction relationship with the hardware devices in the terminal through corresponding programming languages, so as to call resources or transmit information to the various hardware devices in the terminal. In addition, the startup method of the system can be determined through corresponding software setting methods. The startup methods include but not limited to directly opening, opening through an APP, opening in the form of logging in to a corresponding WEB page, etc., to meet the user's operation, monitoring, or debugging of the system, and then implement the corresponding technical solutions described in this application.

[0021] The following will detail the technical solutions provided by each embodiment of this application in conjunction with the accompanying drawings.

[0022] As Figure 1 shown, an audio data processing method for identifying wild environmental sounds provided by an embodiment of this application includes:

[0023] S101: Obtain audio samples collected by a sound collection device.

[0024] Specifically, sound collection devices can be set on power facilities such as transmission wires and poles in outdoor scenarios. The sound collection devices are used to collect various audio data near the power facilities. In the embodiments of this application, the specific models of the sound collection devices are not specifically limited, and any sound collection device in the prior art that can meet the technical solutions of this application can be used.

[0025] Furthermore, by turning on the recording mode of the sound collection device, the audio data near the power facilities can be collected. In the embodiments of this application, for the convenience of making the final generated training sample set, audio samples in a unified format need to be collected.

[0026] By making corresponding settings to the backend program of the sound collection device, it is determined that the sound collection device collects audio samples at a sampling frequency of 44100, in a single channel, and in a coding format of 16-bit PCM. Through the above setting method, it is equivalent to obtaining an audio sample every ten seconds, that is, the duration of each audio sample is ten seconds.

[0027] According to the above settings, the system can obtain the audio samples collected by the sound collection device.

[0028] S102: Identify the audio sample using a pre-trained recognition model to obtain a recognition result, and select, according to the recognition result, the audio samples that conform to the first specified type as environmental audio samples, where the first specified type includes at least one of the following: thunderstorm sound, animal calls.

[0029] First of all, it should be noted that in the embodiments of the present application, millions of audio samples can be obtained through step S101. However, for the recognition task of wild environmental sounds, the number of audio samples that need to be concerned about is extremely small in the massive data. Manual screening is time-consuming and laborious, and relying solely on manual means to perceive the sound type is also prone to misjudgment. Therefore, a data screening scheme needs to be proposed.

[0030] To solve the above problems, the present application introduces a recognition model, that is, the recognition model can be used to identify millions of audio samples to quickly screen out the required type of audio samples.

[0031] In the embodiments of the present application, the pre-trained recognition model can adopt the open-source PANNs network model, and the PANNs network model can be trained through the large-scale audio dataset AudioSet. The large-scale audio dataset AudioSet contains 632 audio categories and more than two million artificially labeled audio categories, and each sound clip is 10 seconds long. These sound clips cover a wide range of sound types such as human and animal sounds, and daily environmental sounds.

[0032] Specifically, the system inputs the obtained audio samples into the above-mentioned pre-trained recognition model to obtain a recognition result, and the recognition result includes the sound types corresponding to each segment of the audio in the audio sample. Based on the fact that the pre-trained recognition model is trained through the AudioSet dataset, various sound types can be recognized within the range of audio categories included in the AudioSet dataset.

[0033] Furthermore, according to the recognition result, for any one of the segments of the audio in the audio sample, the system determines that the segment of the audio contains more than three sound types and deletes the segment of the audio. Since the number of sound types in this segment of the audio is too large and it has no reference value, deleting the audio with more than three sound types is beneficial to obtaining more accurate training samples.

[0034] Further, based on the recognition result, among the remaining audio samples, those with corresponding sound types conforming to the first specified sound type are used as environmental audio samples, where the first specified sound type includes at least one of the following: thunderstorm sound, animal calls. In the embodiments of the present application, thunderstorms and various animals are likely to pose certain hazards to power facilities. Therefore, it is particularly important to select audio samples of these two sound types. It should also be noted that in the embodiments of the present application, since birds have a greater impact and potential hazards on power facilities, audio samples containing various bird calls can be preferentially selected.

[0035] Further, after the system uses, as environmental audio samples, those with corresponding sound types conforming to the first specified type among the remaining audio samples according to the recognition result, the following technical solutions may also be included:

[0036] The system obtains the position coordinates and operating period of the sound collection device, and determines through query that there is a thunderstorm weather during the operating period at the position coordinates. Specifically, the query method can be set as: according to the position coordinates of the sound collection device, query the weather information near the position coordinates, and the coverage time period of the weather information is relatively long. In the embodiments of the present application, the weather information for the past fifteen days is collected; further, according to the operating period of the sound collection device, the weather information during the operating period is selected.

[0037] The system can determine that there is a thunderstorm weather based on the weather information during the operating period. That is, obtain the occurrence period of the thunderstorm weather and obtain the audio samples during the occurrence period.

[0038] At this time, it can be determined that there is an audio signal corresponding to the thunderstorm weather in the audio samples during the occurrence period. To ensure the accuracy of the samples, it is necessary to perform a decibel intensity detection on the audio samples during the occurrence period to obtain a detection result, and based on the detection result, select those with a decibel intensity greater than the first preset threshold as enhanced samples.

[0039] Further, the system adds the enhanced samples to the environmental audio samples. By adding the enhanced samples, the accuracy of the environmental audio samples can be further improved, and thus the training effect of the subsequent recognition model is improved.

[0040] S103: Select audio samples conforming to the second specified type according to the recognition result, where the second specified type includes at least one of the following: silence, wind sound, noise.

[0041] It should be noted that sound waves propagate through the air and reach the sound collection device. There are different propagation paths in different scenarios or spaces. Due to the complexity of the scenario or space, sound waves will undergo refraction, diffraction, and reflection, and finally be superimposed and collected by the sound collection device. Since the influencing factors of the scenario or space are diverse, any change in the relative position among the sound source, various obstacles in the scenario or space, and the sound collection device will cause changes in the collected audio samples, thereby affecting the manifestation form of the sound signal. Therefore, there are significant differences in the audio samples collected in different scenarios or spaces.

[0042] In the outdoor environment where power facilities are located, it is very difficult to identify and determine various sound sources that pose safety hazards due to their great differences and changes. Moreover, the outdoor environment is relatively open, and the propagation of sound waves will diverge. Therefore, the audio samples collected by the sound collection device are severely attenuated.

[0043] Therefore, to solve the above problems, background audio samples are added to this application. The background audio samples here can include sound signals such as silence, wind noise, and background noise that are likely to affect the audio samples. By adding background audio samples to train the recognition model, a higher recognition effect can be achieved.

[0044] Specifically, from the recognition results obtained in step S102, audio samples that meet the second specified type are selected. The second specified type here includes at least one of the following: silence, wind noise, and background noise.

[0045] S104: Perform data augmentation processing on the audio samples that meet the second specified type to obtain background audio samples. The data augmentation processing methods include at least one of the following: waveform displacement and pitch transformation.

[0046] Specifically, the system collects the waveform signals of the audio samples that meet the second specified type, and within a preset time domain range, randomly scrolls the waveform signals of the audio samples that meet the second specified type along the time axis corresponding to the time domain range to obtain the first processed audio sample. The random scrolling processing here is waveform displacement. The changing state of sound over time can be represented by the waveform signal, which is a common representation method of sound in the time domain. Randomly scrolling the waveform signal along the time axis generates a new signal different from the original signal. That is, within a ten-second range, the waveform signal is scrolled. For example, the waveform signal at the fifth second is scrolled to the first second, thereby enhancing the intensity of the waveform signal at the first second to enhance the waveform signal within the entire ten-second range.

[0047] Further, the system collects the pitch values of the first processed audio samples and performs a stretching process on the pitch values according to a second preset threshold to obtain second processed audio samples. Pitch is one of the three major characteristics of sound. Different from sound intensity and timbre, pitch refers to various sounds with different pitch levels, that is, the height of the sound. The second process here is pitch transformation. By stretching the pitch value, the pitch value of the original sound can be increased without affecting the sound speed. Therefore, the duration of the audio sample after pitch transformation will not change.

[0048] Further, the second processed audio samples are used as background audio samples. Due to various complex factors in the outdoor environment, there are problems such as intensity differences and low representativeness in the audio samples that meet the second specified type. In the embodiments of the present application, by introducing the first process and the second process, that is, waveform displacement and pitch change, the characteristic values of the audio samples can be effectively improved, and thus they are more easily recognized. At the same time, as training samples, better training effects can be achieved.

[0049] In addition, since the number of audio samples that meet the second specified type is also very large, it is difficult to perform data augmentation processing on a large number of audio samples that meet the second specified type. Therefore, before performing data augmentation processing, the present application has established a unique audio template for each type in the second specified type through corresponding technical solutions to reduce the number of audio samples that meet the second specified type.

[0050] Specifically, before performing data augmentation processing on the audio samples that meet the second specified type to obtain background audio samples, it further includes:

[0051] The system classifies the audio samples that meet the second specified type according to the sound type according to the second specified type to obtain multiple groups of classified audio samples. In the embodiments of the present application, since the second specified type includes three categories, three groups of classified audio samples are obtained. Among them, each group of classified audio samples includes multiple audio samples.

[0052] Further, for each group of the multiple groups of classified audio samples, the system obtains multiple waveform signals corresponding to the group of classified audio samples, and performs a mean value process on the multiple waveform signals corresponding to the classified audio samples to obtain a processed audio sample, which is used as the uniquely corresponding mean value processed audio sample for the group of classified audio samples and used as the audio template corresponding to the group of classified audio samples. The mean value process here is to superimpose multiple waveform signals and take the mean value of each time point according to the number of waveform signals to finally generate a uniquely corresponding audio sample.

[0053] Further, the system obtains multiple groups of audio templates corresponding to the classified audio samples, that is, three audio templates, and replaces the audio samples that meet the specified type with the multiple groups of audio templates.

[0054] In addition, after performing data augmentation processing on the audio samples that meet the second specified type to obtain background audio samples, it further includes:

[0055] The system collects network audio samples that meet the second specified type according to the second specified type, where the network audio samples are audio samples of the same type as the second specified type collected from the network.

[0056] Furthermore, the system optimizes the background audio samples through the network audio samples of the second specified type, and the optimization methods at least include one of the following: audio splicing, audio superposition.

[0057] Specifically, it further improves the representativeness of the background audio samples. In this application, the background audio samples are optimized by using network audio samples.

[0058] Audio splicing refers to splicing the network audio samples and the background audio samples of the corresponding same type on the time domain axis, that is, increasing the time domain length of the background audio samples, further increasing the number of features in the background audio samples, and improving the subsequent training effect.

[0059] Audio superposition refers to mixing and superposing the network audio samples and the background audio samples of the corresponding same type within the time domain range. This kind of processing can increase the waveform signal strength at the time points with lower feature values in the background audio samples.

[0060] Through the above optimization processing, richer background audio samples can be obtained, which not only expands the number of samples in the sample set, but also expands the diversity of the samples. Through the optimization processing, the recognition accuracy can be effectively improved, and at the same time, the training effect of the model is improved.

[0061] S105: Integrate the environmental audio samples and the background audio samples as the training samples for the wild environmental sound recognition model to be trained.

[0062] After the system integrates the environmental audio samples and the background audio samples as the training samples for the wild environmental sound recognition model to be trained, it further includes:

[0063] Supervise and train the wild environmental sound recognition model to be trained through the training samples to obtain the first wild environmental sound recognition model.

[0064] Supervise and train the wild environmental sound recognition model to be trained through audio samples to obtain the second wild environmental sound recognition model. It should be noted that the audio samples here are the audio samples collected by the above-mentioned sound collection device without any processing.

[0065] To verify the recognition effect of the first wild environmental sound recognition model, the present application introduces a corresponding verification scheme, which specifically includes:

[0066] The system obtains accident case information of wild power facilities, and determines accident cause information and audio information before the accident according to the accident case information.

[0067] Input the audio information before the accident into the first wild environmental sound recognition model and the second wild environmental sound recognition model respectively to obtain a first recognition result and a second recognition result. It should be noted that there can be multiple pieces of the above-mentioned accident case information, and the first recognition result and the second recognition result here can also include multiple recognition results.

[0068] Compare the first recognition result and the second recognition result with the accident cause information respectively to obtain a comparison result, and obtain the recognition accuracy rate of the first wild environmental sound recognition model and the recognition accuracy rate of the second wild environmental sound recognition model according to the comparison result.

[0069] Furthermore, determine whether the recognition accuracy rate of the first wild environmental sound recognition model is greater than a preset multiple of the recognition accuracy rate of the second wild environmental sound recognition model.

[0070] In the embodiment of the present application, the recognition accuracy rate of the first wild environmental sound recognition model should be 20% higher than the recognition accuracy rate of the second wild environmental sound recognition model.

[0071] If so, determine that the training effect corresponding to the training sample is qualified.

[0072] If not, re-collect audio samples through the sound collection device, and expand the training sample through the audio samples to obtain an expanded training sample.

[0073] Retrain the first wild environmental sound recognition model with the expanded training sample until the recognition accuracy rate of the first wild environmental sound model is 20% higher than the recognition accuracy rate of the second wild environmental sound recognition model.

[0074] In one embodiment, as Figure 2 shown, the present application also provides an audio data processing device for recognizing wild environmental sounds, including:

[0075] At least one processor; and,

[0076] A memory communicatively connected to the at least one processor; wherein,

[0077] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the following instructions:

[0078] Obtain an audio sample collected by a sound collection device;

[0079] Identify the audio sample through a pre-trained recognition model to obtain a recognition result, and select an audio sample that conforms to a first specified type as an environmental audio sample according to the recognition result. The first specified type includes at least one of the following: thunderstorm sound, animal call;

[0080] Select an audio sample that conforms to a second specified type according to the recognition result. The second specified type includes at least one of the following: silence, wind sound, noise;

[0081] Perform data augmentation processing on the audio sample that conforms to the second specified type to obtain a background audio sample. The data augmentation processing methods include at least one of the following: waveform displacement, high-pitched transformation;

[0082] Integrate the environmental audio sample and the background audio sample as a training sample for a wild environmental sound recognition model to be trained.

[0083] In one embodiment, the present application also proposes a non-volatile computer storage medium storing computer-executable instructions, and the computer-executable instructions are set as follows:

[0084] Obtain an audio sample collected by a sound collection device;

[0085] Identify the audio sample through a pre-trained recognition model to obtain a recognition result, and select an audio sample that conforms to a first specified type as an environmental audio sample according to the recognition result. The first specified type includes at least one of the following: thunderstorm sound, animal call;

[0086] Select an audio sample that conforms to a second specified type according to the recognition result. The second specified type includes at least one of the following: silence, wind sound, noise;

[0087] Perform data augmentation processing on the audio sample that conforms to the second specified type to obtain a background audio sample. The data augmentation processing methods include at least one of the following: waveform displacement, high-pitched transformation;

[0088] Integrate the environmental audio sample and the background audio sample as a training sample for a wild environmental sound recognition model to be trained.

[0089] Each embodiment in this application is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for relevant content.

[0090] The devices and media provided in the embodiments of this application correspond one-to-one with the methods. Therefore, the devices and media also have beneficial technical effects similar to those of their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be elaborated here.

[0091] Those skilled in the art should understand that the embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0092] This application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0093] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0094] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 steps of the functions specified in one block or multiple blocks.

[0095] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0096] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium.

[0097] Computer-readable media include both permanent and non-permanent, removable and non-removable media and can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0098] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, commodity or device including the element.

[0099] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various modifications and changes can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. An audio data processing method for identifying outdoor environmental sounds, characterized in that, Including: Obtaining audio samples collected by a sound acquisition device; Identifying the audio samples through a pre-trained recognition model to obtain a recognition result, and selecting the audio samples that conform to the first specified type as environmental audio samples according to the recognition result. Specifically, it includes: inputting the audio samples into the pre-trained recognition model to obtain a recognition result, where the recognition result includes the sound types corresponding to each segment of audio in the audio samples; according to the recognition result, for any one of the segments of audio in the audio samples, determining that the segment of audio contains more than three sound types and deleting the segment of audio; according to the recognition result, taking the remaining audio samples whose corresponding sound types conform to the first specified sound type as environmental audio samples; the first specified type includes at least one of the following: thunderstorm sound, animal calls; Obtaining the position coordinates and operating period of the sound acquisition device, and determining through query that there is a thunderstorm weather during the operating period at the position coordinates; obtaining the occurrence period of the thunderstorm weather, and obtaining the audio samples during the occurrence period; performing decibel intensity detection on the audio samples during the occurrence period to obtain a detection result, and according to the detection result, selecting those with a decibel intensity greater than the first preset threshold as enhanced samples; adding the enhanced samples to the environmental audio samples; Selecting the audio samples that conform to the second specified type according to the recognition result, where the second specified type includes at least one of the following: silence, wind sound, noisy sound; Classifying the audio samples that conform to the second specified type according to the sound type to obtain multiple groups of classified audio samples; for each group of the multiple groups of classified audio samples, obtaining multiple waveform signals corresponding to the classified audio samples in the group, and performing mean processing on the multiple waveform signals corresponding to the classified audio samples to obtain processed audio samples, and taking the processed audio samples as the audio templates corresponding to the classified audio samples in the group; obtaining multiple groups of audio templates corresponding to the multiple groups of classified audio samples respectively, and replacing the audio samples that conform to the second specified type with the multiple groups of audio templates; Performing data augmentation processing on the audio samples that conform to the second specified type to obtain background audio samples, and the data augmentation processing methods include at least one of the following: waveform displacement, pitch transformation; where the pitch transformation is to perform a stretching process on the pitch value; Integrating the environmental audio samples and the background audio samples as the training samples for the wild environmental sound recognition model to be trained.

2. The audio data processing method for identifying field ambient sounds according to claim 1, wherein Performing data augmentation processing on the audio samples that conform to the second specified type to obtain background audio samples, specifically including: Collecting the waveform signals of the audio samples that conform to the second specified type, and randomly rolling the waveform signals of the audio samples that conform to the second specified type along the time axis corresponding to the preset time domain range within the preset time domain range to obtain the first processed audio samples; Collect the pitch value of the first processed audio sample, and perform a stretching process on the pitch value according to a second preset threshold to obtain a second processed audio sample; Use the second processed audio sample as the background audio sample.

3. The audio data processing method for identifying outdoor environmental sounds according to claim 1, wherein After performing data augmentation on the audio sample that conforms to the second specified type to obtain a background audio sample, the method further includes: Collect network audio samples that conform to the second specified type according to the second specified type. The network audio samples are audio samples of the same type as the second specified type collected from the network; Optimize the background audio sample with the network audio samples of the second specified type. The optimization methods at least include one of the following: audio splicing, audio superposition.

4. The audio data processing method for identifying wild environmental sounds according to claim 1, characterized in that After integrating the environmental audio sample and the background audio sample as the training sample of the wild environmental sound recognition model to be trained, the method further includes: Perform supervised training on the wild environmental sound recognition model to be trained with the training sample to obtain a first wild environmental sound recognition model; Perform supervised training on the wild environmental sound recognition model to be trained with the audio sample to obtain a second wild environmental sound recognition model; Obtain the accident case information of the wild power facilities, and determine the accident cause information and the audio information before the accident according to the accident case information; Input the audio information before the accident into the first wild environmental sound recognition model and the second wild environmental sound recognition model respectively to obtain a first recognition result and a second recognition result; Compare the first recognition result and the second recognition result with the accident cause information respectively to obtain a comparison result, and obtain the recognition accuracy of the first wild environmental sound recognition model and the recognition accuracy of the second wild environmental sound recognition model according to the comparison result.

5. The audio data processing method for identifying wild environmental sounds according to claim 4, characterized in that, After comparing the first recognition result and the second recognition result with the accident cause information respectively to obtain a comparison result, and obtaining the recognition accuracy of the first wild environmental sound recognition model and the recognition accuracy of the second wild environmental sound recognition model according to the comparison result, the method further includes: Determine whether the recognition accuracy of the first wild environmental sound recognition model is greater than a preset multiple of the recognition accuracy of the second wild environmental sound recognition model; If so, determine that the training effect corresponding to the training sample is qualified; If not, re-collect the audio sample through the sound collection device, and expand the training sample with the audio sample to obtain an expanded training sample; Re-train the first wild environmental sound recognition model with the expanded training sample.

6. An audio data processing device for identifying outdoor environmental sounds, characterized in that, Include: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor. The instructions are executed by the at least one processor so that the at least one processor can execute the following instructions: the audio data processing method for recognizing wild environmental sounds described in claim 1.

7. A non-volatile computer storage medium stores computer-executable instructions, characterized in that, The computer-executable instructions are set to: the audio data processing method for identifying field environmental sounds described in claim 1.

Citation Information

Patent Citations

  • Speech data augmentation

    US20200335086A1