Abnormal audio recognition method and device and electronic equipment

CN120220698APending Publication Date: 2025-06-27XIAN LONGXING INTELLIGENT PATROL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510298577.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the prior art, when using deep neural networks to identify audio data collected by patrol robots, it is difficult to accurately identify abnormal audio data, because the number of abnormal samples is small, resulting in training tending to normal sample classification.

Method used

By obtaining the audio data collected by the inspection robot in real time, calculating its energy spectrum density signal and removing background noise, a feature map is obtained. Then, the feature map is compared with the pre-generated background map, a difference map is generated, and whether the audio is abnormal audio is based on the difference map. The background image is pre-generated based on feature map differences in normal audio datasets.

Benefits of technology

This method can accurately identify abnormal audio data, avoid bias problems caused by the small number of abnormal samples during deep neural network training, and improve the accuracy of abnormal audio recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220698A_ABST
    Figure CN120220698A_ABST
Patent Text Reader

Abstract

The invention discloses an abnormal audio identification method and apparatus, and an electronic device. The method comprises the steps of obtaining first audio data collected by an inspection robot performing inspection according to a target route in real time; under the condition that the audio frequency corresponding to the currently collected first audio data is greater than or equal to the audio frequency threshold value corresponding to the current road section, obtaining a first feature map corresponding to the current first audio data; obtaining a first difference image of the first feature image relative to the background image; based on the first difference chart, determining whether the audio contained in the current first audio data is abnormal audio; wherein the background image is pre-generated based on the difference between the feature maps corresponding to the normal audio data in the normal audio data set of the current road section. By applying the technical scheme provided by the invention, the abnormal audio can be accurately identified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer application technologies, and in particular, to an abnormal audio recognition method, device, and electronic device. Background Art

[0002] With the development of industrial automation and intelligence, inspection robots have been widely used in the inspection of industrial equipment. Controlling the inspection robot to collect audio data in complex and harsh environments, such as high temperature, high humidity, strong noise, etc., the operating state of the equipment can be judged according to the audio data collected by the inspection robot, so as to complete the equipment operating state detection task, and the inspection efficiency and safety are relatively high.

[0003] Currently, the audio data collected by the inspection robot is mainly analyzed and recognized through a deep neural network. The deep neural network requires a large amount of positive and negative sample data for training. However, in actual applications, the number of abnormal samples is small, resulting in the deep neural network training being biased towards normal sample classification and it being difficult to accurately identify abnormal audio data. Summary of the Invention

[0004] The purpose of the present application is to provide an abnormal audio recognition method, device, and electronic device to accurately recognize abnormal audio.

[0005] To solve the above technical problems, the present application provides the following technical solutions:

[0006] In a first aspect, an abnormal audio recognition method is provided, including:

[0007] Obtain first audio data collected in real time by an inspection robot patrolling along a target route, where the target route includes multiple sections, and each section corresponds to its own audio frequency threshold;

[0008] When the audio frequency corresponding to the currently collected first audio data is greater than or equal to the audio frequency threshold corresponding to the current section, obtain a first feature map corresponding to the current first audio data;

[0009] Compare the first feature map with the background map of the current section to obtain a first difference map of the first feature map relative to the background map;

[0010] Based on the first difference map, determine whether the audio included in the current first audio data is abnormal audio;

[0011] Wherein, the background map is pre-generated based on the differences between the feature maps corresponding to the normal audio data in the normal audio data set of the current section.

[0012] Optionally, the obtaining of the first feature map corresponding to the current first audio data includes:

[0013] Calculate the energy spectral density signal of the current first audio data;

[0014] Remove background noise from the current first audio data according to the energy spectral density signal of the current first audio data to obtain second audio data;

[0015] Perform feature extraction and post-processing on the second audio data to obtain a first feature map corresponding to the current first audio data.

[0016] Optionally, the performing feature extraction and post-processing on the second audio data to obtain a first feature map corresponding to the current first audio data includes:

[0017] Perform feature extraction on the second audio data to obtain a first feature;

[0018] Perform probability distribution estimation, normalization processing, and visualization processing on the first feature in sequence to obtain a first feature map corresponding to the current first audio data.

[0019] Optionally, the comparing the first feature map with the background map of the current road section to obtain a first difference map of the first feature map relative to the background map includes:

[0020] Compare the first feature map with the background map of the current road section to generate a binary image;

[0021] Perform dimensionality reduction processing on the binary image to obtain a first difference map of the first feature map relative to the background map.

[0022] Optionally, the determining whether the audio included in the current first audio data is abnormal audio based on the first difference map includes:

[0023] Determine the cosine distance between every two difference points in the first difference map;

[0024] Determine the difference value between the first feature map and the background map based on the cosine distance between every two difference points in the first difference map;

[0025] Determine whether the audio included in the current first audio data is abnormal audio according to the magnitude relationship between the difference value and a preset difference threshold.

[0026] Optionally, the determining whether the audio included in the current first audio data is abnormal audio according to the magnitude relationship between the difference value and a preset difference threshold includes:

[0027] In the case where the difference value is greater than or equal to the preset difference threshold, determine that the audio included in the current first audio data is suspicious audio;

[0028] Determine whether the audio included in the current first audio data is abnormal audio according to the confirmation information of the suspicious audio;

[0029] In the case that the difference value is less than the difference threshold, determine that the audio included in the current first audio data is normal audio.

[0030] Optionally, the method further includes:

[0031] In the case that it is determined that the audio included in the current first audio data is abnormal audio, add the current first audio data to the abnormal audio dataset for abnormal processing based on the abnormal audio data in the abnormal audio dataset;

[0032] In the case that it is determined that the audio included in the current first audio data is normal audio, add the current first audio data to the normal audio dataset of the current section to regenerate the background map of the current section based on the differences between the feature maps corresponding to the normal audio data in the normal audio dataset.

[0033] Optionally, obtain the audio frequency threshold corresponding to each section through the following steps:

[0034] In the case that it is determined that all devices are operating normally, obtain the third audio data collected by the inspection robot in real time according to the target route;

[0035] For each section, in the case that there is third audio data with an audio frequency greater than or equal to the initial audio frequency threshold corresponding to the current section, determine the maximum audio frequency corresponding to the third audio data of the current section as the audio frequency threshold corresponding to the current section;

[0036] In the case that there is no third audio data with an audio frequency greater than or equal to the audio frequency initial threshold in the current section, determine the audio frequency initial threshold as the audio frequency threshold corresponding to the current section.

[0037] Optionally, generate the background map of the current section through the following steps:

[0038] Determine the set composed of the third audio data collected in the current section as the normal audio dataset of the current section;

[0039] For each normal audio data in the normal audio dataset, calculate the energy spectral density signal of the current normal audio data;

[0040] Remove background noise from the current normal audio data according to the energy spectral density signal of the current normal audio data to obtain the fourth audio data;

[0041] Extract features from and post-process the fourth audio data to obtain a second feature map corresponding to the current fourth audio data;

[0042] Generate a background map for the current road section based on the differences between the second feature maps corresponding to each fourth audio data.

[0043] In a second aspect, an abnormal audio recognition device is provided, including:

[0044] A first acquisition module, configured to acquire first audio data collected in real time by an inspection robot inspecting according to a target route, where the target route includes multiple road sections, and each road section corresponds to its own audio frequency threshold;

[0045] A second acquisition module, configured to acquire a first feature map corresponding to the current first audio data when the audio frequency corresponding to the currently acquired first audio data is greater than or equal to the audio frequency threshold corresponding to the current road section;

[0046] An obtaining module, configured to compare the first feature map with the background map of the current road section to obtain a first difference map of the first feature map relative to the background map;

[0047] A determination module, configured to determine whether the audio included in the current first audio data is abnormal audio based on the first difference map;

[0048] Wherein, the background map is pre-generated based on the differences between the feature maps corresponding to the normal audio data in the normal audio data set of the current road section.

[0049] In a third aspect, an electronic device is provided, including:

[0050] A memory, configured to store a computer program;

[0051] A processor, configured to implement the steps of the abnormal audio recognition method as described in the first aspect when executing the computer program.

[0052] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the abnormal audio recognition method as described in the first aspect are implemented.

[0053] In a fifth aspect, a computer program product is provided, where the computer program product includes computer instructions, and the computer instructions are stored in a computer-readable storage medium and are adapted to be read and executed by a processor so that a computer device having the processor executes the steps of the abnormal audio recognition method as described in the first aspect.

[0054] Applying the technical solution provided by the embodiments of the present application, the background image of each section is pre-generated based on the differences between the feature maps corresponding to the normal audio data in the normal audio dataset of the corresponding section, so that the background image of each section can accurately reflect the characteristics of the normal audio data of the corresponding section. Furthermore, for each section, according to the first difference map of the first feature map corresponding to the first audio data with a relatively large audio frequency in the currently collected audio data of the inspection robot with respect to the background image of the current section, it can be accurately determined whether the audio contained in the first audio data is abnormal audio. The number of normal audio data is large, and it is easy to generate the background image of each section based on the normal audio data without the need for training of a deep neural network, which can effectively avoid the problem that when training a deep neural network, due to the small number of abnormal samples, the training tends to classify normal samples and it is difficult to accurately identify abnormal audio data.

[0055] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0057] Figure 1 It is a flowchart of the implementation of an abnormal audio recognition method in the embodiments of the present application;

[0058] Figure 2 It is a schematic diagram of an inspection process in the embodiments of the present application;

[0059] Figure 3 It is a schematic diagram of the feature map corresponding to the first audio data in the embodiments of the present application;

[0060] Figure 4 It is a schematic diagram of the background image of the current section in the embodiments of the present application;

[0061] Figure 5 It is a schematic diagram of the binary image of the feature map corresponding to the first audio data with respect to the background image of the current section in the embodiments of the present application;

[0062] Figure 6 It is a schematic diagram of the determination process of the audio frequency threshold corresponding to each section in the embodiments of the present application;

[0063] Figure 7 It is a schematic diagram of the structure of an abnormal audio recognition system in the embodiments of the present application;

[0064] Figure 8 It is a schematic structural diagram of an abnormal audio recognition device in an embodiment of the present application;

[0065] Figure 9 It is a schematic structural diagram of an electronic device in an embodiment of the present application. Specific embodiments

[0066] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.

[0067] The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are usually of the same type, and the number of objects is not limited. For example, the first object can be one or multiple.

[0068] The core of the present application is to provide an abnormal audio recognition method, which can be applied to scenarios for detecting the operating state of a device.

[0069] See Figure 1 As shown, it is an implementation flowchart of an abnormal audio recognition method provided by an embodiment of the present application, and the method may include the following steps:

[0070] S110: Obtain first audio data collected in real time by a patrol robot patrolling along a target route.

[0071] The target route includes multiple sections, and each section corresponds to its own audio frequency threshold.

[0072] In the embodiments of the present application, it is necessary to patrol the device to judge the operating state of the device. If the device is in a normal operating state, no treatment is required. If the device is in an abnormal state, abnormal treatment needs to be carried out as soon as possible.

[0073] The target route can be pre-planned. The target route includes multiple sections, and each section corresponds to its own audio frequency threshold. The audio frequency thresholds corresponding to different sections may be the same or different. For example, the target route includes section A, section B, and section C. Section A is in area A, section B is in area B, and section C is in area C. The types of devices deployed in different areas may be different, and the generated audio frequencies and sizes may be different.

[0074] During the process of inspecting the equipment, the inspection robot can be controlled to walk along the target route. When the inspection robot walks along the target route, it can collect the first audio data in real time. When the inspection robot walks along the target route, it can save the audio data according to a preset first duration. For example, the audio data is saved every 6s. Therefore, it can be understood that there are multiple pieces of first audio data, and each piece of first audio data is the audio data for a continuous first duration. Each section corresponds to multiple pieces of first audio data.

[0075] By interacting with the inspection robot, the first audio data collected in real time by the inspection robot during the inspection along the target route can be obtained.

[0076] S120: In the case where the audio frequency corresponding to the currently collected first audio data is greater than or equal to the audio frequency threshold corresponding to the current section, obtain the first feature map corresponding to the current first audio data.

[0077] After obtaining the first audio data collected in real time by the inspection robot, the currently collected first audio data can be analyzed to determine the audio frequency corresponding to the current first audio data. Compare the audio frequency corresponding to the current first audio data with the audio frequency threshold corresponding to the current section. If the audio frequency corresponding to the current first audio data is relatively large, greater than or equal to the audio frequency threshold corresponding to the current section, it can be considered that the audio contained in the current first audio data may be abnormal, and the first feature map corresponding to the current first audio data can be obtained.

[0078] The current section is the section where the collection location of the current first audio data is located, and the current first audio data is the currently collected first audio data.

[0079] It should be noted that in the embodiments of the present application, if the audio data is abnormal, or the audio data is abnormal audio data, it can be understood that the audio contained in the audio data is abnormal, or the audio contained in the audio data is abnormal audio, or the equipment deployed in the area where the section corresponding to the audio data is located is abnormal.

[0080] S130: Compare the first feature map with the background map of the current section to obtain the first difference map of the first feature map relative to the background map.

[0081] The background map is pre-generated based on the differences between the feature maps corresponding to the normal audio data in the normal audio data set of the current section.

[0082] In the embodiments of the present application, the normal audio data of each section of the target route can be collected in advance to form a normal audio data set. For each section, a background map of the current section can be generated based on the differences between the feature maps corresponding to the normal audio data in the normal audio data set of the current section. The background map of the current section can be regarded as a comprehensive representation of the audio features of the devices deployed in the area where the current section is located under normal working conditions.

[0083] After obtaining the first feature map corresponding to the current first audio data, the first feature map can be compared with the background map of the current section to obtain a first difference map of the first feature map relative to the background map. The first difference map can represent the difference between the features of the current first audio data and the normal audio features of the current section.

[0084] S140: Based on the first difference map, determine whether the audio included in the current first audio data is abnormal audio.

[0085] After comparing the first feature map corresponding to the currently collected first audio data with the background map of the current section to obtain the first difference map of the first feature map relative to the background map, further based on the first difference map, it can be determined whether the audio included in the current first audio data is abnormal audio.

[0086] For each obtained first audio data, the operations of the above steps S120 - S140 can be performed to determine whether the audio included in each first audio data is abnormal audio.

[0087] Applying the method provided by the embodiments of the present application, the background map of each section is pre - generated based on the differences between the feature maps corresponding to the normal audio data in the normal audio data set of the corresponding section, so that the background map of each section can accurately reflect the features of the normal audio data of the corresponding section. Furthermore, for each section, according to the first difference map of the first feature map corresponding to the first audio data with a relatively large audio frequency of the current section collected by the inspection robot relative to the background map of the current section, it can be accurately determined whether the audio included in the first audio data is abnormal audio. The number of normal audio data is large, and it is easy to generate the background map of each section based on the normal audio data without the need for training of a deep neural network, which can effectively avoid the problem that when training a deep neural network, due to the small number of abnormal samples, it tends to classify normal samples, making it difficult to accurately identify abnormal audio data.

[0088] In some embodiments of the present application, obtaining the first feature map corresponding to the current first audio data may include the following steps:

[0089] Calculate the energy spectral density signal of the current first audio data;

[0090] Remove the background noise from the current first audio data according to the energy spectral density signal of the current first audio data to obtain the second audio data;

[0091] Perform feature extraction and post-processing on the second audio data to obtain the first feature map corresponding to the current first audio data.

[0092] For the convenience of description, the above steps are combined for description.

[0093] In the embodiment of the present application, if the audio frequency corresponding to the currently collected first audio data is greater than or equal to the audio frequency threshold corresponding to the current road section, it can be considered that the audio contained in the current first audio data may be abnormal, and the energy spectral density signal of the current first audio data can be calculated first. The energy spectral density is also called the energy spectrum, which describes how the energy of a signal or time series is distributed with frequency. The energy spectrum is the square of the Fourier transform of the original signal.

[0094] Separate the background audio according to the energy spectral density signal of the current first audio data, and remove the background noise from the current first audio data to obtain the second audio data. Because there is strong background noise in some working environments, which will interfere with the recognition of abnormal audio. For example, in underground mines, high humidity and high dust environments may cause sensor failures, and strong background noise will cover up the abnormal sounds of equipment. Therefore, in the embodiment of the present application, the background noise of the audio data collected by the inspection robot is removed first to reduce the interference of background noise on the recognition of abnormal sounds of equipment.

[0095] Perform feature extraction and post-processing on the second audio data to obtain the first feature map corresponding to the current first audio data. The first feature map can be considered as the feature representation of the current first audio data.

[0096] Optionally, feature extraction can be performed on the second audio data to obtain the first feature, and then probability distribution estimation, normalization processing, and visualization processing are sequentially performed on the first feature to obtain the first feature map corresponding to the current first audio data.

[0097] Among them, the first feature obtained by performing feature extraction on the second audio data may include Mel-Frequency Cepstral Coefficients (MFCC) features. MFCC feature extraction is a method of audio feature extraction, mainly used in fields such as speech processing and music information retrieval, and can effectively capture important characteristics such as the pitch and timbre of sound. The probability distribution estimation performed on the first feature may include delta processing. The delta method is a method in mathematical statistics that uses Taylor expansion to approximate a function and makes inferences based on this, mainly used to estimate the statistical characteristics of parameters such as the mean and variance.

[0098] By performing the above processing on the current first audio data, the first feature map corresponding to the current first audio data can be accurately obtained.

[0099] In some embodiments of the present application, comparing the first feature map with the background map of the current road section to obtain the first difference map of the first feature map relative to the background map may include the following steps:

[0100] Compare the first feature map with the background map of the current road section to generate a binary image;

[0101] Perform dimensionality reduction processing on the binary image to obtain the first difference map of the first feature map relative to the background map.

[0102] For ease of description, the above steps are combined for explanation.

[0103] In the embodiments of the present application, when the audio frequency corresponding to the currently collected first audio data is greater than or equal to the audio frequency threshold corresponding to the current road section, after obtaining the first feature map corresponding to the current first audio data, the first feature map can be compared with the background map of the current road section to calculate the difference between the first feature map and the background map, and a binary image is formed. The first color points in the binary image can represent the difference points between the first feature map and the background map, and the second color points can represent the same points between the first feature map and the background map. The first color can be white, and the second color can be black.

[0104] After generating the binary image, dimensionality reduction processing can be performed on the binary image to obtain the first difference map of the first feature map relative to the background map. For example, after performing dimensionality reduction processing on the binary image, a first difference map of 80*80 is obtained.

[0105] Performing dimensionality reduction processing on the binary image helps to reduce the amount of calculation, reduce redundant features, and reduce the data dimension.

[0106] In some embodiments of the present application, based on the first difference map, determining whether the audio included in the current first audio data is abnormal audio includes:

[0107] Determine the cosine distance between every two difference points in the first difference map;

[0108] Based on the cosine distance between every two difference points in the first difference map, determine the difference value between the first feature map and the background map;

[0109] According to the magnitude relationship between the difference value and the preset difference threshold, determine whether the audio included in the current first audio data is abnormal audio.

[0110] For ease of description, the above steps are combined for explanation.

[0111] In an embodiment of the present application, when the audio frequency corresponding to the currently collected first audio data is greater than or equal to the audio frequency threshold corresponding to the current road section, after obtaining the first difference map of the first feature map corresponding to the current first audio data relative to the background map of the current road section, the coordinates of each difference point in the first difference map can be determined, and then, based on the coordinates of each difference point, the cosine distance between every two difference points can be determined.

[0112] Based on the cosine distance between every two difference points in the first difference map, the difference value between the first feature map and the background map can be determined. Optionally, the average value or median value or sum value of the cosine distances between every two difference points in the first difference map can be determined as the difference value between the first feature map and the background map.

[0113] A difference threshold can be preset. After determining the difference value between the first feature map and the background map, the difference value is compared with the difference threshold. According to the magnitude relationship between the difference value and the difference threshold, it can be determined whether the audio included in the current first audio data is an abnormal different frequency.

[0114] If the difference value is greater than or equal to the difference threshold, it is considered that the difference between the first feature map and the background map is large. The background map is pre-generated based on the differences between the feature maps corresponding to the normal audio data in the normal audio dataset of the current road section. If the difference between the first feature map and the background map is large, it can be considered that the current first audio data corresponding to the first feature map has a large difference from the normal audio data, and the possibility that the audio included in the current first audio data is abnormal audio is high.

[0115] If the difference value is less than the difference threshold, it is considered that the difference between the first feature map and the background map is small. The current first audio data corresponding to the first feature map has a small difference from the normal audio data, and the possibility that the audio included in the current first audio data is abnormal audio is low.

[0116] Based on the cosine distance between every two difference points in the first difference map, the difference value between the first feature map and the background map can be accurately determined, and then, according to the magnitude relationship between the difference value and the difference threshold, it can be accurately determined whether the audio included in the current first audio data is abnormal audio.

[0117] In some embodiments of the present application, determining whether the audio included in the current first audio data is abnormal audio according to the magnitude relationship between the difference value and the preset difference threshold may include the following steps:

[0118] When the difference value is greater than or equal to the preset difference threshold, it is determined that the audio included in the current first audio data is suspicious audio;

[0119] According to the confirmation information of the suspicious audio, it is determined whether the audio included in the current first audio data is abnormal audio;

[0120] When the difference value is less than the difference threshold, it is determined that the audio included in the current first audio data is normal audio.

[0121] For convenience of description, the above steps are combined and described.

[0122] In the embodiment of the present application, after determining the difference value between the first feature map corresponding to the current first audio data and the background map of the current road section, the difference value is compared with the difference threshold. If the difference value is greater than or equal to the difference threshold, the audio included in the current first audio data can be determined as suspicious audio. Whether the audio included in the current first audio data is abnormal audio needs to be further confirmed.

[0123] Optionally, the suspicious audio can be listened to manually to confirm whether it is abnormal audio. According to the confirmation information of the suspicious audio, it is determined whether the audio included in the current first audio data is abnormal audio. Optionally, if it is manually confirmed that the suspicious audio is abnormal audio, it can be determined that the audio included in the current first audio data is abnormal audio. If it is manually confirmed that the suspicious audio is not abnormal audio, it can be determined that the audio included in the current first audio data is normal audio.

[0124] If the difference value is less than the difference threshold, it can be determined that the audio included in the current first audio data is normal audio.

[0125] The difference value is compared with the difference threshold. According to the comparison result, when it is determined that the audio included in the current first audio data is suspicious audio, the suspicious audio is confirmed again to accurately determine whether the audio included in the current first audio data is abnormal audio.

[0126] In some embodiments of the present application, the method may further include the following steps:

[0127] When it is determined that the audio included in the current first audio data is abnormal audio, the current first audio data is added to the abnormal audio dataset to perform abnormal processing based on the abnormal audio data in the abnormal audio dataset;

[0128] When it is determined that the audio included in the current first audio data is normal audio, the current first audio data is added to the normal audio dataset of the current road section to regenerate the background map of the current road section based on the differences between the feature maps corresponding to the normal audio data in the normal audio dataset.

[0129] In the embodiment of the present application, for the currently collected first audio data, based on the first difference map between the first feature map corresponding to the current first audio data and the background map of the current road section, it is determined whether the audio included in the current first audio data is abnormal audio.

[0130] If the audio contained in the current first audio data is abnormal audio, it is considered that the devices deployed in the area where the current road section is located may be in an abnormal operating state. The current first audio data can be added to the abnormal audio dataset to perform abnormal processing on the relevant devices based on the abnormal audio data in the abnormal audio dataset, such as device repair or device replacement, etc., to ensure the stable operation of the devices.

[0131] If the audio contained in the current first audio data is normal audio, it is considered that the devices deployed in the area where the current road section is located are in a normal operating state. The current first audio data can be added to the normal audio dataset of the current road section, so that the number of normal audio data of the current road section gradually increases. In this way, based on the differences between the feature maps corresponding to the normal audio data in the normal audio dataset, the background map of the current road section can be regenerated, and the regenerated background map will become more and more accurate. Based on the regenerated background map of the current road section, abnormal identification of the audio data collected by the inspection robot is helpful to improve the accuracy of abnormal audio identification.

[0132] See Figure 2 As shown below, a possible inspection process is as follows:

[0133] Control the inspection robot to walk along the target route;

[0134] Obtain the first audio data collected by the inspection robot in real time;

[0135] Determine whether the audio frequency corresponding to the current first audio data is greater than or equal to the audio frequency threshold corresponding to the current road section;

[0136] If so, calculate the energy spectral density signal of the current first audio data, remove the background noise in the current first audio data, obtain and save the second audio data. If not, continue to process the next first audio data;

[0137] Extract the MFCC features of the second audio data, perform delta processing, normalization processing, and visualization of the MFCC features on the MFCC features to obtain the first feature map corresponding to the current first audio data;

[0138] Calculate the difference between the first feature map and the background map of the current road section to generate a binary image;

[0139] Perform dimensionality reduction processing on the binary image to obtain a first difference map of 80*80;

[0140] Based on the cosine distance between every two white points in the first difference map, determine the difference value between the first feature map and the background map;

[0141] Determine whether the difference value is greater than or equal to the difference threshold;

[0142] If so, determine that the audio included in the current first audio data is suspicious audio; if not, determine that the audio included in the current first audio data is normal audio, and add the current first audio data to the normal audio data set of the current section.

[0143] Have the staff listen to the suspicious audio to confirm whether it is abnormal audio.

[0144] If it is confirmed that it is not abnormal audio, add the current first audio data to the normal audio data set of the current section to update the background map of the current section.

[0145] If it is confirmed that it is abnormal audio, add the current first audio data to the abnormal audio data set for timely abnormal processing.

[0146] Among them, the feature map corresponding to each first audio data is as Figure 3 shown, the background map of the current section is as Figure 4 shown, and the binary image of the feature map corresponding to each first audio data relative to the background map of the current section is as Figure 5 shown.

[0147] In some embodiments of the present application, the audio frequency threshold corresponding to each section can be obtained through the following steps:

[0148] When it is determined that all devices are operating normally, obtain the third audio data collected in real time by the inspection robot according to the target route.

[0149] For each section, when there is third audio data in the current section whose audio frequency is greater than or equal to the initial audio frequency threshold preset for the current section, determine the maximum audio frequency corresponding to the third audio data in the current section as the audio frequency threshold corresponding to the current section.

[0150] When there is no third audio data in the current section whose audio frequency is greater than or equal to the audio frequency initial threshold, determine the audio frequency initial threshold as the audio frequency threshold corresponding to the current section.

[0151] For the convenience of description, the above steps are combined for description.

[0152] In the embodiments of the present application, before performing the inspection task, the audio frequency threshold corresponding to each section can be determined first.

[0153] The initial audio frequency threshold corresponding to each section can be preset according to experience or historical data. The initial audio frequency thresholds corresponding to different sections can be the same or different.

[0154] When it is determined that all devices are operating normally, the inspection robot can be controlled to walk along the target route to obtain the third audio data collected by the inspection robot in real time. Each third audio data has a corresponding audio frequency.

[0155] For each section, the audio frequency of the third audio data of the current section can be compared with the initial audio frequency threshold corresponding to the current section. If there is third audio data in the current section whose audio frequency is greater than or equal to the initial audio frequency threshold, it can be considered that the audio frequency of the normal audio data in the current section is relatively high, and the initial audio frequency threshold corresponding to the current section can be increased. The maximum audio frequency corresponding to the third audio data of the current section can be determined as the audio frequency threshold corresponding to the current section.

[0156] If there is no third audio data in the current section whose audio frequency is greater than or equal to the initial audio frequency threshold, it can be considered that the initial audio frequency threshold corresponding to the current section is set reasonably, and the initial audio frequency threshold can be determined as the audio frequency threshold corresponding to the current section.

[0157] Before performing the inspection task, first determining the audio frequency threshold corresponding to each section based on the audio frequency of the third audio data collected by the inspection robot in real time helps to subsequently screen out the audio data that may be abnormal based on the audio frequency threshold corresponding to each section.

[0158] Before performing the inspection task, controlling the inspection robot to walk along the target route to collect the third audio data and processing the third audio data can realize the initialization of the environmental audio. Because the audio generated by different types of devices in each environment is different, through initialization, the determined audio frequency threshold corresponding to each section can be made more reasonable.

[0159] As Figure 6 shown, a possible implementation process for determining the audio frequency threshold corresponding to each section is as follows:

[0160] Control the inspection robot to walk along the target route;

[0161] Obtain the third audio data collected by the inspection robot in real time. When the audio frequency of the third audio data reaches the preset audio frequency, such as audio frequency 1, start saving the third audio data. Because the inspection robot will also generate noise with a certain frequency during walking, when the audio frequency of the third audio data is relatively low and does not reach this audio frequency 1, it can be considered that the audio contained in the current third audio data is the walking noise of the inspection robot. Such third audio data has no analysis value and can be filtered out without being saved.

[0162] For each road section, if there is third audio data in the current road section whose audio frequency is greater than or equal to the initial audio frequency threshold corresponding to the preset current road section, then mark the current road section. In an industrial environment, the audio frequencies and magnitudes generated in different regions are different. For example, Workshop A is a welding workshop and generates relatively high audio frequencies, while Workshop B is a packaging workshop and generates relatively low audio frequencies. Marking the road sections with audio data having relatively high audio frequencies can timely adjust the audio frequency thresholds corresponding to the respective road sections;

[0163] Traverse each road section and adjust the initial audio frequency threshold corresponding to the marked road section to obtain the audio frequency threshold corresponding to each road section.

[0164] In some embodiments of the present application, the background map of the current road section can be generated through the following steps:

[0165] Determine the set of third audio data collected in the current road section as the normal audio data set of the current road section;

[0166] For each normal audio data in the normal audio data set, calculate the energy spectral density signal of the current normal audio data;

[0167] According to the energy spectral density signal of the current normal audio data, remove the background noise from the current normal audio data to obtain the fourth audio data;

[0168] Perform feature extraction and post-processing on the fourth audio data to obtain the second feature map corresponding to the current fourth audio data;

[0169] Generate the background map of the current road section based on the differences between the second feature maps corresponding to each fourth audio data.

[0170] For the sake of convenient description, the above steps will be combined and described.

[0171] In the embodiments of the present application, when it is determined that all devices are operating normally, control the inspection robot to walk along the target route and obtain the third audio data collected by the inspection robot in real time. The third audio data is normal audio data.

[0172] For each road section, the set of third audio data collected in the current road section can be determined as the normal audio data set of the current road section.

[0173] For each normal audio data in the normal audio data set, the energy spectral density signal of the current normal audio data can be calculated, and then the background audio can be separated according to the energy spectral density signal of the current normal audio data, and the background noise can be removed from the current normal audio data to obtain the fourth audio data. The current normal audio data is the normal audio data targeted by the current operation.

[0174] Feature extraction and post - processing are performed on the fourth audio data, and a second feature map corresponding to the current fourth audio data can be obtained. The second feature map can be regarded as the feature representation of the current normal audio data.

[0175] Optionally, feature extraction can be performed on the fourth audio data to obtain a second feature, and then probability distribution estimation, normalization processing, and visualization processing are sequentially performed on the second feature to obtain a second feature map corresponding to the current fourth audio data.

[0176] Among them, the second feature obtained by performing feature extraction on the fourth audio data can include MFCC features. The probability distribution estimation performed on the second feature can include delta processing.

[0177] By performing the above - mentioned processing on the current normal audio data, the second feature map corresponding to the current fourth audio data can be accurately obtained.

[0178] Each normal audio data in the normal audio data set of the current road section corresponds to a fourth audio data, each fourth audio data corresponds to a second feature map, and there are differences between the second feature maps corresponding to each fourth audio data. Based on the differences between the second feature maps corresponding to each fourth audio data, a background map of the current road section can be generated. Thus, when performing an inspection task, based on the background map of the current road section, a first difference map of the first feature map corresponding to the first audio data collected by the inspection robot in real - time relative to the background map can be obtained, and then it can be accurately determined whether the audio contained in the first audio data is abnormal audio.

[0179] See Figure 7 As shown, it is a schematic structural diagram of an abnormal audio recognition system provided by an embodiment of the present application. The system includes a data middle - platform, a map data center, an inspection robot control center, an audio data acquisition center, and an algorithm calculation center. A user can interact with the map data center, the inspection robot control center, the audio data acquisition center, and the algorithm calculation center through the data middle - platform.

[0180] Specifically, users can issue instructions to execute inspection tasks to the map data center and the inspection robot control center through the data middle platform. The map data center can provide a walking route to the inspection robot control center, so that the inspection robot control center can control the inspection robot to walk along the walking route and collect audio data. Users can also obtain the audio data during the inspection in real time from the inspection robot control center through the data center. Users can also perform manual intervention in case of emergency and control the actions and return of the inspection robot through the data center and the inspection robot control center. The audio data acquisition center can obtain the audio data collected by the inspection robot through the data middle platform, perform feature extraction and visualization processing on the audio data through the algorithm calculation center, and identify abnormal audio based on its difference from the background map. For suspicious audio, it can be output to the user through the data middle platform, and the user can confirm whether the suspicious audio is abnormal audio.

[0181] The users here can be understood as staff members, technical personnel, etc.

[0182] Corresponding to the above method embodiments, the embodiment of the present application also provides an abnormal audio recognition device. The abnormal audio recognition device described below can be mutually referred to with the abnormal audio recognition method described above.

[0183] See Figure 8 As shown, the device includes:

[0184] A first acquisition module 810, configured to acquire first audio data collected in real time by an inspection robot performing inspections according to a target route, where the target route includes multiple road segments, and each road segment corresponds to its own audio frequency threshold;

[0185] A second acquisition module 820, configured to acquire a first feature map corresponding to the current first audio data when the audio frequency corresponding to the currently acquired first audio data is greater than or equal to the audio frequency threshold corresponding to the current road segment;

[0186] An obtaining module 830, configured to compare the first feature map with the background map of the current road segment to obtain a first difference map of the first feature map relative to the background map;

[0187] A determination module 840, configured to determine whether the audio included in the current first audio data is abnormal audio based on the first difference map;

[0188] Wherein, the background map is pre-generated based on the difference between the feature maps corresponding to the normal audio data in the normal audio dataset of the current road segment.

[0189] By applying the device provided in the embodiments of the present application, the background image of each road section is pre-generated based on the differences between the feature maps corresponding to the normal audio data in the normal audio dataset of the corresponding road section, so that the background image of each road section can accurately reflect the characteristics of the normal audio data of the corresponding road section. Furthermore, for each road section, according to the first difference map of the first feature map corresponding to the first audio data with a relatively large audio frequency in the currently collected audio data of the inspection robot with respect to the background image of the current road section, it is possible to accurately determine whether the audio included in the first audio data is abnormal audio. Since the number of normal audio data is large, it is easy to generate the background image of each road section based on the normal audio data, and there is no need to perform the training of the deep neural network, which can effectively avoid the problem that when the deep neural network is trained, it tends to classify normal samples due to the small number of abnormal samples, and it is difficult to accurately identify abnormal audio data.

[0190] In some embodiments of the present application, the second acquisition module 820 is specifically configured to:

[0191] Calculate the energy spectral density signal of the current first audio data;

[0192] Remove background noise from the current first audio data according to the energy spectral density signal of the current first audio data to obtain second audio data;

[0193] Perform feature extraction and post-processing on the second audio data to obtain the first feature map corresponding to the current first audio data.

[0194] In some embodiments of the present application, the second acquisition module 820 is specifically configured to:

[0195] Perform feature extraction on the second audio data to obtain a first feature;

[0196] Perform probability distribution estimation, normalization processing, and visualization processing on the first feature in sequence to obtain the first feature map corresponding to the current first audio data.

[0197] In some embodiments of the present application, the acquisition module 830 is specifically configured to:

[0198] Compare the first feature map with the background image of the current road section to generate a binary image;

[0199] Perform dimensionality reduction processing on the binary image to obtain the first difference map of the first feature map with respect to the background image.

[0200] In some embodiments of the present application, the determination module 840 is specifically configured to:

[0201] Determine the cosine distance between every two difference points in the first difference map;

[0202] Determine the difference value between the first feature map and the background map based on the cosine distance between every two difference points in the first difference map;

[0203] According to the magnitude relationship between the difference value and a preset difference threshold, determine whether the audio included in the current first audio data is abnormal audio.

[0204] In some embodiments of the present application, the determination module 840 is specifically configured to:

[0205] In the case where the difference value is greater than or equal to the preset difference threshold, determine that the audio included in the current first audio data is suspicious audio;

[0206] According to the confirmation information of the suspicious audio, determine whether the audio included in the current first audio data is abnormal audio;

[0207] In the case where the difference value is less than the difference threshold, determine that the audio included in the current first audio data is normal audio.

[0208] In some embodiments of the present application, it further includes an adding module for:

[0209] In the case where it is determined that the audio included in the current first audio data is abnormal audio, add the current first audio data to the abnormal audio dataset to perform abnormal processing based on the abnormal audio data in the abnormal audio dataset;

[0210] In the case where it is determined that the audio included in the current first audio data is normal audio, add the current first audio data to the normal audio dataset of the current road section to regenerate the background map of the current road section based on the differences between the feature maps corresponding to the normal audio data in the normal audio dataset.

[0211] In some embodiments of the present application, the acquisition module 820 is further configured to obtain the audio frequency threshold corresponding to each road section through the following steps:

[0212] In the case where it is determined that all devices are operating normally, obtain the third audio data collected in real time by the inspection robot according to the target route;

[0213] For each road section, in the case where there is third audio data with an audio frequency greater than or equal to the initial audio frequency threshold corresponding to the current road section in the third audio data, determine the maximum audio frequency corresponding to the third audio data of the current road section as the audio frequency threshold corresponding to the current road section;

[0214] In the case where there is no third audio data with an audio frequency greater than or equal to the audio frequency initial threshold in the current road section, determine the audio frequency initial threshold as the audio frequency threshold corresponding to the current road section.

[0215] In some embodiments of the present application, it further includes a generation module, which is used to generate a background map of the current road section through the following steps:

[0216] Determine the set of the third audio data collected in the current road section as the normal audio data set of the current road section;

[0217] For each normal audio data in the normal audio data set, calculate the energy spectral density signal of the current normal audio data;

[0218] Remove background noise from the current normal audio data according to the energy spectral density signal of the current normal audio data to obtain the fourth audio data;

[0219] Perform feature extraction and post-processing on the fourth audio data to obtain a second feature map corresponding to the current fourth audio data;

[0220] Generate a background map of the current road section based on the differences between the second feature maps corresponding to each fourth audio data.

[0221] In some embodiments of the present application, the generation module is specifically used for:

[0222] Perform feature extraction on the fourth audio data to obtain a second feature;

[0223] Perform probability distribution estimation, normalization processing, and visualization post-processing on the second feature in sequence to obtain a second feature map corresponding to the current fourth audio data.

[0224] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0225] Corresponding to the above method embodiments, an embodiment of the present application further provides an electronic device, including:

[0226] A memory for storing a computer program;

[0227] A processor for implementing the steps of the above abnormal audio recognition method when executing the computer program.

[0228] As Figure 9 shown, it is a schematic structural diagram of the electronic device. The electronic device may include: a processor 10, a memory 11, a communication interface 12, and a communication bus 13. The processor 10, the memory 11, and the communication interface 12 all complete communication with each other through the communication bus 13.

[0229] In an embodiment of the present application, the processor 10 may be a central processing unit (CPU), an application specific integrated circuit, a digital signal processor, a field programmable gate array, or other programmable logic devices, etc.

[0230] The processor 10 may call the program stored in the memory 11. Specifically, the processor 10 may execute the operations in the embodiment of the abnormal audio recognition method.

[0231] The memory 11 is used to store one or more programs. The program may include program code, and the program code includes computer operation instructions. In an embodiment of the present application, the memory 11 stores at least a program for implementing the following functions:

[0232] Obtain first audio data collected in real time by a patrol robot patrolling along a target route. The target route includes multiple sections, and each section corresponds to its own audio frequency threshold;

[0233] When the audio frequency corresponding to the currently collected first audio data is greater than or equal to the audio frequency threshold corresponding to the current section, obtain a first feature map corresponding to the current first audio data;

[0234] Compare the first feature map with the background map of the current section to obtain a first difference map of the first feature map relative to the background map;

[0235] Based on the first difference map, determine whether the audio included in the current first audio data is abnormal audio;

[0236] Wherein, the background map is pre-generated based on the differences between the feature maps corresponding to the normal audio data in the normal audio dataset of the current section.

[0237] In a possible implementation, the memory 11 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function, etc.; the data storage area may store the data created during use.

[0238] In addition, the memory 11 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device or other volatile solid-state storage devices.

[0239] The communication interface 12 may be an interface of a communication module for connecting to other devices or systems.

[0240] Of course, it should be noted that Figure 9 The structure shown does not constitute a limitation on the electronic device in the embodiment of the present application. In actual applications, the electronic device may include more Figure 9more or fewer components shown, or combine certain components.

[0241] Corresponding to the above method embodiments, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above abnormal audio recognition method are implemented.

[0242] In addition, it should be noted that: An embodiment of the present application further provides a computer program product or a computer program. The computer program product or the computer program may include computer instructions, and the computer instructions may be stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor may execute the computer instructions, so that the computer device executes the description of the abnormal audio recognition method in the corresponding embodiment described above. Therefore, the description will not be repeated here. In addition, the description of the beneficial effects of using the same method will not be repeated. For the technical details not disclosed in the computer program product or the computer program embodiment involved in the present application, please refer to the description of the method embodiment of the present application.

[0243] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other.

[0244] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may be added, omitted or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0245] From the description of the above embodiments, those skilled in the art can also clearly understand that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0246] The steps of the methods or algorithms described in combination with the embodiments disclosed in this article can be directly implemented by hardware, software modules executed by a processor, or a combination of both. The software modules can be placed in a random access memory (RAM), memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable disks, compact disc read-only memory (CD-ROM), or any other form of storage medium well-known in the technical field, including several instructions for executing the methods described in various embodiments of this application.

[0247] The embodiments of this application have been described above in conjunction with the accompanying drawings. The description of the above embodiments is only used to help understand the technical solution and its core idea of this application. It should be noted that this application is not limited to the above specific implementation manners. The above specific implementation manners are only illustrative and not restrictive. For those of ordinary skill in the art, without departing from the purpose of this application and the scope protected by the claims, many forms of implementation manners can still be made, and several improvements and modifications can also be made to this application. These implementation manners, improvements, and modifications are all within the protection scope of this application.

Claims

1. A method for identifying abnormal audio, characterized in that: include: Acquire first audio data collected in real time by an inspection robot inspecting a target route, wherein the target route includes a plurality of sections, and each section corresponds to a respective audio frequency threshold; When the audio frequency corresponding to the first audio data currently collected is greater than or equal to the audio frequency threshold corresponding to the current road section, obtaining a first feature graph corresponding to the current first audio data; Compare the first feature map with the background map of the current road section to obtain a first difference map between the first feature map and the background map; Based on the first difference map, determining whether the audio included in the current first audio data is abnormal audio; The background image is pre-generated based on the difference between feature images corresponding to normal audio data in the normal audio data set of the current road section.

2. The method according to claim 1, characterized in that The obtaining of a first feature map corresponding to the current first audio data includes: Calculate the energy spectrum density signal of the current first audio data; According to the energy spectrum density signal of the current first audio data, background noise is removed from the current first audio data to obtain second audio data; Feature extraction and post-processing are performed on the second audio data to obtain a first feature map corresponding to the current first audio data.

3. The method according to claim 2, characterized in that The extracting and post-processing the second audio data to obtain a first feature graph corresponding to the current first audio data includes: Performing feature extraction on the second audio data to obtain a first feature; performing probability distribution estimation, normalization processing, and visualization processing on the first feature in sequence to obtain a first feature graph corresponding to the current first audio data; and / or, The step of comparing the first feature map with the background map of the current road section to obtain a first difference map of the first feature map relative to the background map includes: Compare the first feature map with the background map of the current road section to generate a binary image; Perform dimensionality reduction processing on the binary image to obtain a first difference map between the first feature map and the background map.

4. The method according to any one of claims 1 to 3, characterized in that The determining, based on the first difference map, whether the audio included in the current first audio data is abnormal audio includes: Determine the cosine distance between every two difference points in the first difference map; Determine a difference value between the first feature map and the background map based on a cosine distance between every two difference points in the first difference map; According to the magnitude relationship between the difference value and a preset difference threshold, it is determined whether the audio included in the current first audio data is abnormal audio.

5. The method according to claim 4, characterized in that The determining, based on a magnitude relationship between the difference value and a preset difference threshold, whether the audio included in the current first audio data is abnormal audio includes: When the difference value is greater than or equal to a preset difference threshold, determining that the audio included in the current first audio data is suspicious audio; Determining whether the audio included in the current first audio data is abnormal audio according to the confirmation information of the suspicious audio; When the difference value is less than the difference threshold, it is determined that the audio included in the current first audio data is normal audio.

6. The method according to any one of claims 1 to 3 and 5, characterized in that: The method further comprises: In the case where it is determined that the audio included in the current first audio data is abnormal audio, adding the current first audio data to an abnormal audio data set to perform abnormal processing based on the abnormal audio data in the abnormal audio data set; When it is determined that the audio contained in the current first audio data is normal audio, the current first audio data is added to the normal audio data set of the current road section to regenerate the background image of the current road section based on the difference between the feature maps corresponding to the normal audio data in the normal audio data set.

7. The method according to any one of claims 1 to 3 and 5, characterized in that: Obtain the audio frequency threshold corresponding to each segment through the following steps: When it is determined that all devices are operating normally, obtaining third audio data collected in real time by the inspection robot along the target route; For each section, if there is third audio data in the current section whose audio frequency is greater than or equal to a preset initial audio frequency threshold corresponding to the current section, the maximum audio frequency corresponding to the third audio data of the current section is determined as the audio frequency threshold corresponding to the current section; In the case that the current section does not contain third audio data whose audio frequency is greater than or equal to the initial audio frequency threshold, the initial audio frequency threshold is determined as the audio frequency threshold corresponding to the current section.

8. The method according to claim 7, characterized in that The background image of the current road section is generated by following the steps below: Determine a set consisting of third audio data collected on the current road section as a normal audio data set of the current road section; For each normal audio data in the normal audio data set, calculating an energy spectrum density signal of the current normal audio data; removing background noise from the current normal audio data according to the energy spectrum density signal of the current normal audio data to obtain fourth audio data; Performing feature extraction and post-processing on the fourth audio data to obtain a second feature map corresponding to the current fourth audio data; Based on the difference between the second feature maps corresponding to each fourth audio data, a background map of the current road section is generated.

9. An abnormal audio recognition device, characterized in that: include: A first acquisition module is used to acquire first audio data collected in real time by an inspection robot inspecting a target route, wherein the target route includes a plurality of sections, and each section corresponds to a respective audio frequency threshold; A second acquisition module is used to acquire a first feature map corresponding to the current first audio data when the audio frequency corresponding to the currently collected first audio data is greater than or equal to the audio frequency threshold corresponding to the current road section; An obtaining module, configured to compare the first feature map with the background map of the current road section to obtain a first difference map of the first feature map relative to the background map; A determination module, configured to determine whether the audio included in the current first audio data is abnormal audio based on the first difference map; The background image is pre-generated based on the difference between feature images corresponding to normal audio data in the normal audio data set of the current road section.

10. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the abnormal audio identification method according to any one of claims 1 to 8 when executing the computer program.