An environmental sound recognition and classification system and method

By screening multiple neural network recognition models in the environmental sound recognition classification system and identifying and classifying them according to environmental feature data, the problem of long training time of deep neural network recognition model in the prior art is solved, and the classification efficiency of environmental sound recognition and classification and system response speed are improved.

CN116386662BActive Publication Date: 2025-07-01GUANGDONG VOCATIONAL COLLEGE OF POST & TELECOM
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202310097384.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-07
Publication Date
2025-07-01
Estimated Expiration
2043-02-07

AI Technical Summary

Technical Problem

In the prior art, the long training time required to use a single deep neural network recognition model to identify different types of sounds, resulting in low classification efficiency of environmental sound recognition.

Method used

By establishing an environmental sound recognition classification system between the monitoring terminal and the server, acquiring environmental sound data and environmental feature data, and filtering multiple different types of neural network recognition models based on environmental feature data, the environmental sound data is identified and classified.

Benefits of technology

This method can use the screened neural network recognition model to identify and classify, overcome the problem of long training time of deep neural network recognition model, and improve the classification efficiency of environmental sound recognition and classification and system response speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116386662B_ABST
    Figure CN116386662B_ABST
Patent Text Reader

Abstract

The present invention provides an environmental sound recognition and classification system and method. The environmental sound recognition and classification system includes: a monitoring terminal, configured to obtain environmental sound data and environmental feature data corresponding to the environmental sound data and feed them back to a server; the server stores multiple neural network recognition models of different types and is configured to receive the environmental sound data and the environmental feature data, and screen a neural network recognition model according to the environmental feature data to perform recognition and classification on the environmental sound data. The present invention can specifically use the screened neural network recognition model to perform recognition and classification on the environmental sound data, overcoming the problem of the long training time required for using a single deep neural network recognition model to recognize different types of sounds in the prior art, and improving the recognition and classification efficiency of environmental sounds.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of sound recognition, and more particularly, to an environmental sound recognition and classification system and method. Background Art

[0002] Sounds in the environment are diverse. For example, urban environmental sounds generally include air conditioner sounds, car honks, children playing sounds, dog barks, drilling sounds, engine idling sounds, gunshots, jackhammers, police sirens, and street music sounds, etc. Agricultural farm sounds include wind sounds, human voices, bird calls, and insect sounds, etc. Identifying and classifying environmental sounds helps to monitor the environment in real time.

[0003] After the applicant's retrieval, some typical prior arts were found. For example, the Chinese invention patent with the application number CN2013100438026 discloses a voice-based animal recognition method and device, which extracts the acoustic feature parameters of animal calls through voice spectrum analysis and matches them with the database model to identify the surrounding animal species and their quantity distribution. Especially in the wild, it can achieve the purpose of seeking advantages and avoiding disadvantages, and the operation experience has entertainment and interestingness. Another example is the Chinese invention patent with the application number CN2021104375143, which discloses a voice system based on object recognition. It uses a first sound collection module to collect environmental sounds, collects the sounds of staff or small animals entering the substation from a large range, and performs tracking processing on the collected sounds through a second sound collection module, making the sound collection and recognition more accurate and the data processing faster. The processing module irradiates lights according to the movement prediction of small animals, and can quickly expel small animals. Still another example is the Chinese invention patent with the application number CN015102260826, which discloses an animal sound recognition method based on double features of spectrograms. Through the training of random forests, the category corresponding to the sound signal to be recognized in the sound sample library is obtained and the result is output, improving the recognition rate of various low signal-to-noise ratio animal sounds in different sound environments.

[0004] It can be seen that there are still many unsolved technical problems (such as improving the efficiency of environmental sound recognition and classification, etc.) in the practical application of sound recognition, and there are still many solutions that have not been proposed. Summary of the Invention

[0005] Based on this, in order to improve the efficiency of environmental sound recognition and classification, the present invention provides an environmental sound recognition and classification system and method, and its specific technical solutions are as follows:

[0006] An environmental sound recognition and classification system includes a monitoring terminal and a server.

[0007] The monitoring terminal is used to obtain environmental sound data and environmental feature data corresponding to the environmental sound data and feedback them to the server.

[0008] The server stores multiple neural network recognition models of different types, which are used to receive environmental sound data and environmental feature data, and screen the neural network recognition models according to the environmental feature data to identify and classify the environmental sound data.

[0009] By acquiring environmental sound data and the corresponding environmental feature data, and screening the neural network recognition models according to the environmental feature data, the environmental sound recognition and classification system can specifically use the screened neural network recognition model to identify and classify the environmental sound data, overcoming the problem of the long training time required for using a single deep neural network recognition model to identify different types of sounds in the prior art, and improving the efficiency of environmental sound recognition and classification.

[0010] Further, the monitoring terminal includes a first acquisition module and a second acquisition module.

[0011] The first acquisition module is used to acquire environmental sound data and feedback it to the server; the second acquisition module is used to acquire the environmental feature data corresponding to the environmental sound data and feedback it to the server.

[0012] Further, the server includes an identification module, a screening module, and a classification module.

[0013] The identification module is used to identify the possible environmental types according to the environmental feature data; the screening module is used to screen the neural network recognition model corresponding to the possible environmental type according to the possible environmental type.

[0014] The classification module is used to identify and classify the environmental sound data according to the screened neural network recognition model corresponding to the possible environmental type.

[0015] Further, the identification module includes:

[0016] A first calculation unit, which is used to calculate the probability values p between the on-site environment where the monitoring terminal is located and multiple pre-stored environmental types respectively according to the environmental feature data;

[0017] An identification unit, which is used to use the preset environmental type with the probability value p greater than the preset probability threshold as the possible environmental type.

[0018] Further, the classification module includes:

[0019] A second calculation unit, which is used to calculate the probability average value

[0020] A classification unit, which is used to according to the probability average value And the probability value p of possible environmental types divides the screened neural network recognition models corresponding to possible environmental types into a preferred recognition model and an alternative recognition model;

[0021] An identification unit, configured to identify and classify environmental sound data according to the preferred recognition model and the alternative recognition model.

[0022] Furthermore, an environmental sound identification and classification method is applied to the environmental sound identification and classification system, and it includes the following steps:

[0023] Obtain environmental sound data and environmental feature data corresponding to the environmental sound data and feedback them to the server;

[0024] The server receives the environmental sound data and the environmental feature data, and screens the neural network recognition model according to the environmental feature data to identify and classify the environmental sound data.

[0025] Furthermore, the specific method for screening the neural network recognition model according to the environmental feature data to identify and classify the environmental sound data includes the following steps:

[0026] Identify possible environmental types according to the environmental feature data;

[0027] Screen the neural network recognition model corresponding to the possible environmental type according to the possible environmental type;

[0028] Identify and classify the environmental sound data according to the screened neural network recognition model corresponding to the possible environmental type.

[0029] Furthermore, the specific method for identifying possible environmental types according to the environmental feature data includes the following steps:

[0030] Calculate the probability value p between the on-site environment where the monitoring terminal is located and multiple pre-stored environmental types respectively according to the environmental feature data;

[0031] Take the pre-set environmental type with the probability value p greater than the pre-set probability threshold as the possible environmental type.

[0032] Furthermore, the specific method for identifying and classifying the environmental sound data according to the screened neural network recognition model corresponding to the possible environmental type includes the following steps:

[0033] Calculate the probability average value

[0034] According to the probability average value And the probability value p of possible environmental types divides the screened neural network recognition models corresponding to possible environmental types into a preferred recognition model and an alternative recognition model;

[0035] Identify and classify the environmental sound data according to the preferred recognition model and the alternative recognition model.

[0036] Furthermore, the present invention provides a computer-readable storage medium storing a computer program, which implements the environmental sound recognition and classification method when the computer program is executed. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The present invention can be further understood from the following description in conjunction with the drawings. The components in the drawings are not necessarily drawn to scale, but the emphasis is placed on showing the principles of the embodiments. In different views, the same reference numerals designate corresponding parts.

[0038] Figure 1 is a schematic diagram of the overall structure of an environmental sound recognition and classification system in an embodiment of the present invention;

[0039] Figure 2 is a schematic diagram of the overall process of an environmental sound recognition and classification method in an embodiment of the present invention;

[0040] Figure 3 is a schematic diagram of the overall process of an environmental sound recognition and classification method in another embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with its embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not limit the protection scope of the present invention.

[0042] It should be noted that when an element is referred to as being "fixed to" another element, it can be directly on the other element or there may also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are only for the purpose of illustration and do not represent the only implementation.

[0043] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0044] The "first" and "second" in the present invention do not represent specific quantities and orders, but are only used for name distinction.

[0045] In the prior art, generally, a deep neural network trained with a training set of hundreds of thousands or even millions of samples is used to identify and classify environmental sounds. Training such a deep neural network may take several days or even weeks computationally. And using the trained deep neural network to identify and classify the input environmental sounds also often takes a long time. Therefore, when using a deep neural network to identify and classify environmental sounds, there is a problem of low working efficiency and it needs to be improved.

[0046] To improve the efficiency of identifying and classifying environmental sounds, as Figure 1 shown, in an embodiment of the present invention, an environmental sound identification and classification system is provided, which includes a monitoring terminal and a server.

[0047] The monitoring terminal is used to obtain environmental sound data and environmental feature data corresponding to the environmental sound data and feedback them to the server.

[0048] Specifically, the monitoring terminal includes a first acquisition module and a second acquisition module.

[0049] The first acquisition module is used to obtain environmental sound data and feedback it to the server; the second acquisition module is used to obtain environmental feature data corresponding to the environmental sound data and feedback it to the server.

[0050] The first acquisition module includes, but is not limited to, a sound collector, which is installed in the on-site environment to be monitored and is used to collect environmental sound data of the on-site environment.

[0051] The second acquisition module includes, but is not limited to, image collectors such as cameras and scanners, which are installed in the on-site environment to be monitored and are used to collect environmental feature data of the on-site environment.

[0052] The environmental feature data includes, but is not limited to, environmental two-dimensional plane images or three-dimensional stereo images, which are used to characterize the environmental type.

[0053] Here, the environmental type includes, but is not limited to, farms, shopping malls, borders, streets, schools, rural areas, forest farms, and factories, etc.

[0054] The server stores multiple neural network recognition models of different types, which are used to receive environmental sound data and environmental feature data, and screen the neural network recognition models according to the environmental feature data to identify and classify the environmental sound data.

[0055] The multiple neural network recognition models of different types refer to the trained neural network recognition models used to identify different types of sounds.

[0056] For each neural network recognition model, it is trained using sound datasets of several sound types. In this way, each neural network recognition model is used to recognize different types of environmental sound data. As Figure 1 shown, preferably, the system further includes a cloud in communication connection with the server, and the sound dataset is obtained from the cloud and transmitted to the server.

[0057] After receiving the environmental feature data, the server filters out the trained neural network recognition model that matches the environmental feature data, that is, the trained neural network recognition model obtained by filtering can be used to recognize and classify the environmental sound data.

[0058] Since the trained neural network recognition model obtained by filtering is specifically trained using sound datasets of several sound types, the time spent in its training process is less than that of the deep neural network.

[0059] Similarly, using the trained neural network recognition model obtained by filtering to recognize and classify the environmental sound data can also improve the response speed and calculation speed of the system, and overcome the problem that the existing sound recognition system requires a large amount of computing power resources to recognize sounds using the deep neural network recognition model.

[0060] In summary, the environmental sound recognition and classification system obtains environmental sound data and the corresponding environmental feature data, filters the neural network recognition model according to the environmental feature data, and specifically uses the filtered neural network recognition model to recognize and classify the environmental sound data. It can not only overcome the problem of the long training time required for using a single deep neural network recognition model to recognize multiple different types of sounds in the prior art, improve the recognition and classification efficiency of environmental sounds, but also overcome the problem that the existing sound recognition system requires a large amount of computing power resources to recognize sounds using the deep neural network recognition model, and improve the response speed and calculation speed of the system.

[0061] In one embodiment, the server includes an identification module, a filtering module, and a classification module.

[0062] The identification module is used to identify possible environmental types according to the environmental feature data; the filtering module is used to filter the neural network recognition model corresponding to the possible environmental types according to the possible environmental types.

[0063] The classification module is used to recognize and classify the environmental sound data according to the filtered neural network recognition model corresponding to the possible environmental types.

[0064] Specifically, the identification module includes a first calculation unit and an identification unit.

[0065] The first computing unit is used to calculate the probability value p between the on-site environment where the monitoring terminal is located and multiple pre-stored environment types respectively according to the environmental characteristic data.

[0066] The recognition unit is used to take the pre-set environment type with the probability value p greater than the pre-set probability threshold as the possible environment type.

[0067] In the server, multiple pre-stored environment types are stored. The pre-stored environment types include but are not limited to farms, shopping malls, borders, streets, schools, rural areas, forest farms, factories, etc. The first computing unit calculates the similarity between the on-site environment where the monitoring terminal is located and multiple pre-stored environment types respectively according to the environmental characteristic data, and calculates the probability value according to the similarity.

[0068] In the server, multiple standard characteristic data corresponding to each pre-stored environment type are pre-stored, such as multiple pre-stored two-dimensional plane images or pre-stored three-dimensional stereo images; at this time, the environmental characteristic data corresponds to the standard characteristic data pre-stored in the server, and is an environmental two-dimensional plane image or a three-dimensional stereo image.

[0069] By calculating the similarity between the environmental characteristic data and the standard characteristic data, the probability value p between the on-site environment where the monitoring terminal is located and multiple pre-stored environment types is calculated.

[0070] In some cases, the probability values between the environmental characteristic data obtained by the monitoring terminal and multiple predicted environment types are different. For the pre-set environment type corresponding to the probability value less than or equal to the pre-set probability threshold, it can be considered that it is not consistent with the on-site environment. That is to say, the pre-set environment type with the probability value greater than the pre-set probability threshold is taken as the possible environment type.

[0071] After determining the possible environment type of the on-site environment, obtain the sound types included in the possible environment type, and screen out the trained neural network recognition model that matches according to the sound types included in the possible environment type, that is, the trained neural network recognition model that is screened can be used to identify and classify the environmental sound data.

[0072] By setting the pre-set probability threshold, using the probability value p between the on-site environment where the monitoring terminal is located and multiple pre-stored environment types and the pre-set probability threshold, screening multiple pre-set environment types, obtaining the possible environment type that matches the on-site environment, and finally screening out the trained neural network recognition model that matches according to the sound types included in the possible environment type, and using the trained neural network recognition model that is screened to identify and classify the environmental sound data, the screened neural network recognition model can be used to identify and classify the environmental sound data specifically, improving the recognition and classification efficiency of environmental sounds, the response speed and calculation speed of the system.

[0073] In one of the embodiments, the classification module includes a second calculation unit, a classification unit, and an identification unit.

[0074] The second calculation unit is used to calculate the probability average value

[0075] The classification unit is used to, according to the probability average value and the probability value p of the possible environment type, classify the filtered neural network identification models corresponding to the possible environment type into a preferred identification model and an alternative identification model.

[0076] The identification unit is used to identify and classify the environmental sound data according to the preferred identification model and the alternative identification model.

[0077] Wherein, N represents the number of possible environment types, and p i represents the probability value between the i-th possible environment type and the on-site environment.

[0078] After taking the preset environment types with the probability value p greater than the preset probability threshold as the possible environment types, there may be multiple possible environment types that match the on-site environment. The probability values corresponding to the multiple possible environment types are different from each other.

[0079] Generally speaking, the larger the probability value corresponding to the possible environment type, the greater the probability that the sound types included in the collected environmental sound data are a subset of the sound types included in the possible environment type. For example, when the possible environment type is a street, the sound types it includes are air conditioner sound, car horn sound, children playing sound, dog barking sound, drilling sound, engine idling sound, gunfire sound, hand drill, police siren sound, and street music sound. At this time, the sound types included in the collected environmental sound data may be one or several of the air conditioner sound, car horn sound, children playing sound, dog barking sound, drilling sound, engine idling sound, gunfire sound, hand drill, police siren sound, and street music sound.

[0080] The smaller the probability value corresponding to the possible environment type, one or several of the sound types included in the collected environmental sound data may not be among the sound types included in the possible environment type.

[0081] That is to say, the larger the probability value corresponding to the possible environment type, the higher the efficiency of identifying and classifying the environmental sound data according to the trained neural network identification model filtered out according to the sound types included in the possible environment type.

[0082] If the collected environmental sound data is randomly and / or sequentially identified and classified using the screened neural network recognition models corresponding to possible environmental types, it is more likely that more than half of the screened neural network recognition models corresponding to possible environmental types are required to completely identify and classify the environmental sound data.

[0083] Therefore, the screened neural network recognition models corresponding to possible environmental types are divided into preferred recognition models and alternative recognition models. First, use the preferred recognition models to identify and classify environmental sounds. In the case where the preferred recognition models cannot identify and classify environmental sounds, then use the alternative recognition models to identify and classify environmental sounds. Compared with randomly and / or sequentially using the screened neural network recognition models corresponding to possible environmental types to identify and classify the collected environmental sound data, the efficiency of identifying and classifying environmental sound data can be improved as a whole.

[0084] That is to say, according to the probability average value and the probability value p of the possible environmental type, the screened neural network recognition models corresponding to possible environmental types are divided into preferred recognition models and alternative recognition models. Using the preferred recognition models and alternative recognition models to identify and classify environmental sound data can further improve the efficiency of recognition and classification.

[0085] In one embodiment, as Figure 2 shown, an environmental sound recognition and classification method is applied to the environmental sound recognition and classification system, and it includes the following steps:

[0086] S1, obtain environmental sound data and environmental feature data corresponding to the environmental sound data and feedback them to the server.

[0087] S2, the server receives the environmental sound data and environmental feature data, and screens the neural network recognition models according to the environmental feature data to identify and classify the environmental sound data.

[0088] Preferably, use the voice training sets of a preset number of types to train the neural network recognition models. Each type of voice training set includes multiple audio samples.

[0089] Specifically, divide the preset M voice types into N groups; where N is less than or equal to M.

[0090] After dividing the M voice types into N groups, construct N neural network recognition models, and use the voice training sets of the voice types corresponding to the groups to train the neural network recognition models. That is to say, each group of voice types corresponds to a neural network recognition model

[0091] By dividing M preset sound types into N groups and training N neural network recognition models, the problem of the long training time required to recognize multiple different types of sounds using a single deep neural network recognition model in the prior art can be overcome.

[0092] At the same time, by constructing N neural network recognition models and training the corresponding neural network recognition models using the sound training sets of several sound types in the corresponding groups, the problem of the large computing power resources required to construct and train a deep neural network recognition model can also be avoided.

[0093] Using the selected and trained neural network recognition models to identify and classify environmental sound data can also improve the response speed and computing speed of the system, and overcome the problem of the large computing power resources required for a sound recognition system to identify sounds using a deep neural network recognition model in the prior art.

[0094] That is to say, the environmental sound recognition and classification method obtains environmental sound data and environmental feature data corresponding to the environmental sound data, screens the neural network recognition models according to the environmental feature data, and specifically uses the selected neural network recognition models to identify and classify the environmental sound data, which can not only overcome the problem of the long training time required to recognize multiple different types of sounds using a single deep neural network recognition model in the prior art, improve the recognition and classification efficiency of environmental sounds, but also overcome the problem of the large computing power resources required for a sound recognition system to identify sounds using a deep neural network recognition model in the prior art, and improve the response speed and computing speed of the system.

[0095] In one embodiment, in step S2, as Figure 3 shown, the specific method for screening the neural network recognition models according to the environmental feature data to identify and classify the environmental sound data includes the following steps:

[0096] S20, identifying the possible environmental types according to the environmental feature data.

[0097] S21, screening the neural network recognition models corresponding to the possible environmental types according to the possible environmental types.

[0098] S22, identifying and classifying the environmental sound data according to the selected neural network recognition models corresponding to the possible environmental types.

[0099] Specifically, the specific method for identifying the possible environmental types according to the environmental feature data includes the following steps:

[0100] S200, respectively calculating the probability values p between the on-site environment where the monitoring terminal is located and multiple pre-stored environmental types according to the environmental feature data.

[0101] S201. Take the preset environmental types with probability value p greater than the preset probability threshold as possible environmental types.

[0102] The server pre-stores multiple standard feature data corresponding to each pre-stored environmental type, such as multiple pre-stored two-dimensional plane images or pre-stored three-dimensional solid images; at this time, the environmental feature data corresponds to the standard feature data pre-stored in the server, which is a two-dimensional plane image or a three-dimensional solid image of the on-site environment.

[0103] By calculating the similarity between the environmental feature data and the standard feature data, calculate the probability value p between the on-site environment where the monitoring terminal is located and multiple pre-stored environmental types. That is, the probability value p between the on-site environment where the monitoring terminal is located and the pre-stored environmental type is the similarity between the environmental feature data of the on-site environment and the standard feature data of the preset environmental type.

[0104] More specifically, as a preferred technical solution, each pre-stored environmental type corresponds to multiple standard feature data.

[0105] The standard feature data are two-dimensional plane images and three-dimensional plane images of the pre-stored environmental type taken from different angles.

[0106] The environmental feature data are two-dimensional plane images and three-dimensional solid images of the on-site environment.

[0107] The specific method for calculating the similarity between the environmental feature data and the standard feature data includes the following steps:

[0108] The first step is to extract the contour lines of the environmental feature data and the standard feature data.

[0109] The second step is to obtain the number of item types included in the environmental feature data and the standard feature data respectively according to the contour lines of the environmental feature data and the standard feature data. The item types include but are not limited to buildings, plants, public facilities, vehicles, personnel, roads, ridges, overpasses, etc.

[0110] The third step is to calculate the similarity Sim = S1 / S2 between the environmental feature data and the standard feature data according to the number of item types S1 included in the environmental feature data and the number of item types S2 included in the standard feature data.

[0111] By calculating the number of item types included in the environmental feature data and the standard feature data to determine the similarity between the environmental feature data and the standard feature data, the calculation process of the probability value p between the on-site environment where the monitoring terminal is located and the pre-stored environmental type can be simplified. Of course, for the number of item types S2 included in the standard feature data, it can also be determined by technical personnel and input into the server according to the two-dimensional plane images and three-dimensional plane images of the pre-stored environmental type taken from different angles.

[0112] Since both the preset probability threshold and the item type can be set according to actual needs, they will not be elaborated here.

[0113] By setting a preset probability threshold, using the probability value p between the on-site environment where the monitoring terminal is located and multiple pre-stored environment types, and the preset probability threshold, multiple preset environment types are screened to obtain possible environment types that match the on-site environment. Finally, according to the sound types included in the possible environment types, a trained neural network recognition model that matches is screened out. By using the screened and trained neural network recognition model to identify and classify environmental sound data, the screened neural network recognition model can be specifically used to identify and classify environmental sound data, improving the efficiency of identifying and classifying environmental sounds, the response speed of the system, and the calculation speed.

[0114] In one embodiment, the specific method for identifying and classifying environmental sound data according to the screened neural network recognition model corresponding to the possible environment type includes the following steps:

[0115] S220, calculate the average probability

[0116] S221, according to the average probability and the probability value p of the possible environment type, the screened neural network recognition model corresponding to the possible environment type is divided into a preferred recognition model and an alternative recognition model.

[0117] S222, identify and classify the environmental sound data according to the preferred recognition model and the alternative recognition model.

[0118] Specifically, the neural network recognition model corresponding to the possible environment type with a probability value greater than the average probability is divided into a preferred recognition model, and the neural network recognition model corresponding to the possible environment type with a probability value less than or equal to the average probability is divided into an alternative recognition model.

[0119] Dividing the screened neural network recognition model corresponding to the possible environment type into a preferred recognition model and an alternative recognition model, first using the preferred recognition model to identify and classify environmental sounds. In the case where the preferred recognition model cannot identify and classify environmental sounds, then using the alternative recognition model to identify and classify environmental sounds. Compared with randomly and / or sequentially using the screened neural network recognition model corresponding to the possible environment type to identify and classify the collected environmental sound data, the efficiency of identifying and classifying environmental sound data can be improved as a whole.

[0120] That is to say, according to the average probability The probability value p of the possible environmental types divides the screened neural network recognition models corresponding to the possible environmental types into a preferred recognition model and an alternative recognition model. Using the preferred recognition model and the alternative recognition model to identify and classify environmental sound data can further improve the efficiency of recognition and classification.

[0121] In one embodiment, the present invention provides a computer-readable storage medium storing a computer program, which implements the environmental sound recognition and classification method when the computer program is executed.

[0122] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0123] The above embodiments only represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the invention patent should be subject to the appended claims.

Claims

1. An environmental sound recognition and classification system, characterized in that, The described environmental sound recognition and classification system includes: A monitoring terminal, which is used to obtain environmental sound data and environmental feature data corresponding to the environmental sound data and feedback them to the server; A server, which stores multiple neural network recognition models of different types, and is used to receive the environmental sound data and environmental feature data, and screen the neural network recognition models according to the environmental feature data to recognize and classify the environmental sound data; The server includes: A recognition module, which is used to recognize the possible environmental types according to the environmental feature data; A screening module, which is used to screen the neural network recognition model corresponding to the possible environmental type according to the possible environmental type; A classification module, which is used to recognize and classify the environmental sound data according to the screened neural network recognition model corresponding to the possible environmental type; The recognition module includes: A first calculation unit, which is used to calculate the probability value p between the on-site environment where the monitoring terminal is located and multiple pre-stored environmental types respectively according to the environmental feature data; A recognition unit, which is used to use the pre-set environmental type with the probability value p greater than the pre-set probability threshold as the possible environmental type; The classification module includes: A second calculation unit for calculating the average probability A taxonomic unit for classifying neural network recognition models corresponding to possible environmental types according to the average probability value and the probability value p of the possible environmental type into a preferred recognition model and an alternative recognition model; A recognition unit, which is used to recognize and classify the environmental sound data according to the preferred recognition model and the alternative recognition model; Among them, the server stores multiple pre-stored environmental types, and the first calculation unit calculates the similarity between the on-site environment where the monitoring terminal is located and multiple pre-stored environmental types respectively according to the environmental feature data, and calculates the probability value according to the similarity.

2. The environmental sound recognition and classification system according to claim 1, characterized in that, The monitoring terminal includes: A first acquisition module, which is used to acquire environmental sound data and feedback it to the server; A second acquisition module, which is used to acquire environmental feature data corresponding to the environmental sound data and feedback it to the server.

3. An environmental sound recognition and classification method, applied to the environmental sound recognition and classification system according to any one of claims 1-2, characterized in that, The environmental sound recognition and classification method includes the following steps: Acquire environmental sound data and environmental feature data corresponding to the environmental sound data and feedback them to the server; The server receives the environmental sound data and environmental feature data, and screens the neural network recognition model according to the environmental feature data to recognize and classify the environmental sound data; The specific method of screening the neural network recognition model according to the environmental feature data to recognize and classify the environmental sound data includes the following steps: Recognize the possible environmental types according to the environmental feature data; Screen the neural network recognition model corresponding to the possible environmental type according to the possible environmental type; Recognize and classify the environmental sound data according to the screened neural network recognition model corresponding to the possible environmental type; The specific method of recognizing the possible environmental types according to the environmental feature data includes the following steps: Calculate the probability value p between the on-site environment where the monitoring terminal is located and multiple pre-stored environmental types respectively according to the environmental feature data; Use the pre-set environmental type with the probability value p greater than the pre-set probability threshold as the possible environmental type; The specific method of recognizing and classifying the environmental sound data according to the screened neural network recognition model corresponding to the possible environmental type includes the following steps: Calculate the average probability According to the probability average value And classify the screened neural network recognition models corresponding to possible environmental types into preferred recognition models and alternative recognition models according to the probability value p of possible environmental types; Recognize and classify the environmental sound data according to the preferred recognition model and the alternative recognition model; Among them, the server stores multiple pre-stored environment types, calculates the similarity between the on-site environment where the monitoring terminal is located and the multiple pre-stored environment types according to the environmental characteristic data respectively, and calculates a probability value according to the similarity.

4. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed, it implements the environmental sound recognition and classification method as described in claim 3.

Citation Information

Patent Citations

  • Audio frequency setting method and audio frequency setting device

    CN106817653A

  • Speech endpoint detection method based on location information

    CN107564546A

  • Voice recognition method and device, storage medium, electronic equipment and vehicle

    CN114420163A

  • Automatically selecting a sound recognition model for an environment based on audio data and image data associated with the environment

    US20240153524A1