A sound event detection method, device, apparatus and storage medium

By combining the acoustic features of audio data with a pre-constructed sound event relationship graph, and using a graph convolutional neural network for feature fusion, the problem of low sound event detection accuracy is solved, and the application scope of detection scenarios is expanded.

CN115985294BActive Publication Date: 2025-12-23AUTOMOBILE RES INST OF TSINGHUA UNIV IN SUZHOU XIANGCHENG
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211298179.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-21
Publication Date
2025-12-23
Estimated Expiration
2042-10-21

AI Technical Summary

Technical Problem

Existing sound event detection technologies suffer from poor detection accuracy in polyphonic sound event detection scenarios, especially when multiple sound events overlap in time, resulting in limited detection range and low accuracy.

Method used

By acquiring the acoustic features of the audio data to be detected and combining them with the sound event relationship features determined by a pre-constructed sound event relationship graph based on the statistical results of the audio dataset, a graph convolutional neural network is used for feature extraction and fusion to improve the detection accuracy.

Benefits of technology

While improving the accuracy of sound event detection, it expands the application scope of detection scenarios and enhances the detection capability in complex sound environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115985294B_ABST
    Figure CN115985294B_ABST
Patent Text Reader

Abstract

The application discloses a sound event detection method, device and equipment and a storage medium. The method comprises the following steps: acquiring to-be-detected audio data, and extracting acoustic characteristics of the to-be-detected audio data; determining a sound event detection result of the to-be-detected audio data according to the acoustic characteristics and a pre-determined sound event relationship characteristic; wherein the sound event relationship characteristic is determined based on a pre-constructed sound event relationship graph; and the sound event relationship graph is determined according to a statistical result of sound events in an audio data set. The technical scheme solves the problems of limited application range and low accuracy of sound event detection, and can effectively expand the application range of the detection scene while improving the accuracy of sound event detection.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech detection, and in particular to a sound event detection method, device, equipment and storage medium. BACKGROUND

[0002] With the development of intelligent and networked vehicles, the sensing ability is required to be higher and higher. In addition to visual technology, sound signals can also provide a lot of useful information and have the advantages of not needing light, not being disturbed by rain, and not being disturbed by dark conditions. The audio detection device with sound event detection technology can be widely used in the cabin of intelligent vehicles due to its low cost, small size, convenient installation, strong reliability, not easy to damage, and simple maintenance.

[0003] At present, the sound event detection scheme is mainly based on a deep learning model, such as obtaining a sound event detection result through a convolutional neural network and a recurrent neural network. The convolutional neural network is good at capturing local features of a sound signal and can well classify the sound signal. The combination of CNN and RNN forms a convolutional recurrent neural network, which can obtain more ideal detection performance and has become the mainstream model for sound event detection.

[0004] However, the prior art extracts independent features through a convolutional recurrent neural network and other models for sound event detection, that is, a single sound event is detected in an audio signal. The prior art does not consider the relationship between sound events in a sound scene, so there is still a certain gap in the sound event detection accuracy in a polyphonic sound event detection scene, especially when multiple sound events overlap in time. SUMMARY

[0005] The present application provides a sound event detection method, device, equipment and storage medium to solve the problems of limited application range and low accuracy of sound event detection, which can effectively expand the application range of the detection scene while improving the accuracy of sound event detection.

[0006] According to an aspect of the present application, a sound event detection method is provided, which comprises:

[0007] obtaining to-be-detected audio data and extracting acoustic features of the to-be-detected audio data;

[0008] determining a sound event detection result of the to-be-detected audio data according to the acoustic features and a predetermined sound event relationship feature;

[0009] The sound event relationship feature is determined based on a pre-constructed sound event relationship graph, and the sound event relationship graph is determined according to the statistical results of sound events in an audio data set.

[0010] According to another aspect of the present application, there is provided a sound event detection apparatus, comprising:

[0011] an acoustic feature extraction module configured to obtain to-be-detected audio data and extract acoustic features of the to-be-detected audio data;

[0012] a sound event detection result determination module configured to determine a sound event detection result of the to-be-detected audio data according to the acoustic features and predetermined sound event relationship features;

[0013] wherein the sound event relationship features are determined based on a pre-constructed sound event relationship graph, and the sound event relationship graph is determined according to statistical results of sound events in an audio data set.

[0014] According to another aspect of the present application, there is provided an electronic device, comprising:

[0015] at least one processor; and

[0016] a memory in communication with the at least one processor; wherein

[0017] the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the sound event detection method according to any one of the embodiments of the present application.

[0018] According to another aspect of the present application, there is provided a computer readable storage medium storing computer instructions for enabling a processor to perform the sound event detection method according to any one of the embodiments of the present application.

[0019] The technical solution of the embodiments of the present application determines a sound event detection result of to-be-detected audio data by comprehensively considering acoustic features of the to-be-detected audio data and prior sound event relationship features, thereby solving the problems of limited application range and low accuracy of sound event detection, and effectively expanding the application range of the detection scene while improving the accuracy of sound event detection.

[0020] It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0021] In order to make the technical solution in the embodiments of the present application clearer, the following will briefly introduce the drawings needed in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort on the basis of these drawings.

[0022] Figure 1A is a flow chart of a sound event detection method according to the first embodiment of the present application;

[0023] Figure 1B is a sound event relationship diagram according to the first embodiment of the present application;

[0024] Figure 2 is a flow chart of a sound event detection method according to the second embodiment of the present application;

[0025] Figure 3 is a structural schematic diagram of a sound event detection device according to the third embodiment of the present application;

[0026] Figure 4 is a structural schematic diagram of an electronic device implementing the sound event detection method according to the present application. DETAILED DESCRIPTION

[0027] In order to make the technical solution in the embodiments of the present application clearer, the following will briefly introduce the drawings needed in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort on the basis of these drawings.

[0028] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices. The acquisition, storage, use, processing and the like of data in the technical solution of the present application comply with the relevant provisions of national laws and regulations.

[0029] Embodiment One

[0030] Figure 1A A flowchart of a sound event detection method is provided for Embodiment One of the present application. The present embodiment can be applied to a sound event detection scenario in a vehicle cabin. The method can be executed by a sound event detection device, which can be implemented in the form of hardware and / or software. The device can be configured in an electronic device. As shown in Figure 1A the method comprises:

[0031] In S110, the sound event detection device acquires audio data to be detected and extracts acoustic features of the audio data to be detected.

[0032] The present embodiment can be executed by a vehicle-mounted system. The vehicle-mounted system can acquire audio data to be detected in the vehicle cabin through a microphone or other voice pickup device. The vehicle-mounted system can perform voice signal processing on the audio data to be detected and extract acoustic features such as amplitude and phase of the audio data to be detected. The vehicle-mounted system can directly apply the acoustic features output by the voice signal processing, or further process the acoustic features obtained by the voice signal processing based on a convolutional neural network to extract deeper acoustic features.

[0033] In S120, the sound event detection device determines a sound event detection result of the audio data to be detected based on the acoustic features and pre-determined sound event relationship features.

[0034] It is easy to understand that the sound event can be a sound carrying physical event information, such as a yawn sound, a sneeze sound, and a mobile phone ringtone, etc. The sound event relationship features can be pre-stored in the vehicle-mounted system. The sound event relationship features can be determined based on a pre-constructed sound event relationship graph. The sound event relationship graph can be determined based on statistical results of sound events in an audio data set. The audio data set can be a public audio data set or an audio data set collected and produced by a developer as needed. The audio data set can include audio data and label information associated with the audio data. The audio data in the audio data set can be collected in a vehicle cabin, and the label information of the audio data can include the type of sound events in the audio data, the start and end time of each sound event, etc. Using a computer, a server, or other electronic device, the sound events in the audio data set can be counted to obtain the number of occurrences of each type of sound event, the number of simultaneous occurrences of each type of sound event, and the number of sequential occurrences of each type of sound event, etc. Based on the statistical results of the sound events in the audio data set, the electronic device can construct a sound event relationship graph to represent the association between each type of sound event.

[0035] Figure 1B is a sound event relationship diagram provided according to the present application. After obtaining the sound event relationship, the electronic device can drawFigure 1B The electronic device can extract a sound event relationship feature implied by the sound event relationship graph to assist in the acoustic feature of the to-be-detected audio data, so as to realize sound event detection in the to-be-detected audio data.

[0036] The vehicle-mounted system can fuse the extracted acoustic feature and the pre-determined sound event relationship feature. According to the fused feature, the vehicle-mounted system can determine a sound event detection result of the to-be-detected audio data based on a sound event classification model. The sound event detection result can include a type of sound event existing in the to-be-detected audio data and a probability of occurrence of each type of sound event.

[0037] In one possible solution, the determination process of the sound event relationship feature includes:

[0038] Obtaining label information of each audio data in the audio data set;

[0039] According to the label information of each audio data, determining a time sequence relationship between each type of sound event and a probability of occurrence matched with the time sequence relationship;

[0040] According to the time sequence relationship and the probability of occurrence matched with the time sequence relationship, constructing a sound event relationship graph;

[0041] According to the sound event relationship graph, determining a sound event relationship feature based on a graph convolutional neural network.

[0042] The label information can include distribution information of sound events in the audio data, i.e., a time sequence relationship and start and end times of the sound events. According to the distribution information of the sound events, the electronic device can statistically determine a time sequence relationship between each type of sound event in the audio data set and a probability of occurrence of each time sequence relationship. For example, there are 10,000 pieces of audio data in the audio data set, among which 500 pieces of audio data each contain an I-type sound event and an II-type sound event. Among the 500 pieces of audio data, 300 pieces of I-type sound events occur before II-type sound events, and 200 pieces of II-type sound events occur before I-type sound events. The electronic device can statistically determine that the probability of occurrence of I-type sound events before II-type sound events is 0.6, and the probability of occurrence of II-type sound events before I-type sound events is 0.4.

[0043] The electronic device can take each type of sound event as a node, construct an edge between each node according to a time sequence relationship between the sound events, and take a probability of occurrence matched with the time sequence relationship as a weight of the edge between the nodes, so as to obtain a sound event relationship graph. The time sequence relationship between the sound events can be directed or undirected, and therefore, the sound event relationship graph can be a directed graph or an undirected graph.

[0044] After obtaining the sound event relation graph, the electronic device can determine an adjacency matrix, a degree matrix, and a feature vector matrix of nodes according to the sound event relation graph. According to the adjacency matrix, the degree matrix, and the feature vector matrix of the nodes, relation feature extraction is performed based on a graph convolutional neural network, and the electronic device can obtain a sound event relation feature.

[0045] On the basis of the above scheme, the label information includes a sound event type existing in the audio data and start and end times of the sound event.

[0046] The label information of each audio data is used to determine a time sequence relationship between sound events of each type and an occurrence probability matching the time sequence relationship, including:

[0047] The occurrence number of each type of sound event is determined according to the sound event type existing in each audio data, and the time sequence relationship between sound events of each type is determined according to the sound event type existing in each audio data and the start and end times of the sound event in each audio data.

[0048] The occurrence probability matching the time sequence relationship is determined according to the time sequence relationship between sound events of each type and the occurrence number of sound events of each type.

[0049] The electronic device such as a computer or a server can obtain label information of audio data from an audio data set. According to the label information, the occurrence number of each type of sound event, the total number of sound event occurrences, the number of simultaneous occurrences of multiple sound events in each audio data, and the number of sequential occurrences are counted. Specifically, the electronic device can record each type of sound event as L i , i = 0, 1, 2, 3,..., which represents a sound event type index. The case where two types of sound events exist in one audio data is recorded as L ij , i, j = 0, 1, 2, 3,..., and i ≠ j, both i and j represent a sound event type index. L ij represents that the sound event L i and the sound event L j exist in the same audio data, that is, both the simultaneous occurrence and the sequential occurrence of the two sound events are considered. In order to ensure that the calculated sound event sequence in the same audio data conforms to the real rules, the duration of a single audio data can be set within a certain time range, for example, 30s-60s, to ensure the integrity and relevance of the audio data information. According to the sound event type existing in each audio data, the occurrence number of each type of sound event is counted and recorded as X i , which represents the occurrence number of the sound event L iThe number of times of occurrence. The number of times of occurrence of each type of sound event is added to obtain the total number of times of occurrence of sound events, X represents the total number of times of occurrence of sound events. The number of times of occurrence of both types of sound events is denoted as X ij , which represents the number of times of occurrence of sound event L i and sound event L j in the same audio.

[0050] The electronic device can calculate the occurrence probability of each type of sound event according to the number of times of occurrence of each type of sound event, and then calculate the probability of simultaneous occurrence and sequential occurrence of various sound events through the prior probability formula. Specifically, the occurrence probability Y i of each type of sound event is calculated, the occurrence probability Y i of sound event L i is calculated according to the number of times of occurrence X i and the total number of times of occurrence X, and the occurrence probability Y i of sound event L i is calculated. Y ij = X i / X. The probability P ij of simultaneous occurrence and sequential occurrence of various sound events is calculated according to the number of times of occurrence X i of each type of sound event and the number of times of occurrence X j of both types of sound events, the probability P ij of occurrence of sound event L j under the premise of occurrence of sound event L i is calculated, P ij (L i |L ij ) = X ji / X ij . It should be noted that P i and P j are not necessarily equal, and the meanings expressed are also different, P ji represents the probability of occurrence of sound event L j under the premise of occurrence of sound event L i .

[0051] In the present scheme, optionally, the sound event relationship graph is constructed according to the time sequence relationship and the occurrence probability matched with the time sequence relationship, comprising:

[0052] Taking each type of sound event as a node, the directed edges of each node are determined according to the time sequence relationship;

[0053] The occurrence probability matched with the time sequence relationship is taken as the weight of each directed edge to generate the sound event relationship graph.

[0054] Electronic devices can use various types of sound events as nodes, temporal relationships as directed edges connecting these nodes, and the probability of occurrence matching the temporal relationship as the weight of each directed edge, thereby completing tasks such as... Figure 1B The construction of the sound event relationship diagram shown.

[0055] In a preferred embodiment, before constructing the sound event relationship graph based on temporal relationships and the occurrence probabilities matching those temporal relationships, the method further includes:

[0056] If the type of sound event associated with the current time sequence relationship is in the pre-set set of important events, and the probability of occurrence matching the current time sequence relationship is less than the first preset probability threshold, then the probability of occurrence matching the current time sequence relationship is corrected according to the preset probability correction principle.

[0057] If there is a bidirectional temporal relationship between the sound event types associated with the current temporal relationship, and the probability of occurrence of matching the bidirectional temporal relationship is less than the second preset probability threshold, then the probability of occurrence of the bidirectional temporal relationship is corrected to 0.

[0058] In real-world applications, there are sound events that occur with low probability but are of high importance, such as vehicle collision sounds or vehicle explosion sounds. If sound event L... j If a rare but important event is included in the set of important events, then it can be classified as sound event L. j The probability of occurrence of the associated temporal relationship is adjusted, for example, by setting a large value of 1 for the probability of occurrence, in order to improve the reliability of sound event prediction.

[0059] There is a bidirectional temporal relationship between two types of sound events in the audio data, for example... Figure 1B Coughing and talking sounds, cell phone ringtones and cell phone vibrations. Sound Event L i Sound event L under the premise of occurrence j The probability of occurrence P ij Harmony and Sound Event L j Sound event L under the premise of occurrence i The probability of occurrence P ji When both are higher than the second preset probability threshold, it indicates that the two types of sound events have a strong correlation. If the probability of occurrence P... ij Very low, but P ji But it's very high, and Y j If it is very small, it indicates that the sound event L j This is a rare event; a sound event L occurred. i Sound event L may not occur. j But the sound event L j It has a high probability of causing sound event L i This occurs. Therefore, the sound event L can be preserved. jPoint to the sound event L i The directed edge of the sound event L j Whether the event L

[0060] If the probability P ij And P ji Both are lower than the second preset probability threshold, it means that the correlation of the two types of sound events is low. In order to avoid the adverse effect of the low occurrence probability of the two-way time sequence relationship on the sound event detection, and to avoid the overfitting of the model in the graph convolution network training process, the electronic device can filter the occurrence probability by using the second preset probability threshold, and correct the occurrence probability of the two-way time sequence relationship whose occurrence probability is less than the second preset probability threshold to 0, that is, the correction principle of the occurrence probability of the two-way time sequence relationship can be represented as:

[0061] Wherein, α represents the second preset probability threshold, α∈(0,1).

[0062] The above scheme can avoid the influence of accidental sound events on the sound event correlation system, and ensure the effectiveness of the sound event correlation system about rare but important events.

[0063] The technical scheme determines the sound event detection result of the to-be-detected audio data by comprehensively considering the acoustic features of the to-be-detected audio data and the prior sound event relationship features, solves the problems of limited application range and low accuracy of sound event detection, and can effectively expand the application range of the detection scene while improving the accuracy of sound event detection.

[0064] Embodiment two

[0065] Figure 2 A flowchart of a sound event detection method provided by the second embodiment of the present application is provided, and the embodiment is refined based on the above-mentioned embodiment. As shown in the figure, the method comprises: Figure 2

[0066] S210, acquiring to-be-detected audio data, and determining the signal processing features of the to-be-detected audio data.

[0067] In the scheme, the signal processing features can include phase features and amplitude features. The sampling frequency of the to-be-detected audio data can be a preset frequency, for example, 44.1KHz. The vehicle-mounted system can perform frame processing on the to-be-detected audio data, for example, set the frame length T to 512 sampling points and the frame shift J to 256 sampling points. The frame data can be represented by the following formula:

[0068] x k (m)=x[kJ:KJ+M-1],1<m<M,k∈[0,Z];

[0069] ​Where k represents the audio frame index value, and Z represents the frame number. Let M represent the speech vector of the audio data to be detected at frame k, M represent the length of the window function, and m represent the frequency band.

[0070] The in-vehicle system can perform windowing processing on the framed data. For example, the window function can be a Hamming window, and the expression for the Hanning window function can be shown below:

[0071]

[0072] Where 'a' represents a coefficient, and 'a' can typically be set to 0.53836, L win Indicates window length, l win Indicates frequency band.

[0073] Frame-by-frame windowing operation on the frame-by-frame data can yield the frame-by-frame windowed data y. k (m)=x k (m)w(l win ), where x k (m) represents the frame data to be windowed.

[0074] After obtaining the framed and windowed data, the vehicle system can use a fast algorithm of Discrete Fourier Transform (FFT) to transform the framed and windowed data from the time domain to the frequency domain, thereby obtaining the frequency domain signal S of the audio data to be detected. k (l fft S k (l fft The calculation formula for ) can be Among them, 1 <m<M,1≤l fft ≤L fft L fft The length of the Fast Fourier Transform is represented by l. fft Indicates frequency components, Let f be a complex number, representing the frequency domain signal after FFT transformation, with the real part representing the l-th frequency band in the m-th frequency band of the k-th frame. fft The amplitude of each frequency component, where the imaginary part represents the l-th frequency component in the m-th frequency band of the k-th frame. fft The phase of each frequency component.

[0075] The vehicle-mounted system can perform a modulus extraction operation on the frequency domain signal to obtain the amplitude characteristics U of the audio data to be detected. k (l fft )=|S k (l fft Simultaneously, the phase angle of the frequency domain signal can be extracted to obtain the phase feature θ = argtan(S) of the audio data to be detected. k (l fft )).

[0076] S220, determine the acoustic feature of the to-be-detected audio data based on the convolutional neural network according to the signal processing feature of the to-be-detected audio data.

[0077] In order to further extract deep acoustic features, the vehicle-mounted system can combine the amplitude feature and the phase feature into a signal processing feature, and use the signal processing feature as the input of the convolutional neural network to extract the acoustic feature of the to-be-detected audio data. The convolutional neural network can be constructed based on network structures such as AlexNet, VGG-Net, or ResNet, or can be customized according to actual application needs. Specifically, the convolutional neural network can be flexibly constructed according to convolutional layers, pooling layers, standardization layers, fully connected layers, dimension exchange layers, and reshaping layers. Network parameters such as convolution kernel size, step size, pooling method, activation function type, and optimizer type can also be customized. The convolutional neural network can be pre-trained and deployed on the vehicle-mounted system to extract the acoustic feature of the to-be-detected audio data.

[0078] S230, fuse the acoustic feature and the sound event relationship feature to obtain a joint feature.

[0079] The vehicle-mounted system can fuse the acoustic feature output by the convolutional neural network with the sound event relationship feature. The vehicle-mounted system can directly splice and combine the acoustic feature and the sound event relationship feature to obtain a joint feature. The vehicle-mounted system can also perform matrix calculation on the acoustic feature and the sound event relationship feature to obtain a joint feature. In one specific example, the sound event relationship feature can be represented as an N x M matrix, where N represents the number of sound event types and M represents the feature dimension. In order to fully fuse the sound event relationship feature, the acoustic feature can be an M x N matrix. Based on the above scheme, the feature fusion method can include the following three methods:

[0080] (1) Adjust the sound event relationship feature and the acoustic feature to the same dimension by matrix transposition, and then add the corresponding elements in the sound event relationship feature and the acoustic feature with the same dimension to obtain a joint feature;

[0081] (2) Adjust the sound event relationship feature and the acoustic feature to the same dimension by matrix transposition, and then multiply the corresponding elements in the sound event relationship feature and the acoustic feature with the same dimension to obtain a joint feature;

[0082] (3) Directly multiply the sound event relationship feature and the acoustic feature in the matrix to obtain a joint feature.

[0083] This scheme can fuse the prior sound event relationship feature and the acoustic feature to enrich the features of the sound event, thereby improving the detection probability of the sound event.

[0084] S240, according to the joint feature, determining the sound event detection result of the to-be-detected audio data based on the sound event classification network.

[0085] The vehicle-mounted system can take the joint feature as the input of the sound event classification network, and determine the sound event detection result of the to-be-detected audio data according to the output result of the sound event classification network. The sound event classification network can be pre-trained and used to realize the identification of sound events. The sound event classification network can be built based on a neural network or a convolutional neural network. The output result of the sound event classification network can include the occurrence probability of each type of sound event. The vehicle-mounted system can take the sound event type with the maximum occurrence probability as the sound event detection result of the audio frame of the to-be-detected audio data, or take all sound event types with an occurrence probability greater than a preset probability threshold as the sound event detection result of the audio frame of the to-be-detected audio data. For example, the output result of the sound event classification network includes the occurrence probability of 7 types of sound events, wherein the occurrence probability of type I sound event and type II sound event is greater than the preset probability threshold, and the sound event detection result includes type I sound event and type II sound event, that is, there are type I sound event and type II sound event in the audio frame.

[0086] The technical scheme comprehensively determines the sound event detection result of the to-be-detected audio data by combining the acoustic feature of the to-be-detected audio data and the prior sound event relationship feature, solves the problems of limited application range and low accuracy of sound event detection, and can effectively expand the application range of the detection scene while improving the accuracy of sound event detection.

[0087] Embodiment three

[0088] Figure 3 A structure schematic diagram of a sound event detection device provided by the third embodiment of the present application is shown in FIG. 3. As shown in the figure, the device includes: Figure 3

[0089] The acoustic feature extraction module 310 is configured to acquire the to-be-detected audio data and extract the acoustic feature of the to-be-detected audio data.

[0090] The detection result determination module 320 is configured to determine the sound event detection result of the to-be-detected audio data according to the acoustic feature and the pre-determined sound event relationship feature.

[0091] The sound event relationship feature is determined based on a pre-constructed sound event relationship graph, and the sound event relationship graph is determined according to the statistical result of sound events in an audio data set.

[0092] ​In the scheme, optionally, the device further comprises a relationship feature determination module, comprising:

[0093] A label information acquisition unit is configured to acquire label information of each audio data in the audio data set.

[0094] A relationship and probability determination unit is configured to determine a time sequence relationship between each type of sound event and an occurrence probability matching the time sequence relationship according to the label information of each audio data.

[0095] A relationship graph construction unit is configured to construct a sound event relationship graph according to the time sequence relationship and the occurrence probability matching the time sequence relationship.

[0096] A relationship feature determination unit is configured to determine a sound event relationship feature based on a graph convolutional neural network according to the sound event relationship graph.

[0097] On the basis of the above scheme, the label information comprises a type of sound event existing in the audio data and start and end times of the sound event.

[0098] The relationship and probability determination unit comprises:

[0099] A time sequence relationship determination subunit is configured to determine a number of occurrences of each type of sound event according to the type of sound event existing in each audio data, and determine a time sequence relationship between each type of sound event according to the type of sound event existing in each audio data and the start and end times of the sound event in each audio data.

[0100] An occurrence probability determination subunit is configured to determine an occurrence probability matching the time sequence relationship according to the time sequence relationship between each type of sound event and the number of occurrences of each type of sound event.

[0101] In a feasible scheme, the relationship graph construction unit comprises:

[0102] A directed edge determination subunit is configured to determine directed edges of each node according to the time sequence relationship, with each type of sound event as the node.

[0103] A relationship graph generation subunit is configured to generate the sound event relationship graph by taking the occurrence probability matching the time sequence relationship as a weight of each directed edge.

[0104] In a preferred scheme, the relationship feature determination module further comprises:

[0105] A first occurrence probability correction unit is configured to correct the occurrence probability matching the current time sequence relationship according to a preset probability correction principle, if a type of sound event associated with the current time sequence relationship is in a set of pre-set important events and the occurrence probability matching the current time sequence relationship is less than a first preset probability threshold.

[0106] The second occurrence probability correction unit is configured to correct the occurrence probabilities of the bidirectional time sequence relationship to 0 if there is a bidirectional time sequence relationship between the sound event types associated with the current time sequence relationship, and the occurrence probabilities matched with the bidirectional time sequence relationship are all less than a second preset probability threshold.

[0107] Optionally, the acoustic feature extraction module comprises:

[0108] The signal processing feature determination unit is configured to acquire the to-be-detected audio data and determine a signal processing feature of the to-be-detected audio data, wherein the signal processing feature comprises a phase feature and an amplitude feature.

[0109] The acoustic feature determination unit is configured to determine an acoustic feature of the to-be-detected audio data based on a convolutional neural network according to the signal processing feature of the to-be-detected audio data.

[0110] On the basis of the above scheme, the detection result determination module comprises:

[0111] The joint feature determination unit is configured to fuse the acoustic feature and the sound event relationship feature to obtain a joint feature.

[0112] The detection result determination unit is configured to determine a sound event detection result of the to-be-detected audio data based on a sound event classification network according to the joint feature.

[0113] The sound event detection device provided by the embodiment can execute the sound event detection method provided by any of the embodiments of the present application, and has the corresponding function modules and beneficial effects of the execution method.

[0114] Embodiment four

[0115] Figure 4 A structural schematic diagram of an electronic device 410 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices (such as headsets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present application described and / or claimed in this document.

[0116] As Figure 4As shown, the electronic device 410 includes at least one processor 411, and a memory, such as a read-only memory (ROM) 412, a random access memory (RAM) 413, etc., connected to the at least one processor 411 in communication. The memory stores computer programs executable by the at least one processor 411, and the processor 411 can perform various appropriate actions and processes according to the computer programs stored in the read-only memory (ROM) 412 or loaded from the storage unit 418 into the random access memory (RAM) 413. In the RAM 413, various programs and data required for the operation of the electronic device 410 can also be stored. The processor 411, the ROM 412, and the RAM 413 are connected to each other through a bus 414. An input / output (I / O) interface 415 is also connected to the bus 414.

[0117] Various components in the electronic device 410 are connected to the I / O interface 415, including an input unit 416, such as a keyboard, a mouse, etc., an output unit 417, such as various types of displays, a speaker, etc., a storage unit 418, such as a magnetic disk, an optical disk, etc., and a communication unit 419, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 419 allows the electronic device 410 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0118] The processor 411 can be various general and / or special purpose processing components having processing and computing capabilities. Some examples of the processor 411 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 411 performs various methods and processes described above, such as the sound event detection method.

[0119] In some embodiments, the sound event detection method can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as the storage unit 418. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 410 via the ROM 412 and / or the communication unit 419. When the computer program is loaded onto the RAM 413 and executed by the processor 411, one or more steps of the sound event detection method described above can be performed. Alternatively, in other embodiments, the processor 411 can be configured to perform the sound event detection method by any other appropriate means, such as by means of firmware.

[0120] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a load programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0121] Computer programs used to implement the processes of the application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer program, when executed, can cause instructions defined in the flow charts and / or block diagrams to be implemented on the computer or other programmable apparatus. The computer programs can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, partially on a machine and partially on a remote machine or entirely on a remote machine or server.

[0122] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store computer programs for use by or in connection with an instruction execution system, apparatus, or device. Computer-readable storage media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0123] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0124] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0125] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. A server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service.

[0126] It should be understood that various forms of flow shown above can be used, with steps reordered, added, or removed. For example, the steps recited in the present invention can be performed in parallel, in series, or in a different order, without limitation herein, as long as the desired results of the technical solutions of the present invention can be achieved.

[0127] The specific embodiments described above are not intended to be limiting, but rather to illustrate the principles of the present invention. Variations, combinations, sub-combinations, and modifications can be made to the described embodiments within the spirit and scope of the present invention. Accordingly, other implementations are within the scope of the following claims.

Claims

1. A method for detecting sound events, characterized in that, The method includes: Acquire the audio data to be detected and extract its acoustic features; Based on acoustic characteristics and pre-determined sound event relationship characteristics, determine the sound event detection results of the audio data to be detected; The sound event relationship features are determined based on a pre-constructed sound event relationship graph; the sound event relationship graph is determined based on the statistical results of sound events in the audio dataset. The process of determining the relationship features of the sound events includes: Obtain the label information of each audio data in the audio dataset; the label information includes the type of sound event present in the audio data and the start and end times of the sound event; Based on the types of sound events present in each audio data, determine the frequency of occurrence of each type of sound event; and based on the types of sound events present in each audio data and the start and end times of the sound events in each audio data, determine the temporal relationship between the types of sound events. Based on the temporal relationship between different types of sound events and the frequency of occurrence of each type of sound event, determine the probability of occurrence that matches the temporal relationship; If the type of sound event associated with the current time sequence relationship is in the preset set of important events, and the probability of occurrence matching the current time sequence relationship is less than the first preset probability threshold, then the probability of occurrence matching the current time sequence relationship is corrected to the preset maximum probability of occurrence. If there is a bidirectional temporal relationship between the sound event types associated with the current temporal relationship, and the probability of occurrence of matching the bidirectional temporal relationship is less than the second preset probability threshold, then the probability of occurrence of the bidirectional temporal relationship is corrected to 0. Construct a sound event relationship graph based on temporal relationships and the probability of occurrence matching temporal relationships; Based on the sound event relationship graph, the sound event relationship features are determined using a graph convolutional neural network.

2. The method according to claim 1, characterized in that, The step of constructing a sound event relationship graph based on temporal relationships and the probability of occurrence matching the temporal relationships includes: Using various types of sound events as nodes, the directed edges of each node are determined based on temporal relationships; The probability of occurrence matching the temporal relationship is used as the weight of each directed edge to generate a sound event relationship graph.

3. The method according to claim 1, characterized in that, The process of acquiring the audio data to be detected and extracting its acoustic features includes: Acquire the audio data to be detected and determine the signal processing features of the audio data to be detected; wherein, the signal processing features include phase features and amplitude features; Based on the signal processing characteristics of the audio data to be detected, the acoustic characteristics of the audio data to be detected are determined using a convolutional neural network.

4. The method according to claim 3, characterized in that, The step of determining the sound event detection result of the audio data to be detected based on acoustic features and pre-determined sound event relationship features includes: By fusing acoustic features with sound event relationship features, joint features are obtained; Based on joint features and a sound event classification network, the sound event detection results of the audio data to be detected are determined.

5. A sound event detection device, characterized in that, include: The acoustic feature extraction module is used to acquire the audio data to be detected and extract the acoustic features of the audio data to be detected. The detection result determination module is used to determine the sound event detection results of the audio data to be detected based on acoustic features and pre-determined sound event relationship features; wherein, the sound event relationship features are determined based on a pre-constructed sound event relationship graph; and the sound event relationship graph is determined based on the statistical results of sound events in the audio dataset. The device further includes a relation feature determination module, comprising: The tag information acquisition unit is used to acquire the tag information of each audio data in the audio dataset; the tag information includes the type of sound event present in the audio data and the start and end time of the sound event; The relationship and probability determination unit is used to determine the occurrence frequency of each type of sound event based on the types of sound events present in each audio data; and to determine the temporal relationship between each type of sound event based on the types of sound events present in each audio data and the start and end times of the sound events in each audio data; and to determine the occurrence probability matching the temporal relationship based on the temporal relationship between each type of sound event and the occurrence frequency of each type of sound event. The first occurrence probability correction unit is used to correct the occurrence probability matching the current time sequence relationship to the preset maximum occurrence probability if the sound event type associated with the current time sequence relationship is in the preset important event set and the occurrence probability matching the current time sequence relationship is less than the first preset probability threshold. The second probability correction unit is used to correct the probability of occurrence of the two-way timing relationship to 0 if there is a two-way timing relationship between the sound event types associated with the current timing relationship and the probability of occurrence of the two-way timing relationship is less than the second preset probability threshold. The relationship graph construction unit is used to construct a sound event relationship graph based on temporal relationships and the occurrence probability matching the temporal relationships; The relationship feature determination unit is used to determine the relationship features of sound events based on the sound event relationship graph and a graph convolutional neural network.

6. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the sound event detection method according to any one of claims 1-4.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the sound event detection method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Audio recognition method and device, computer equipment and computer readable storage medium

    CN113593606A