A sound class detection method, system, medium, and electronic device
Patent Information
- Application Number
- CN202510606419.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2045-05-12
AI Technical Summary
[0015]The sound category detection method provided in this application can process audio to obtain audio features of several audio segments, and use the image and text features of the target sound to obtain an audio-text matrix and an audio-image matrix. After deep fusion of the audio-text matrix and the audio-image matrix, the target sound corresponding to each audio segment is further obtained. Target sound detection of audio segments is achieved through audio-text-image multimodal information.
Smart Images

Figure CN120544601B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of artificial intelligence technology and relates to a sound category detection method, system, medium, and electronic device. Background Technology
[0002] In recent years, speech recognition technology has been widely used in applications such as sound monitoring, emotion recognition, identity verification, and fault sound monitoring. Sound detection is a crucial preprocessing step in speech recognition and has a significant impact on its performance. However, current sound detection methods primarily focus on detecting and recognizing single sounds. Therefore, providing a sound detection method that can detect different sound categories in audio under mixed sound conditions is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0003] The purpose of this application is to provide a sound category detection method, system, medium, and electronic device for detecting target sounds in audio segments.
[0004] In a first aspect, this application provides a sound category detection method, characterized in that the sound category detection method includes: acquiring audio features of several audio segments based on audio; acquiring text features and image features corresponding to at least two target sounds; acquiring an audio-text matrix based on the audio features and the text features; acquiring an audio-image matrix based on the audio features and the image features; performing deep fusion of the audio-text matrix and the audio-image matrix to obtain a correlation matrix; and acquiring the target sound corresponding to each audio segment based on the correlation matrix.
[0005] In one implementation of the first aspect, the process of obtaining audio features of several audio segments based on audio includes: using the VAD algorithm to cut the audio into several audio segments; and using an audio encoder to obtain the audio features of the several audio segments.
[0006] In one implementation of the first aspect, the process of obtaining image features corresponding to at least two target sounds includes: obtaining several images corresponding to the target sounds using a stable diffusion model; and obtaining image features corresponding to the several images using an image encoder.
[0007] In one implementation of the first aspect, the process of obtaining an audio-text matrix based on the audio features and the text features includes: performing matrix multiplication on the audio features and the text features, querying text information using the audio, and obtaining the audio-text matrix.
[0008] In one implementation of the first aspect, the process of obtaining an audio-image matrix based on the audio features and the image features includes: performing matrix multiplication on the audio features and the image features, querying image information using audio, and obtaining the audio-image matrix.
[0009] In one implementation of the first aspect, the process of deep fusing the audio-text matrix and the audio-image matrix to obtain a correlation matrix includes: inputting the audio-text matrix and the audio-image matrix into a cross-attention mechanism for deep correlation fusing to obtain the correlation matrix.
[0010] In one implementation of the first aspect, the process of obtaining the target sound corresponding to each audio segment based on the correlation matrix includes: classifying and mapping the correlation matrix to obtain the probability distribution corresponding to each target sound; and taking the category with the highest probability value among each audio segment as the corresponding target sound.
[0011] Secondly, this application provides a sound category detection system, the sound category detection system comprising: an audio feature acquisition module, used to acquire audio features of several audio segments based on audio; a text feature and image feature acquisition module, used to acquire text features and image features corresponding to at least two target sounds; an audio-text matrix acquisition module, used to acquire an audio-text matrix based on the audio features and the text features; an audio-image matrix acquisition module, used to acquire an audio-image matrix based on the audio features and the image features; a correlation matrix acquisition module, used to perform deep fusion of the audio-text matrix and the audio-image matrix to obtain a correlation matrix; and a target sound acquisition module, used to acquire the target sound corresponding to each audio segment based on the correlation matrix.
[0012] Thirdly, this application provides an electronic device, the electronic device comprising: a memory storing a computer program thereon; and a processor communicatively connected to the memory for executing the computer program to implement the above-described sound category detection method.
[0013] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by an electronic device, implements the above-described sound category detection method.
[0014] As described above, the sound category detection method of this application has the following beneficial effects:
[0015] The sound category detection method provided in this application can process audio to obtain audio features of several audio segments, and use the image and text features of the target sound to obtain an audio-text matrix and an audio-image matrix. After deep fusion of the audio-text matrix and the audio-image matrix, the target sound corresponding to each audio segment is further obtained. Target sound detection of audio segments is achieved through audio-text-image multimodal information. Attached Figure Description
[0016] Figure 1 The diagram shown illustrates an application scenario of the sound category detection method described in this application embodiment.
[0017] Figure 2 The diagram shows a process schematic of the sound category detection method described in the embodiments of this application.
[0018] Figure 3 The diagram shown is a schematic representation of the process of acquiring audio features as described in an embodiment of this application.
[0019] Figure 4 The diagram shown is a schematic representation of the process of acquiring image features as described in an embodiment of this application.
[0020] Figure 5 The diagram shown is a schematic representation of the process of acquiring the target sound as described in an embodiment of this application.
[0021] Figure 6 The diagram shown is a structural schematic of the sound category detection system described in an embodiment of this application.
[0022] Figure 7 The diagram shown is a structural schematic of the electronic device described in an embodiment of this application.
[0023] Component designation explanation
[0024] 11 mobile phones
[0025] 12 tablet computers
[0026] 13 Laptops
[0027] 2. Sound Category Detection System
[0028] 21 Audio Feature Acquisition Module
[0029] 22. Text and Image Feature Acquisition Module
[0030] 23 Audio Text Matrix Acquisition Module
[0031] 24 Audio-Image Matrix Acquisition Module
[0032] 25. Correlation Matrix Acquisition Module
[0033] 26 Target Sound Acquisition Module
[0034] 3 Electronic devices
[0035] 31 Memory
[0036] 32 processors
[0037] Steps S11 to S16
[0038] Steps S21 to S22
[0039] Steps S31 to S32
[0040] Steps S41 to S42 Detailed Implementation
[0041] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, unless otherwise specified, the following embodiments and features in the embodiments can be combined with each other.
[0042] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this application. Therefore, the drawings only show the components related to this application and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.
[0043] In recent years, speech recognition technology has been widely used in applications such as sound monitoring, emotion recognition, identity verification, and fault sound monitoring. Sound detection is a crucial preprocessing step in speech recognition and has a significant impact on its performance. However, current sound detection methods primarily focus on detecting and recognizing single sounds. Therefore, providing a sound detection method that can detect different sound categories in audio under mixed sound conditions is a problem that urgently needs to be solved by those skilled in the art.
[0044] To address at least the aforementioned problems, the following embodiments of this application provide a sound category detection method, which can be applied to, for example... Figure 1 The electronic device shown.
[0045] The electronic devices described in this application may include mobile phones 11 with wireless charging function, tablet computers 12, laptop computers 13, wearable devices, in-vehicle devices, augmented reality (AR) / virtual reality (VR) devices, ultra-mobile personal computers (UMPC), netbooks, personal digital assistants (PDAs), etc. The embodiments of this application do not impose any restrictions on the specific types of electronic devices.
[0046] For example, the electronic device may be a station (STAION, ST) in a WLAN with wireless charging capability, a cellular phone, cordless phone, Session Initiation Protocol (SIP) phone, Wireless Local Loop (WLL) station, Personal Digital Assistant (PDA) device, handheld device with wireless charging capability, computing device or other processing device, computer, laptop computer, handheld communication device, handheld computing device, and / or other devices for communication over a wireless system, as well as next-generation communication systems, such as mobile terminals in 5G networks, mobile terminals in future evolved Public Land Mobile Networks (PLMNs), or mobile terminals in future evolved Non-terrestrial Networks (NTNs).
[0047] For example, the electronic device can communicate with networks and other devices wirelessly. The aforementioned wireless communication can use any communication standard or protocol, including but not limited to Global System for Mobile Communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), BT, GNSS, WLAN, NFC, FM, and / or IR technologies. The GNSS may include Global Positioning System (GPS), Global Navigation Satellite System (GLONASS), BeiDou Navigation Satellite System (BDS), Quasi-Zenith Satellite System (QZSS), and / or Satellite Based Augmentation Systems (SBAS).
[0048] The technical solutions in the embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0049] Figure 2 This is a schematic diagram illustrating the process of a sound category detection method in one embodiment of this application. For example... Figure 2 As shown, the sound category detection method includes:
[0050] S11, Obtain the audio features of several audio segments based on the audio.
[0051] S12, acquire at least two text features and image features corresponding to the target sounds.
[0052] S13, Obtain an audio-text matrix based on the audio features and the text features.
[0053] S14, Obtain an audio-image matrix based on the audio features and the image features.
[0054] S15, perform deep fusion of the audio text matrix and the audio image matrix to obtain a correlation matrix.
[0055] S16, obtain the target sound corresponding to each audio segment based on the correlation matrix.
[0056] As described above, the sound category detection method provided in this application can process audio to obtain audio features of several audio segments, and use the image and text features of the target sound to obtain an audio-text matrix and an audio-image matrix. After deep fusion of the audio-text matrix and the audio-image matrix, the target sound corresponding to each audio segment is further obtained. Target sound detection of audio segments is achieved through audio-text-image multimodal information.
[0057] Figure 3 This is a schematic diagram illustrating the process of acquiring audio features in one embodiment of this application. For example... Figure 3 As shown, the process of obtaining audio features from several audio segments includes:
[0058] S21, the audio is segmented into several audio segments using the VAD (Voice Activity Detection) algorithm. The VAD algorithm performs speech annotation on the audio, finding the start and end points of speech and extracting the signal features needed for speech emotion recognition. The VAD algorithm can identify and eliminate long periods of silence in the audio, achieving the separation of effective speech signals from useless or noise signals.
[0059] For example, the VAD sound endpoint detection algorithm is used to cut a long audio file into M audio segments.
[0060] S22, the audio features of the several audio segments are obtained using an audio encoder. The audio encoder is the audio encoder in Audio CLIP. Audio CLIP combines deep learning with natural language processing to achieve cross-modal audio understanding. The audio encoder in Audio CLIP can extract audio features from audio segments to obtain their audio characteristics.
[0061] For example, the audio encoder in Audio CLIP is used to extract features from M audio segments to obtain M audio features.
[0062] In one embodiment of this application, the process of obtaining text features corresponding to at least two target sounds includes: setting the target sounds to be detected, and using a text encoder to obtain the text features corresponding to the target sounds. The target sounds include, for example, "bird calls," "whistles," and "human voices." The text encoder is a text encoder in Audio CLIP. The text encoder in Audio CLIP can extract text features from sounds to obtain the text features corresponding to the sounds.
[0063] For example, select N target sounds, and use the text encoder in Audio CLIP to extract features from the N target sounds to obtain N text features.
[0064] Figure 4 This is a schematic diagram illustrating the process of acquiring image features in one embodiment of this application. For example... Figure 4 As shown, the process of obtaining image features corresponding to at least two target sounds includes:
[0065] S31, several images corresponding to the target sound are obtained using a stable diffusion model. The stable diffusion model is a generative model. The stable diffusion model can generate images with relevant content based on the target sound.
[0066] For example, the stablediffusion model is used to generate images related to the content of the target sound, and N images are obtained.
[0067] S32, the image features corresponding to the plurality of images are obtained using an image encoder. The image encoder is an image encoder in Audio CLIP. The image encoder in Audio CLIP can obtain the corresponding image features based on the image generated from the content of the target sound.
[0068] For example, the image encoder in Audio CLIP is used to extract features from N images to obtain N image features.
[0069] In one embodiment of this application, the process of obtaining an audio-text matrix based on the audio features and the text features includes: performing matrix multiplication on the audio features and the text features, querying text information using audio, and obtaining the audio-text matrix.
[0070] For example, M audio features and N text features are multiplied by a matrix to obtain an audio-text matrix. The shape of the audio-text matrix is (M, N). Using audio as the primary modality of interest, information from the text is queried using the audio to obtain an audio-text relevance matrix.
[0071] In one embodiment of this application, the process of obtaining an audio-image matrix based on the audio features and the image features includes: performing matrix multiplication on the audio features and the image features, querying image information using audio, and obtaining the audio-image matrix.
[0072] For example, M audio features and N image features are multiplied by a matrix to obtain an audio-image matrix. The shape of the audio-text matrix is (M, N). Using audio as the primary modality of interest, information about the image is queried using the audio to obtain an audio-image relevance matrix.
[0073] In one embodiment of this application, the process of deep fusing the audio text matrix and the audio image matrix to obtain a correlation matrix includes: inputting the audio text matrix and the audio image matrix into a cross-attention mechanism for deep correlation fusing to obtain the correlation matrix. The cross-attention mechanism is a cross-attention layer between the encoder and decoder. The deep correlation fusing process in the cross-attention mechanism is as follows: in the cross-attention layer, the decoder adjusts the encoder output to obtain encoder information related to the current decoding position. Each position of the decoder generates a query vector Q to calculate attention weights for all positions of the encoder. All positions of the encoder generate a set of key vectors K and value vectors V. A dot product operation is performed using the query vector Q and the key vector K, and attention weights are obtained using a softmax function. The attention weights are multiplied by the value vector V, and the sum of the multiplication results is used to obtain the encoder-adjusted output.
[0074] For example, the audio-text matrix is used as Q, and the audio-image matrix is used as K and V. QKV is input into cross-attention for deep relevance fusion, and the relevance matrix is output. The shape of the relevance matrix is (M, N).
[0075] Figure 5 This is a schematic diagram illustrating the process of acquiring target sound in one embodiment of this application. For example... Figure 5 As shown, the process of obtaining the corresponding target sound for each audio segment based on the correlation matrix includes:
[0076] S41, perform classification mapping on the correlation matrix to obtain the probability distribution corresponding to each target sound. Use the softmax function to perform classification mapping on the correlation matrix and output the probability distribution value of each target sound.
[0077] For example, the correlation matrix is activated by the softmax function along the N dimension to obtain the probability distribution of each target sound.
[0078] S42, the category with the highest probability value among the audio segments is taken as the corresponding target sound.
[0079] For example, the index with the highest probability value in each audio segment is taken as the target sound, and each audio segment corresponds to one target sound, so the detection results of M audio segments can be obtained in the end.
[0080] The sound category detection method of this application will be specifically illustrated through an embodiment below. The specific implementation process is as follows:
[0081] Step 1: Input a long audio file, truncate it, and obtain audio features. Use the VAD (Voice Endpoint Detection) method to truncate the long audio file into M audio segments, and use the audio encoder in Audio CLIP to obtain M audio features.
[0082] Step 2: Set N target sounds to be detected, and use the text encoder in Audio CLIP to obtain N text features.
[0083] Step 3: Generate images based on the target sound and obtain corresponding image features. N images corresponding to the target sound content are generated using the stablediffusion model, and N image features are obtained using the image encoder in Audio CLIP.
[0084] Step 4: Obtain the audio-text matrix based on audio and text features. Perform matrix multiplication between the M audio features and N text features, taking the audio as the primary modality of interest, and use the audio to query text information to obtain an audio-text matrix of shape (M, N).
[0085] Step 5: Obtain the audio-image matrix based on audio and image features. Perform matrix multiplication on M audio features and N image features, taking audio as the primary mode of interest, and use the audio to query image information to obtain an audio-image matrix of shape (M, N).
[0086] Step 6: Deeply fuse the audio-text matrix and the audio-image matrix to obtain the correlation matrix. Using the audio-text matrix as Q and the audio-image matrix as K and V, input QKV into cross-attention for deep correlation fusion, outputting a correlation matrix of shape (M, N).
[0087] Step 7: Obtain the corresponding target sound for each audio segment based on the correlation matrix. Activate the correlation matrix using the softmax function along the N-dimensional axis to obtain the probability distribution for each target sound. Use the index with the highest probability value among the audio segments as the target sound. Each audio segment corresponds to one target sound, ultimately yielding the detection results for M audio segments.
[0088] In summary, the sound category detection method described in this application can process audio to obtain audio features of several audio segments, and use the image and text features of the target sound to obtain an audio-text matrix and an audio-image matrix. After deep fusion of the audio-text matrix and the audio-image matrix, the target sound corresponding to each audio segment is further obtained. Target sound detection of audio segments is achieved through audio-text-image multimodal information.
[0089] The scope of protection of the sound category detection method described in this application is not limited to the execution order of the steps listed in this embodiment. Any solution implemented by adding, subtracting, or replacing steps in the prior art based on the principles of this application is included within the scope of protection of this application.
[0090] This application also provides a sound category detection system, which can implement the sound category detection method described in this application. However, the implementation device of the sound category detection method described in this application includes, but is not limited to, the structure of the sound category detection system listed in this embodiment. All structural modifications and substitutions of the prior art made in accordance with the principles of this application are included within the protection scope of this application.
[0091] Figure 6 The diagram shown is a structural schematic of a sound category detection system according to an embodiment of this application. Figure 6 As shown, the sound category detection system 2 includes: an audio feature acquisition module 21, a text feature and image feature acquisition module 22, an audio-text matrix acquisition module 23, an audio-image matrix acquisition module 24, a correlation matrix acquisition module 25, and a target sound acquisition module 26.
[0092] The system includes the following modules: Audio Feature Acquisition Module 21, used to acquire audio features of several audio segments; Text Feature and Image Feature Acquisition Module 22, used to acquire text features and image features corresponding to at least two target sounds; Audio-Text Matrix Acquisition Module 23, used to acquire an audio-text matrix based on the audio features and text features; Audio-Image Matrix Acquisition Module 24, used to acquire an audio-image matrix based on the audio features and image features; Correlation Matrix Acquisition Module 25, used to perform deep fusion of the audio-text matrix and the audio-image matrix to obtain a correlation matrix; and Target Sound Acquisition Module 26, used to acquire the target sound corresponding to each audio segment based on the correlation matrix.
[0093] It should be noted that, Figure 6 The modules in the sound category detection system 2 shown are... Figure 2 The steps in the sound category detection method correspond one-to-one, and will not be elaborated here.
[0094] In the embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, or methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of modules / units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or units may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of apparatuses or modules or units may be electrical, mechanical, or other forms.
[0095] The modules / units described as separate components may or may not be physically separate. The components shown as modules / units may or may not be physical modules; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules / units can be selected to achieve the objectives of the embodiments of this application, depending on actual needs. For example, the functional modules / units in the various embodiments of this application may be integrated into one processing module, or each module / unit may exist physically separately, or two or more modules / units may be integrated into one module / unit.
[0096] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0097] This application also provides a computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the sound category detection method provided in this application. Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing a processor. The program can be stored in a computer-readable storage medium, which is a non-transitory medium, such as random access memory, read-only memory, flash memory, hard disk, solid-state drive, magnetic tape, floppy disk, optical disk, and any combination thereof. The above storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital video disc (DVD)), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0098] This application embodiment may also provide an electronic device. Figure 7 The diagram shown is a structural schematic of an electronic device 3 according to an embodiment of this application. Figure 7 As shown, in this embodiment, the electronic device 3 includes a memory 31 and a processor 32.
[0099] The memory 31 is used to store computer programs. In some possible implementations, the memory 31 may include various media capable of storing program code, such as ROM, RAM, magnetic disk, USB flash drive, memory card, or optical disk.
[0100] In this embodiment, memory 31 may include a computer system readable medium in the form of volatile memory, such as RAM and / or cache memory. Electronic device 3 may further include other removable / non-removable, volatile / non-volatile computer system storage media. Memory 31 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of this application.
[0101] The processor 32 is connected to the memory 31 and is used to execute the computer program stored in the memory 31 so that the electronic device 3 performs the sound category detection method.
[0102] For example, processor 32 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc. In other embodiments, processor 32 may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0103] The descriptions of the processes or structures corresponding to the above figures each have their own emphasis. For parts of a process or structure that are not described in detail, please refer to the relevant descriptions of other processes or structures.
[0104] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in this application should still be covered by the claims of this application.
Claims
1. A sound category detection method, characterized in that, The sound category detection method includes: Based on the audio, obtain the audio features of several audio segments; Obtain text features and image features corresponding to at least two target sounds; The process of obtaining image features corresponding to at least two target sounds includes: obtaining several images corresponding to the target sounds using a stable diffusion model; and obtaining image features corresponding to the several images using an image encoder. An audio-text matrix is obtained based on the audio features and the text features; An audio-image matrix is obtained based on the audio features and the image features; The audio-text matrix and the audio-image matrix are deeply fused to obtain a correlation matrix; the audio-text matrix and the audio-image matrix are then input into a cross-attention mechanism for deep correlation fusion to obtain the correlation matrix. Based on the correlation matrix, the target sound corresponding to each audio segment is obtained; the correlation matrix of shape (M, N) is classified and mapped in the N dimension to obtain the probability distribution corresponding to each target sound; the category with the largest probability value in each audio segment is taken as the corresponding target sound to obtain the target sound detection results in M audio segments.
2. The sound category detection method according to claim 1, characterized in that, The process of obtaining audio features from several audio segments includes: The audio is divided into several audio segments using the VAD algorithm; The audio features of the several audio segments are obtained using an audio encoder.
3. The sound category detection method according to claim 1, characterized in that, The process of obtaining an audio-text matrix based on the audio features and the text features includes: performing matrix multiplication on the audio features and the text features, querying text information using the audio, and obtaining the audio-text matrix.
4. The sound category detection method according to claim 1, characterized in that, The process of obtaining an audio-image matrix based on the audio features and the image features includes: performing matrix multiplication on the audio features and the image features, querying image information using the audio, and obtaining the audio-image matrix.
5. The sound category detection method according to claim 1, characterized in that, The process of deep fusing the audio-text matrix and the audio-image matrix to obtain the correlation matrix includes: inputting the audio-text matrix and the audio-image matrix into a cross-attention mechanism for deep correlation fusing to obtain the correlation matrix.
6. The sound category detection method according to claim 1, characterized in that, The process of obtaining the corresponding target sound for each audio segment based on the correlation matrix includes: The correlation matrix is classified and mapped to obtain the probability distribution corresponding to each target sound; The category with the highest probability value among the audio segments is taken as the corresponding target sound.
7. A sound category detection system, characterized in that, The sound category detection system includes: The audio feature acquisition module is used to acquire the audio features of several audio segments based on the audio. The text feature and image feature acquisition module is used to acquire text features and image features corresponding to at least two target sounds; wherein, the process of acquiring image features corresponding to at least two target sounds includes: acquiring several images corresponding to the target sounds using a stable diffusion model; and acquiring image features corresponding to the several images using an image encoder; An audio-text matrix acquisition module is used to acquire an audio-text matrix based on the audio features and the text features; an audio-image matrix acquisition module is used to acquire an audio-image matrix based on the audio features and the image features; a correlation matrix acquisition module is used to perform deep fusion of the audio-text matrix and the audio-image matrix to obtain a correlation matrix; and input the audio-text matrix and the audio-image matrix into a cross-attention mechanism for deep correlation fusion to obtain the correlation matrix. The target sound acquisition module is used to acquire the target sound corresponding to each audio segment based on the correlation matrix; perform classification mapping on the N-dimensional aspect of the (M, N) correlation matrix to acquire the probability distribution corresponding to each target sound; and take the category with the highest probability value in each audio segment as the corresponding target sound to obtain the target sound detection results in M audio segments.
8. An electronic device, characterized in that, The electronic device includes: A memory on which computer programs are stored; A processor, communicatively connected to the memory, is used to execute the computer program to implement the sound category detection method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by an electronic device, the program implements the sound category detection method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Sound detection method and related equipment
CN115862682A
Sound event detection method and system, storage medium and electronic equipment
CN117912495A