Sound category detection method and system, medium and electronic equipment
By constructing the audio text matrix and the audio image matrix and deeply fusion, the sound category detection problem in the case of mixed sounds is solved, and the target sound detection of audio clips is realized, which improves the accuracy and efficiency of detection.
Patent Information
- Application Number
- CN202510606419.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-05-12
AI Technical Summary
The existing sound detection methods mainly identify a single sound, and it is difficult to effectively detect different sound categories in the audio when multiple sounds are mixed.
By obtaining the audio characteristics of the audio clip and the image characteristics and text characteristics of the target sound, an audio text matrix and audio image matrix are constructed, and deeply integrated, the correlation matrix is obtained using the cross attention mechanism, and finally the target sound is determined through classification mapping.
The target sound detection of audio clips in the case of mixed sounds is achieved, and the accuracy and efficiency of sound category detection are improved.
Smart Images

Figure CN120544601A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of artificial intelligence technology and relates to a sound category detection method, system, medium and electronic device. Background Art
[0002] In recent years, speech recognition technology has been widely used in scenarios such as sound monitoring, emotion recognition, identity verification, and fault sound detection. Sound detection is a crucial pre-processing step for speech recognition and significantly impacts its performance. However, current sound detection primarily focuses on detecting and recognizing a single sound. Therefore, developing a sound detection method that can detect different sound categories in audio even when multiple sounds are mixed is a pressing issue for those skilled in the art. Summary of the Invention
[0003] The purpose of this application is to provide a sound category detection method, system, medium and electronic device for detecting a target sound in an audio clip.
[0004] In a first aspect, the present application provides a sound category detection method, characterized in that the sound category detection method includes: obtaining audio features of several audio clips based on audio; obtaining text features and image features corresponding to at least two target sounds; obtaining an audio-text matrix based on the audio features and the text features; obtaining an audio-image matrix based on the audio features and the image features; deeply fusing the audio-text matrix and the audio-image matrix to obtain a correlation matrix; and obtaining the target sound corresponding to each of the audio clips based on the correlation matrix.
[0005] In an implementation of the first aspect, the process of obtaining audio features of several audio segments based on audio includes: using a VAD algorithm to cut the audio into several audio segments; and using an audio encoder to obtain audio features of the several audio segments.
[0006] In an implementation of the first aspect, the process of obtaining image features corresponding to at least two target sounds includes: obtaining several images corresponding to the target sounds using a stable diffusion model; and obtaining image features corresponding to the several images using an image encoder.
[0007] In an implementation of the first aspect, the process of obtaining an audio-text matrix based on the audio features and the text features includes: performing matrix multiplication on the audio features and the text features, using audio to query text information, and obtaining the audio-text matrix.
[0008] In an implementation of the first aspect, the process of obtaining an audio-image matrix based on the audio features and the image features includes: performing matrix multiplication on the audio features and the image features, using audio to query image information, and obtaining the audio-image matrix.
[0009] In an implementation of the first aspect, the process of deeply fusing the audio text matrix and the audio image matrix to obtain a correlation matrix includes: inputting the audio text matrix and the audio image matrix into a cross-attention mechanism for deep correlation fusion to obtain the correlation matrix.
[0010] In an implementation of the first aspect, the process of obtaining the corresponding target sound of each audio segment based on the correlation matrix includes: performing classification mapping on the correlation matrix to obtain the probability distribution corresponding to each target sound; and taking the category with the largest probability value in each audio segment as the corresponding target sound.
[0011] In a second aspect, the present application provides a sound category detection system, which includes: an audio feature acquisition module for acquiring audio features of several audio clips based on audio; a text feature and image feature acquisition module for acquiring text features and image features corresponding to at least two target sounds; an audio-text matrix acquisition module for acquiring an audio-text matrix based on the audio features and the text features; an audio-image matrix acquisition module for acquiring an audio-image matrix based on the audio features and the image features; a correlation matrix acquisition module for deeply fusing the audio-text matrix and the audio-image matrix to obtain a correlation matrix; and a target sound acquisition module for acquiring the target sound corresponding to each of the audio clips based on the correlation matrix.
[0012] In a third aspect, the present application provides an electronic device, comprising: a memory storing a computer program; and a processor communicatively connected to the memory, for executing the computer program to implement the above-mentioned sound category detection method.
[0013] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the above-mentioned sound category detection method when executed by an electronic device.
[0014] As described above, the sound category detection method described in this application has the following beneficial effects:
[0015] The sound category detection method provided in the embodiments of this application can process audio to obtain audio features of several audio clips, and then use the image and text features of the target sound to obtain an audio-text matrix and an audio-image matrix. After deep fusion of the audio-text matrix and the audio-image matrix, the target sound corresponding to each audio clip can be further obtained. The target sound detection of an audio clip is achieved by using audio-text-image multimodal information. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 Shown is a schematic diagram of an application scenario of the sound category detection method described in an embodiment of the present application.
[0017] Figure 2 Shown is a process diagram of the sound category detection method described in an embodiment of the present application.
[0018] Figure 3 Shown is a schematic diagram of the process of obtaining audio features described in an embodiment of the present application.
[0019] Figure 4 Shown is a schematic diagram of the process of acquiring image features described in an embodiment of the present application.
[0020] Figure 5 Shown is a schematic diagram of the process of acquiring the target sound according to an embodiment of the present application.
[0021] Figure 6 Shown is a structural schematic diagram of the sound category detection system described in an embodiment of the present application.
[0022] Figure 7 Shown is a structural schematic diagram of an electronic device described in an embodiment of the present application.
[0023] Component number description
[0024] 11. Cell Phone
[0025] 12 tablets
[0026] 13 laptops
[0027] 2 Sound Category Detection System
[0028] 21 Audio Feature Acquisition Module
[0029] 22 Text feature and image feature acquisition module
[0030] 23 Audio text matrix acquisition module
[0031] 24 Audio image matrix acquisition module
[0032] 25 Correlation Matrix Acquisition Module
[0033] 26 Target Sound Acquisition Module
[0034] 3 Electronic devices
[0035] 31 Memory
[0036] 32 processors
[0037] Steps S11 to S16
[0038] Steps S21-S22
[0039] Steps S31-S32
[0040] Steps S41-S42 DETAILED DESCRIPTION
[0041] The following describes the embodiments of the present application through specific examples. Those skilled in the art can easily understand the other advantages and effects of the present application from the content disclosed in this specification. The present application can also be implemented or applied through other different specific embodiments. The details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that the following embodiments and features in the embodiments can be combined with each other unless they conflict.
[0042] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present application. Therefore, the illustrations only show components related to the present application and are not drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component can be changed at will, and the component layout type may also be more complicated.
[0043] In recent years, speech recognition technology has been widely used in scenarios such as sound monitoring, emotion recognition, identity verification, and fault sound detection. Sound detection is a crucial pre-processing step for speech recognition and significantly impacts its performance. However, current sound detection primarily focuses on detecting and recognizing a single sound. Therefore, developing a sound detection method that can detect different sound categories in audio even when multiple sounds are mixed is a pressing issue for those skilled in the art.
[0044] At least to address the above problems, the following embodiments of the present application provide a sound category detection method, which can be applied to Figure 1 Electronic devices shown.
[0045] The electronic devices described in this application may include mobile phones 11, tablet computers 12, laptop computers 13, wearable devices, vehicle-mounted devices, augmented reality (AR) / virtual reality (VR) devices, ultra-mobile personal computers (UMPC), netbooks, personal digital assistants (PDA), etc. with wireless charging function. The embodiments of this application do not impose any restrictions on the specific types of electronic devices.
[0046] For example, the electronic device may be a station (STAION, ST) in a WLAN with a wireless charging function, a cellular phone, a cordless phone, a Session Initiation Protocol (SIP) phone, a Wireless Local Loop (WLL) station, a Personal Digital Assistant (PDA) device, a handheld device with a wireless charging function, a computing device or other processing device, a computer, a laptop computer, a handheld communication device, a handheld computing device, and / or other devices for communicating on a wireless system and a next-generation communication system, such as a mobile terminal in a 5G network, a mobile terminal in a future-evolved Public Land Mobile Network (PLMN), or a mobile terminal in a future-evolved Non-terrestrial Network (NTN).
[0047] For example, the electronic device can communicate with a network and other devices via wireless communication. The wireless communication can use any communication standard or protocol, including but not limited to Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), BT, GNSS, WLAN, NFC, FM, and / or IR technology. The GNSS can include Global Positioning System (GPS), Global Navigation Satellite System (GLONASS), BeiDou navigation Satellite System (BDS), Quasi-Zenith Satellite System (QZSS) and / or Satellite Based Augmentation Systems (SBAS).
[0048] The technical solutions in the embodiments of the present application will be described in detail below with reference to the accompanying drawings in the embodiments of the present application.
[0049] Figure 2 The following is a process diagram of a sound category detection method according to an embodiment of the present application. Figure 2 As shown, the sound category detection method includes:
[0050] S11: Acquire audio features of several audio clips based on the audio.
[0051] S12: Acquire text features and image features corresponding to at least two target sounds.
[0052] S13: Acquire an audio-text matrix according to the audio features and the text features.
[0053] S14: Acquire an audio-image matrix according to the audio features and the image features.
[0054] S15, deeply fusing the audio text matrix and the audio image matrix to obtain a correlation matrix.
[0055] S16: Obtain target sounds corresponding to the audio segments based on the correlation matrix.
[0056] As can be seen from the above description, the sound category detection method provided in the embodiments of this application can process audio to obtain the audio features of several audio clips, and use the image features and text features of the target sound to obtain an audio-text matrix and an audio-image matrix. After deep fusion of the audio-text matrix and the audio-image matrix, the target sound corresponding to each audio clip can be further obtained. The target sound detection of the audio clip is achieved through audio-text-image multimodal information.
[0057] Figure 3 The following is a schematic diagram of the process of obtaining audio features in one embodiment of the present application. Figure 3 As shown in FIG, the process of obtaining audio features of several audio clips according to audio includes:
[0058] S21. Utilize a VAD (Voice Activity Detection) algorithm to extract the audio into several audio segments. The VAD algorithm can perform voice annotation processing on the audio, locate the starting and ending points of speech, and extract the signal features required for speech emotion recognition. The VAD algorithm can identify and eliminate long periods of silence in the audio, thereby separating valid speech signals from useless speech signals or noise signals.
[0059] Exemplarily, a VAD sound endpoint detection algorithm is used to cut a long audio into M audio segments.
[0060] S22: Obtain audio features of the plurality of audio segments using an audio encoder. The audio encoder is the audio encoder in Audio CLIP. Audio CLIP combines deep learning with natural language processing to achieve cross-modal audio understanding. The audio encoder in Audio CLIP can extract audio features from the audio segments to obtain the audio features of the audio segments.
[0061] Exemplarily, the audio encoder in Audio CLIP is used to perform feature extraction on M audio clips to obtain M audio features.
[0062] In one embodiment of the present application, the process of obtaining text features corresponding to at least two target sounds includes: setting the target sound to be detected, and using a text encoder to obtain the text features corresponding to the target sound. The target sound may be, for example, "bird calls," "whistles," or "human voices." The text encoder is the text encoder in Audio CLIP. The text encoder in Audio CLIP can extract text features from the sound and obtain the text features corresponding to the sound.
[0063] Exemplarily, N target sounds are selected, and features of the N target sounds are extracted using the text encoder in Audio CLIP to obtain N text features.
[0064] Figure 4 The following is a schematic diagram of the process of obtaining image features in one embodiment of the present application. Figure 4 As shown, the process of obtaining image features corresponding to at least two target sounds includes:
[0065] S31, using a stable diffusion model to obtain a plurality of images corresponding to the target sound. The stable diffusion model is a generative model that can generate images of related content based on the target sound.
[0066] Exemplarily, a stable diffusion model is used to generate images related to the content of the target sound, and N images are obtained.
[0067] S32: Obtain image features corresponding to the plurality of images using an image encoder. The image encoder is an image encoder in Audio CLIP. The image encoder in Audio CLIP can generate images based on the content of the target sound and obtain corresponding image features.
[0068] Exemplarily, the image encoder in Audio CLIP is used to perform feature extraction on N images to obtain N image features.
[0069] In one embodiment of the present application, the process of obtaining an audio-text matrix based on the audio features and the text features includes: performing matrix multiplication on the audio features and the text features, using the audio to query text information, and obtaining the audio-text matrix.
[0070] For example, matrix multiplication is performed on M audio features and N text features to obtain an audio-text matrix. The shape of the audio-text matrix is (M, N). Taking audio as the primary modality of interest, the audio-text correlation matrix is obtained by querying the text information using audio.
[0071] In one embodiment of the present application, the process of obtaining an audio-image matrix based on the audio features and the image features includes: performing matrix multiplication on the audio features and the image features, using audio to query image information, and obtaining the audio-image matrix.
[0072] For example, matrix multiplication is performed on M audio features and N image features to obtain an audio-image matrix. The shape of the audio-text matrix is (M, N). Taking audio as the primary modality of interest, audio is used to query image information to obtain an audio-image correlation matrix.
[0073] In one embodiment of the present application, the audio text matrix and the audio image matrix are deeply fused to obtain a correlation matrix, which includes the following process: inputting the audio text matrix and the audio image matrix into a cross-attention mechanism for deep correlation fusion to obtain the correlation matrix. The cross-attention mechanism is cross-attention, which is a cross-attention layer between the encoder and the decoder. The process of deep correlation fusion in the cross-attention mechanism is: in the cross-attention layer, the decoder adjusts the attention of the encoder output to obtain encoder information related to the current decoding position. Each position of the decoder generates a query vector Q to calculate the attention weight for all positions of the encoder. All positions of the encoder will generate a set of key vectors K and value vectors V, perform dot product operations using the query vector Q and the key vector K, and obtain the attention weight through the softmax function, multiply the attention weight by the value vector V, and obtain the output of the encoder adjustment by summing the multiplication results.
[0074] For example, the audio text matrix is used as Q, the audio image matrix is used as K and V, and QKV is input into the cross-attention method for deep correlation fusion, and the correlation matrix is output. The shape of the correlation matrix is (M, N).
[0075] Figure 5 The following is a schematic diagram of the process of obtaining the target sound in one embodiment of the present application. Figure 5 As shown, the process of obtaining the target sound corresponding to each of the audio segments based on the correlation matrix includes:
[0076] S41, classify and map the correlation matrix to obtain the probability distribution corresponding to each target sound. Classify and map the correlation matrix using a softmax function to output the probability distribution value of each target sound.
[0077] Exemplarily, a softmax function is activated on the correlation matrix in the dimension N to obtain the probability distribution corresponding to each target sound.
[0078] S42: taking the category with the largest probability value in each of the audio segments as the corresponding target sound.
[0079] Exemplarily, the index with the largest probability value in each audio segment is taken as the target sound, and each audio segment corresponds to one target sound, and finally the detection results of M audio segments can be obtained.
[0080] The following is an example of an embodiment to illustrate the sound category detection method of the present application. The specific implementation process is as follows:
[0081] Step 1: Input a long audio file, cut it, and obtain audio features. Use the VAD sound endpoint detection method to cut the long audio file into M audio segments, and use the audio encoder in Audio CLIP to obtain M audio features.
[0082] Step 2: Set N target sounds to be detected and use the text encoder in Audio CLIP to obtain N text features.
[0083] Step 3: Generate images based on the target sound and obtain the corresponding image features. Generate N images corresponding to the target sound content using the stable diffusion model, and use the image encoder in Audio CLIP to obtain N image features.
[0084] Step 4: Obtain an audio-text matrix based on audio and text features. Perform matrix multiplication on the M audio features and the N text features, taking audio as the primary modality of interest and using audio to query text information to obtain an audio-text matrix with a shape of (M, N).
[0085] Step 5: Obtain an audio-image matrix based on audio and image features. Perform matrix multiplication on the M audio features and the N image features, taking audio as the primary modality of interest and using audio to query image information to obtain an audio-image matrix with a shape of (M, N).
[0086] Step 6: Deeply fuse the audio text matrix and the audio image matrix to obtain a correlation matrix. Use the audio text matrix as Q, the audio image matrix as K and V, and input QKV into the cross-attention method for deep correlation fusion, outputting a correlation matrix with a shape of (M, N).
[0087] Step 7: Obtain the target sound corresponding to each audio segment based on the correlation matrix. Apply a softmax function to the correlation matrix over N dimensions to obtain the probability distribution corresponding to each target sound. The index with the highest probability value in each audio segment is used as the target sound. Each audio segment corresponds to one target sound, ultimately yielding detection results for M audio segments.
[0088] In summary, the sound category detection method described in the embodiments of this application can process audio to obtain audio features of several audio clips, and then use the image features and text features of the target sound to obtain an audio-text matrix and an audio-image matrix. After deep fusion of the audio-text matrix and the audio-image matrix, the target sound corresponding to each audio clip can be further obtained. Target sound detection in audio clips is achieved through audio-text-image multimodal information.
[0089] The protection scope of the sound category detection method described in the embodiment of the present application is not limited to the execution order of the steps listed in this embodiment. All solutions implemented by adding, subtracting, or replacing steps in the existing technology based on the principles of the present application are included in the protection scope of the present application.
[0090] An embodiment of the present application also provides a sound category detection system, which can implement the sound category detection method described in the present application. However, the implementation device of the sound category detection method described in the present application includes but is not limited to the structure of the sound category detection system listed in this embodiment. Any structural deformation and replacement of the existing technology made according to the principles of the present application are included in the scope of protection of the present application.
[0091] Figure 6 Shown is a structural diagram of a sound category detection system in one embodiment of the present application. Figure 6 As shown, the sound category detection system 2 includes: an audio feature acquisition module 21, a text feature and image feature acquisition module 22, an audio text matrix acquisition module 23, an audio image matrix acquisition module 24, a correlation matrix acquisition module 25, and a target sound acquisition module 26.
[0092] Among them, the audio feature acquisition module 21 is used to obtain the audio features of several audio clips based on the audio. The text feature and image feature acquisition module 22 is used to obtain the text features and image features corresponding to at least two target sounds. The audio-text matrix acquisition module 23 is used to obtain the audio-text matrix based on the audio features and the text features. The audio-image matrix acquisition module 24 is used to obtain the audio-image matrix based on the audio features and the image features. The correlation matrix acquisition module 25 is used to deeply fuse the audio-text matrix and the audio-image matrix to obtain a correlation matrix. The target sound acquisition module 26 is used to obtain the target sound corresponding to each of the audio clips based on the correlation matrix.
[0093] It should be noted that Figure 6 The modules in the sound category detection system 2 are Figure 2 The steps in the sound category detection method correspond to each other and are not described in detail here.
[0094] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices or methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of modules / units is only a logical function division. There may be other division methods in actual implementation. For example, multiple modules or units can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules or units, which can be electrical, mechanical or other forms.
[0095] The modules / units described as separate components may or may not be physically separate, and the components displayed as modules / units may or may not be physical modules, that is, they may be located in one place or distributed across multiple network elements. Some or all of the modules / units may be selected according to actual needs to achieve the purpose of the embodiments of the present application. For example, the functional modules / units in the various embodiments of the present application may be integrated into a processing module, or each module / unit may exist physically separately, or two or more modules / units may be integrated into a single module / unit.
[0096] Those skilled in the art should further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0097] The present application also provides a computer-readable storage medium having a computer program stored thereon, which implements the sound category detection method provided by the present application embodiment when the computer program is executed by the processor. A person skilled in the art will understand that all or part of the steps in the method for implementing the above embodiment can be completed by instructing the processor through a program, and the program can be stored in a computer-readable storage medium, and the storage medium is a non-transitory medium, such as a random access memory, a read-only memory, a flash memory, a hard disk, a solid-state drive, a magnetic tape, a floppy disk, an optical disc, and any combination thereof. The above storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a digital video disc (DVD)), or a semiconductor medium (e.g., a solid-state drive (SSD)).
[0098] An embodiment of the present application may also provide an electronic device. Figure 7 The diagram shows the structure of the electronic device 3 in one embodiment of the present application. Figure 7 As shown, in this embodiment, the electronic device 3 includes a memory 31 and a processor 32.
[0099] The memory 31 is used to store computer programs. In some possible implementations, the memory 31 may include various media capable of storing program codes, such as ROM, RAM, a magnetic disk, a USB flash drive, a memory card, or an optical disk.
[0100] In the embodiment of the present application, the memory 31 may include a computer system readable medium in the form of a volatile memory, such as RAM and / or cache memory. The electronic device 3 may further include other removable / non-removable, volatile / non-volatile computer system storage media. The memory 31 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of each embodiment of the present application.
[0101] The processor 32 is connected to the memory 31 and is configured to execute the computer program stored in the memory 31 so as to enable the electronic device 3 to perform the sound category detection method.
[0102] Exemplarily, the processor 32 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc. In other embodiments, the processor 32 may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0103] The descriptions of the processes or structures corresponding to the above figures have different emphases. For parts that are not described in detail in a certain process or structure, please refer to the relevant descriptions of other processes or structures.
[0104] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Anyone skilled in the art may modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or alterations made by one of ordinary skill in the art without departing from the spirit and technical concepts disclosed in this application shall be covered by the claims of this application.
Claims
1. A method for detecting a sound category, characterized in that: The sound category detection method comprises: Obtain audio features of several audio clips based on the audio; Acquiring text features and image features corresponding to at least two target sounds; The process of obtaining image features corresponding to at least two target sounds includes: obtaining a plurality of images corresponding to the target sounds using a stable diffusion model; obtaining image features corresponding to the plurality of images using an image encoder; Acquire an audio-text matrix according to the audio features and the text features; Acquire an audio image matrix according to the audio features and the image features; Deeply fusing the audio text matrix and the audio image matrix to obtain a correlation matrix; inputting the audio text matrix and the audio image matrix into a cross attention mechanism to perform deep correlation fusion to obtain the correlation matrix; Based on the correlation matrix, the target sound corresponding to each of the audio clips is obtained; the correlation matrix with a shape of (M, N) is classified and mapped in the dimension of N to obtain the probability distribution corresponding to each of the target sounds; the category with the largest probability value in each of the audio clips is taken as the corresponding target sound to obtain the target sound detection results in the M audio clips.
2. The sound category detection method according to claim 1, characterized in that: The process of obtaining audio features of several audio clips based on audio includes: Using a VAD algorithm to cut the audio into a plurality of audio segments; The audio features of the plurality of audio segments are obtained using an audio encoder.
3. The sound category detection method according to claim 1, wherein: The process of obtaining an audio-text matrix according to the audio features and the text features includes: performing matrix multiplication on the audio features and the text features, using audio to query text information, and obtaining the audio-text matrix.
4. The sound category detection method according to claim 1, wherein: The process of obtaining an audio-image matrix according to the audio features and the image features includes: performing matrix multiplication on the audio features and the image features, querying image information using audio, and obtaining the audio-image matrix.
5. The sound category detection method according to claim 1, wherein: The process of deeply fusing the audio text matrix and the audio image matrix to obtain a correlation matrix includes: inputting the audio text matrix and the audio image matrix into a cross-attention mechanism for deep correlation fusion to obtain the correlation matrix.
6. The sound category detection method according to claim 1, characterized in that: The process of obtaining the target sound corresponding to each of the audio segments based on the correlation matrix includes: Performing classification mapping on the correlation matrix to obtain a probability distribution corresponding to each of the target sounds; The category with the largest probability value in each of the audio segments is taken as the corresponding target sound.
7. A sound category detection system, characterized in that: The sound category detection system comprises: An audio feature acquisition module is used to acquire audio features of several audio clips based on the audio; A text feature and image feature acquisition module is configured to acquire text features and image features corresponding to at least two target sounds; wherein the process of acquiring image features corresponding to at least two target sounds comprises: acquiring a plurality of images corresponding to the target sounds using a stable diffusion model; and acquiring image features corresponding to the plurality of images using an image encoder; An audio-text matrix acquisition module is configured to acquire an audio-text matrix based on the audio features and the text features; an audio-image matrix acquisition module is configured to acquire an audio-image matrix based on the audio features and the image features; a correlation matrix acquisition module is configured to perform deep fusion of the audio-text matrix and the audio-image matrix to acquire a correlation matrix; and the audio-text matrix and the audio-image matrix are input into a cross-attention mechanism for deep correlation fusion to acquire the correlation matrix. The target sound acquisition module is used to obtain the target sound corresponding to each of the audio clips based on the correlation matrix; classify and map the correlation matrix with a shape of (M, N) in the dimension of N to obtain the probability distribution corresponding to each of the target sounds; and take the category with the largest probability value in each of the audio clips as the corresponding target sound to obtain the target sound detection results in the M audio clips.
8. An electronic device, characterized in that: The electronic device comprises: a memory having a computer program stored thereon; A processor is communicatively connected to the memory, and is configured to execute the computer program to implement the sound category detection method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by an electronic device, the sound category detection method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Sound detection method and related equipment
CN115862682A
Sound event detection method and system, storage medium and electronic equipment
CN117912495A
Sound event classification method, electronic equipment and storage medium
CN118351839A
Voice output method, device and equipment and storage medium thereof
CN119181355A
Speech recognition method and device, electronic equipment and medium
CN119811394A