An audio recognition method, an electronic device, and a readable storage medium

By clustering audio features and selecting standard audio features, an audio feature library is constructed, solving the problems of storage and low recognition efficiency in online song recognition systems, and achieving efficient audio recognition.

CN115203342BActive Publication Date: 2026-07-24TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
Filing Date
2022-07-28
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

In existing technologies, online song recognition systems suffer from severe storage problems due to their large song libraries, and their recognition efficiency is low.

Method used

By clustering preset audio features to form multiple audio feature groups, and selecting standard audio features to form an audio feature library, the amount of data is reduced, and matching and recognition are performed using standard audio features.

Benefits of technology

While ensuring the reliability of audio recognition, it significantly reduces the amount of data in the audio feature library, avoids storage problems, and improves recognition efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115203342B_ABST
    Figure CN115203342B_ABST
Patent Text Reader

Abstract

The application discloses an audio recognition method, an electronic device and a readable storage medium. The method comprises the following steps: acquiring preset audios, and extracting preset audio features corresponding to the preset audios respectively; performing clustering processing on the preset audio features to obtain a plurality of audio feature groups; selecting standard audio features in each audio feature group respectively, and constructing an audio feature library by using the standard audio features; acquiring a to-be-recognized audio sent by a terminal; wherein the to-be-recognized audio is acquired by a corresponding sound collecting device of the terminal; extracting to-be-recognized audio features of the to-be-recognized audio; determining a target audio feature most similar to the to-be-recognized audio features from the standard audio features in the audio feature library based on the to-be-recognized audio features; and sending target audio information corresponding to the target audio feature to the terminal. The method can greatly reduce the data amount of the audio feature library while ensuring the reliability of audio recognition through clustering and extraction of standard audio features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing technology, and in particular to an audio recognition method, electronic device, and computer-readable storage medium. Background Technology

[0002] Online song recognition refers to a service where users request a song by humming their own lyrics or capturing audio clips played on other devices, and the service then identifies the corresponding song information. Song recognition requires a massive song library to support the entire service. The number of original songs alone is enormous, and when you add covers of popular songs by various artists, the amount of data in the song library becomes incalculable. Such a massive song library can lead to serious storage problems. Summary of the Invention

[0003] In view of this, the purpose of this application is to provide an audio recognition method, electronic device and computer-readable storage medium that can significantly reduce the amount of data in the audio feature library while ensuring the reliability of audio recognition.

[0004] To address the aforementioned technical problems, in a first aspect, this application provides an audio recognition method, comprising:

[0005] Obtain preset audio and extract the preset audio features corresponding to each preset audio;

[0006] The preset audio features are clustered to obtain multiple audio feature groups;

[0007] Standard audio features are selected from each of the aforementioned audio feature groups, and an audio feature library is constructed using the standard audio features.

[0008] Acquire the audio to be identified sent by the terminal; wherein the audio to be identified is acquired by the audio receiving device corresponding to the terminal;

[0009] Extract the audio features to be identified from the audio to be identified;

[0010] Based on the audio feature to be identified, a target audio feature that is most similar to the audio feature to be identified is determined from each of the standard audio features in the audio feature library;

[0011] The target audio information corresponding to the target audio feature is sent to the terminal.

[0012] Optionally, the clustering process of the preset audio features to obtain multiple audio feature groups includes:

[0013] The preset audio features are subjected to feature similarity clustering to obtain multiple initial audio feature groups;

[0014] For each preset audio feature within each initial audio feature group, clustering based on beat count is performed to obtain the audio feature group corresponding to the initial audio feature group.

[0015] Optionally, the step of performing clustering processing based on the number of beats on each preset audio feature within each initial audio feature group to obtain the audio feature group corresponding to the initial audio feature group includes:

[0016] Determine the number of beats corresponding to each preset audio feature within each initial audio feature group;

[0017] Based on a preset range of beat counts, the preset audio features are grouped using the beat counts to obtain an intermediate audio feature group;

[0018] Determine the preset audio name for each of the preset audio features;

[0019] Using the preset audio name, each of the intermediate audio feature groups is filtered according to the preset name filtering rules to obtain the audio feature group.

[0020] Optionally, the preset name filtering rule is a rule that filters out preset audio features corresponding to the target audio name;

[0021] The target audio name is a name that can represent that the preset audio is obtained by concatenating at least two audio files;

[0022] And / or,

[0023] The target audio name is a name that can characterize the features of the preset audio that cannot be accurately represented by the preset audio features.

[0024] Optionally, selecting standard audio features within each of the audio feature groups includes:

[0025] The standard audio feature is selected based on the audio attribute information corresponding to each preset audio feature in each of the audio feature groups.

[0026] Optionally, constructing an audio feature library using the standard audio features includes:

[0027] Generate the correspondence between the standard audio features and the corresponding standard audio information;

[0028] The audio feature library is constructed using the aforementioned correspondence and the aforementioned standard audio features.

[0029] Optionally, determining the target audio feature most similar to the audio feature to be identified from among the standard audio features in the audio feature library based on the audio feature to be identified includes:

[0030] Calculate the similarity data between the audio feature to be identified and each of the standard audio features;

[0031] The maximum similarity data is determined, and the standard audio feature corresponding to the maximum similarity data is determined as the target audio feature;

[0032] Accordingly, sending the target audio information corresponding to the target audio feature to the terminal includes:

[0033] The target audio information corresponding to the target audio feature is obtained using the correspondence, and the target audio information is sent to the terminal.

[0034] Optionally, the audio to be identified is the audio obtained by the recording device from the user's humming voice;

[0035] or,

[0036] The audio to be identified is the audio obtained by the radio device after collecting audio played by other electronic devices.

[0037] Secondly, this application also provides an electronic device, including a memory and a processor, wherein:

[0038] The memory is used to store computer programs;

[0039] The processor is used to execute the computer program to implement the above-described audio recognition method.

[0040] Thirdly, this application also provides a computer-readable storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the above-described audio recognition method.

[0041] The audio recognition method provided in this application involves: acquiring preset audio and extracting preset audio features corresponding to each preset audio; clustering the preset audio features to obtain multiple audio feature groups; selecting standard audio features within each audio feature group and constructing an audio feature library using the standard audio features; acquiring the audio to be recognized sent by the terminal; wherein the audio to be recognized is acquired by the corresponding audio receiving device of the terminal; extracting the audio features to be recognized from the audio to be recognized; based on the audio features to be recognized, determining the target audio feature most similar to the audio features to be recognized among the standard audio features in the audio feature library; and sending the target audio information corresponding to the target audio feature to the terminal.

[0042] As can be seen, this method does not store all preset audio features in a song library, but rather clusters them into multiple audio feature groups. The preset audio features within each group have high similarity, and one is selected as the representative of that group, i.e., the standard audio feature, forming the audio feature library. During the recognition process, the terminal uses a recording device to acquire the audio to be recognized and extracts the corresponding audio features. The audio features to be recognized are matched with the standard audio features in the audio feature library to obtain the most similar target audio feature, which is considered to correspond to the same preset audio. Therefore, the target audio information corresponding to the target audio feature is fed back to the terminal. Through clustering and extraction of standard audio features, the reliability of audio recognition is ensured while significantly reducing the data size of the audio feature library, avoiding serious storage problems.

[0043] In addition, this application also provides an electronic device and a computer-readable storage medium that have the same beneficial effects as described above. Attached Figure Description

[0044] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0045] Figure 1 A schematic diagram of the hardware framework applicable to an audio recognition method provided in an embodiment of this application;

[0046] Figure 2 A schematic diagram of the hardware framework applicable to another audio recognition method provided in this application embodiment;

[0047] Figure 3 A flowchart illustrating an audio recognition method provided in an embodiment of this application;

[0048] Figure 4 A classification effect diagram provided for an embodiment of this application;

[0049] Figure 5 Another classification effect diagram provided for an embodiment of this application;

[0050] Figure 6 This is a schematic diagram illustrating a specific audio feature library construction process provided in an embodiment of this application. Detailed Implementation

[0051] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0052] To facilitate understanding, the hardware framework used in the audio recognition method provided in this application embodiment will be introduced first. Please refer to... Figure 1 , Figure 1 This is a schematic diagram of the hardware framework applicable to an audio recognition method provided in an embodiment of this application. The electronic device 100 may include a processor 101 and a memory 102, and may further include one or more of a multimedia component 103, an information input / output (I / O) interface 104, and a communication component 105.

[0053] The processor 101 controls the overall operation of the electronic device 100 to complete all or part of the steps in the audio recognition method. The memory 102 stores various types of data to support the operation of the electronic device 100. This data may include, for example, instructions for any application or method operating on the electronic device 100, as well as application-related data. The memory 102 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as one or more of Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. In this embodiment, the memory 102 stores at least programs and / or data for implementing the following functions:

[0054] Obtain preset audio and extract the preset audio features corresponding to each preset audio;

[0055] Clustering is performed on the preset audio features to obtain multiple audio feature groups;

[0056] Standard audio features are selected from each audio feature group, and an audio feature library is constructed using these standard audio features.

[0057] The audio to be recognized is sent by the terminal; wherein the audio to be recognized is acquired by the corresponding audio receiving device of the terminal.

[0058] Extract the audio features to be identified from the audio to be identified;

[0059] Based on the audio features to be identified, the target audio features most similar to the audio features to be identified are determined among various standard audio features;

[0060] Send the target audio information corresponding to the target audio features to the terminal.

[0061] Multimedia component 103 may include a screen and an audio component. The screen may be, for example, a touchscreen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in memory 102 or transmitted via communication component 105. The audio component also includes at least one speaker for outputting audio signals. I / O interface 104 provides an interface between processor 101 and other interface modules, such as a keyboard, mouse, buttons, etc. These buttons may be virtual or physical buttons. Communication component 105 is used for wired or wireless communication between electronic device 100 and other devices. Wireless communication may include Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, or 4G, or a combination of these. Therefore, the corresponding communication component 105 may include a Wi-Fi component, a Bluetooth component, or an NFC component.

[0062] Electronic device 100 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform an audio recognition method.

[0063] certainly, Figure 1The structure of the electronic device 100 shown does not constitute a limitation on the electronic device in the embodiments of this application. In practical applications, the electronic device 100 may include more than Figure 1 More or fewer components as shown, or combinations of certain components.

[0064] It is understood that the number of electronic devices is not limited in the embodiments of this application; multiple electronic devices can work together to complete the audio recognition method. In one possible implementation, please refer to... Figure 2 , Figure 2 This is a schematic diagram of the hardware framework applicable to another audio recognition method provided in this application embodiment. Figure 2 It can be seen that the hardware framework may include: a first electronic device 11 and a second electronic device 12, which are connected to each other via a network 13.

[0065] In this embodiment, the hardware structure of the first electronic device 11 and the second electronic device 12 can be referred to Figure 1 The electronic device 100 refers to two electronic devices 100 that interact with each other. Furthermore, this embodiment does not limit the form of the network 13; that is, the network 13 can be a wireless network (such as Wi-Fi, Bluetooth, etc.) or a wired network.

[0066] In this embodiment, the first electronic device 11 and the second electronic device 12 can be the same type of electronic device, such as both being servers; alternatively, they can be different types of electronic devices, for example, the first electronic device 11 can be a smartphone or other smart terminal, and the second electronic device 12 can be a server. In one possible implementation, both the first electronic device 11 and the second electronic device 12 are servers. The first electronic device 11 generates an audio feature library using preset audio and sends it to the second electronic device 12. The second electronic device 12 acquires the audio to be identified and uses the audio feature library to obtain the target audio features.

[0067] For details, please refer to Figure 3 , Figure 3 This is a flowchart illustrating an audio recognition method provided in an embodiment of this application. The method in this embodiment includes:

[0068] S101: Obtain preset audio and extract the preset audio features corresponding to each preset audio.

[0069] In this context, preset audio refers to the audio used as the audio recognition result. Its specific content and quantity are not limited; for example, it can be the original song, a cover song, or instrumental music. Each preset audio has different content, which can be represented using corresponding preset audio features. The form and extraction method of the preset audio features are not limited; they only need to be able to represent the content, form, and other characteristics of the preset audio. In online audio recognition applications, to improve recognition speed, preset audio features corresponding to each preset audio are pre-extracted and compiled into an audio feature library, so that the audio feature library can be directly used for subsequent audio recognition.

[0070] S102: Cluster the preset audio features to obtain multiple audio feature groups.

[0071] S103: Select standard audio features within each audio feature group and use these standard audio features to construct an audio feature library.

[0072] For ease of explanation, the two steps above will be combined.

[0073] Because there are a large number of preset audio files, and many of them are similar (for example, an original song may have multiple very similar cover versions, which can be considered similar to the original), there may also be cover versions with significant alterations and less similarity. These cover versions may not be considered similar to the original. Directly using all preset audio features to construct an audio feature library would result in an excessively large dataset, leading to severe storage problems.

[0074] To address this issue, this application does not directly utilize all preset audio features to construct an audio feature library. Specifically, through clustering, multiple audio feature groups can be obtained. Clustering treats similar preset audio features as belonging to the same group, and the preset audio corresponding to similar preset audio features can also be considered similar audio. In applications involving fuzzy audio recognition, these similar audios can be identified as the same audio. This audio can be any one of the preset audios corresponding to the audio feature group, and can be called the standard audio. The feature corresponding to the standard audio is the standard audio feature. Therefore, when generating the audio feature library, only the standard audio features can be used. Each standard audio feature represents all preset audios corresponding to an audio feature group. Because these preset audios are very similar (usually similar covers of the same song, or similar performances of the same piece of music), regardless of which audio the subsequent terminal wants to identify, the extracted feature is most similar to the standard audio features, achieving accurate identification. Please refer to [reference needed]. Figure 4 , Figure 4This is a classification effect diagram provided in an embodiment of this application. The preset audio features corresponding to songs B and C, and their related preset audio features, are respectively divided into two audio feature groups. Furthermore, the preset audio features corresponding to the three different versions of song A (i.e., A-1, A-2, and A-3) are respectively divided into three different audio feature groups.

[0075] It is understandable that the number of audio feature groups and the content of each feature group can vary depending on the different clustering processes. In one implementation, a large portion of the audio with similar preset audio features consists of original and cover versions. Some cover versions are simply re-singing or re-performing the original audio by different performers or singers, while others have also been adapted, often with a change in musical style, which is usually reflected in a change in rhythm, i.e., a change in BPM (beats per minute). Although these modified audios may be similar to the original audio in terms of preset audio features, they are actually quite different. To improve recognition accuracy, they can be divided into different audio feature groups. Therefore, when performing clustering, the preset audio features can first be clustered based on feature similarity to obtain multiple initial audio feature groups. The preset audio features within each initial audio feature group may differ significantly in terms of the number of beats. Therefore, clustering based on the number of beats is further performed on each preset audio feature within each initial audio feature group. That is, within each initial audio feature group, grouping is performed again based on BPM to obtain the audio feature group corresponding to the initial audio feature group.

[0076] This embodiment does not limit the specific implementation of BPM-based clustering. In one embodiment, preset audio features with BPM differences greater than a preset threshold can be divided into different audio feature groups. If a preset audio feature has a BPM difference of no more than the preset threshold from all other preset audio features, then its BPM can be determined to be closest to the average BPM of each different audio feature group, and it can be assigned to that closest audio feature group. The average BPM can be the average or weighted average of the BPMs corresponding to all features within the audio feature group.

[0077] In another implementation, different audio styles can be considered to correspond to different BPM ranges. Therefore, clustering can be performed based on a preset beat count range (i.e., BPM range). Specifically, the beat count corresponding to each preset audio feature within each initial audio feature group is determined, and the preset audio features are grouped by beat count based on the preset beat count range to obtain audio feature groups.

[0078] Furthermore, after clustering based on beat count, the resulting data can be called intermediate audio feature groups. Due to certain factors, such as the audio format (e.g., electronic music), the preset audio features corresponding to preset audio may not accurately represent its content, and using them to construct an audio feature library will reduce audio recognition accuracy. Alternatively, some preset audio may be obtained by splicing multiple audio files; users may not want to recognize these types of audio but rather the original audio, thus also reducing recognition accuracy. After obtaining the intermediate audio feature groups, the preset audio names for each preset audio feature are determined. The audio names can reflect, to some extent, whether the preset audio is one of the two types of audio mentioned above. Using the preset audio names, and based on preset name filtering rules, each intermediate audio feature group is filtered to obtain the audio feature group. It should be noted that the filtering object can be all features in the intermediate audio feature group, or it can be some of the features, specifically one or more features.

[0079] Specifically, the preset name filtering rules are rules that filter out preset audio features corresponding to target audio names; that is, rules that filter out preset audio features corresponding to preset audio with target audio names. Specifically, the target audio name is a name that can represent that the preset audio is obtained by concatenating at least two audio files, such as names containing words like plus sign, "medley," or "collection." Alternatively, the target audio name is a name that can represent characteristics of the preset audio that the preset audio features cannot accurately represent, such as names containing words like "electronic music" or "Electronic Ver." Please refer to [reference needed]. Figure 5 , Figure 5 Another classification effect diagram provided in this application embodiment shows that if the preset audio features corresponding to each version of song A were originally divided into an initial audio feature group, then after BPM clustering and name filtering, three audio feature groups are obtained: an audio feature group composed of preset audio features of A-1, A-3, and A-5 with similar BPM values; an audio feature group composed solely of preset audio features of A-2 with similar BPM values; and an audio feature group composed of preset audio features of A-7, A-8, and A-11 with similar BPM values.

[0080] After dividing the audio feature groups, a feature can be selected from each group as a standard audio feature to represent the entire group. It is understood that the standard audio features will inevitably be dissimilar, yet they will cover all the preset audio data. Therefore, the audio feature library formed by these features can provide accurate audio recognition with a relatively small data volume. The selection method for the standard audio feature is not limited; for example, it can be randomly selected. In one implementation, the standard audio feature can be selected based on the audio attribute information corresponding to each preset audio feature within each audio feature group. The specific content of the audio attribute information can be set as needed, and may include one or more attribute data, such as generation date, historical total playback volume, total playback volume in a recent period, performer, etc. Based on the audio attribute information, the most representative one can be selected as the standard audio feature. The most representative could be the earliest generated, the one with the most playbacks, or the one with the most professional or famous performer.

[0081] After obtaining the standard audio information, a correspondence between the standard audio features and the corresponding standard audio information is generated, and the correspondence and standard audio features are used to construct an audio feature library.

[0082] Please refer to Figure 6 , Figure 6 This is a schematic diagram illustrating a specific audio feature library construction process provided in this application embodiment. First, songs (i.e., preset songs) are extracted from the library, and corresponding song features are extracted using a deep learning model as preset audio features. Clustering is performed based on feature similarity to obtain initial audio feature groups. Further, each initial audio feature group is classified based on BPM to obtain multiple intermediate audio feature groups. Then, based on filtering rules, the intermediate audio feature groups are subjected to rule filtering processing to obtain audio feature groups. (Standard) song features, i.e., standard audio features, are extracted from each audio feature group, and these standard audio features are used to construct a song feature library, i.e., the audio feature library.

[0083] S104: Obtain the audio to be recognized sent by the terminal.

[0084] The audio to be identified is acquired by the corresponding audio receiving device of the terminal. This device can be a component of the terminal itself, or it can be a standalone device that communicates with the terminal and transmits audio signals to it. The specific form of the terminal is not limited; for example, it can be a smartphone. Specifically, the audio to be identified can be audio obtained by the audio receiving device capturing the user's humming, or it can be audio obtained by the audio receiving device capturing audio played from other electronic devices. In other words, the audio to be identified can be based on human humming or ambient sound.

[0085] S105: Extract the audio features to be identified from the audio to be identified.

[0086] It should be noted that the method for extracting the audio features to be identified should be the same as the method for extracting the preset audio features mentioned above. This embodiment does not limit the specific extraction method. An audio feature extraction model can be trained and the trained model can be used to extract the audio features to be identified.

[0087] S106: Based on the audio features to be identified, determine the target audio feature that is most similar to the audio features to be identified among the various standard audio features in the audio feature library.

[0088] S107: Send the target audio information corresponding to the target audio features to the terminal.

[0089] By identifying the target audio feature most similar to the audio feature to be identified from among the standard audio features in the audio feature library, the preset audio corresponding to the audio feature to be identified can be accurately determined. Specifically, if an audio feature library is constructed using the correspondence between standard audio features and their corresponding standard audio information, the similarity data between the audio feature to be identified and each standard audio feature can be calculated, and the maximum similarity data can be determined. The standard audio feature corresponding to the maximum similarity data is determined as the target audio feature, and the target audio information corresponding to the target audio feature is obtained using the correspondence and sent to the target terminal.

[0090] The audio recognition method provided in this application does not store all preset audio features in a song library. Instead, it clusters them into multiple audio feature groups. The preset audio features within each group have high similarity, and one is selected as the representative of the audio feature group, i.e., the standard audio feature, to form the audio feature library. During the recognition process, the terminal uses a recording device to acquire the audio to be recognized and extracts the corresponding audio features. The audio features to be recognized are matched with the standard audio features in the audio feature library to obtain the most similar target audio feature. These two features are considered to correspond to the same preset audio, and the target audio information corresponding to the target audio feature is then fed back to the terminal. Through clustering and extraction of standard audio features, the reliability of audio recognition is ensured while significantly reducing the data volume of the audio feature library, avoiding serious storage problems.

[0091] The following describes the computer-readable storage medium provided in the embodiments of this application. The computer-readable storage medium described below can be referred to in correspondence with the audio recognition method described above.

[0092] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described audio recognition method.

[0093] The computer-readable storage medium may include various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0094] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0095] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0096] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0097] Finally, it should be noted that in this document, relationships such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "include," "contain," or any other variations are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus.

[0098] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. An audio recognition method, characterized in that, include: Obtain preset audio and extract the preset audio features corresponding to each preset audio; The preset audio features are subjected to feature similarity clustering to obtain multiple initial audio feature groups; For each preset audio feature within each initial audio feature group, clustering based on beat count is performed to obtain the audio feature group corresponding to the initial audio feature group; Standard audio features are selected from each of the aforementioned audio feature groups, and an audio feature library is constructed using the standard audio features. Acquire the audio to be identified sent by the terminal; wherein the audio to be identified is acquired by the audio receiving device corresponding to the terminal; Extract the audio features to be identified from the audio to be identified; Based on the audio feature to be identified, a target audio feature that is most similar to the audio feature to be identified is determined from each of the standard audio features in the audio feature library; The target audio information corresponding to the target audio feature is sent to the terminal.

2. The audio recognition method according to claim 1, characterized in that, The step of performing clustering based on beat count on each preset audio feature within each initial audio feature group to obtain the audio feature group corresponding to the initial audio feature group includes: Determine the number of beats corresponding to each preset audio feature within each initial audio feature group; Based on a preset range of beat counts, the preset audio features are grouped using the beat counts to obtain an intermediate audio feature group; Determine the preset audio name for each of the preset audio features; Using the preset audio name, each of the intermediate audio feature groups is filtered according to the preset name filtering rules to obtain the audio feature group.

3. The audio recognition method according to claim 2, characterized in that, The preset name filtering rule is a rule that filters out preset audio features corresponding to the target audio name; The target audio name is a name that can represent that the preset audio is obtained by concatenating at least two audio files; And / or, The target audio name is a name that can characterize the features of the preset audio that cannot be accurately represented by the preset audio features.

4. The audio recognition method according to claim 1, characterized in that, The step of selecting standard audio features within each of the audio feature groups includes: The standard audio feature is selected based on the audio attribute information corresponding to each preset audio feature in each of the audio feature groups.

5. The audio recognition method according to claim 1, characterized in that, The construction of the audio feature library using the standard audio features includes: Generate the correspondence between the standard audio features and the corresponding standard audio information; The audio feature library is constructed using the aforementioned correspondence and the aforementioned standard audio features.

6. The audio recognition method according to claim 5, characterized in that, The step of determining the target audio feature most similar to the audio feature to be identified from among the standard audio features in the audio feature library based on the audio feature to be identified includes: Calculate the similarity data between the audio feature to be identified and each of the standard audio features; The maximum similarity data is determined, and the standard audio feature corresponding to the maximum similarity data is determined as the target audio feature; Accordingly, sending the target audio information corresponding to the target audio feature to the terminal includes: The target audio information corresponding to the target audio feature is obtained using the correspondence, and the target audio information is sent to the terminal.

7. The audio recognition method according to claim 1, characterized in that, The audio to be identified is the audio obtained by the recording device from the user's humming voice; or, The audio to be identified is the audio obtained by the radio device after collecting audio played by other electronic devices.

8. An electronic device, characterized in that, Includes memory and processor, wherein: The memory is used to store computer programs; The processor is configured to execute the computer program to implement the audio recognition method as described in any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that, Used to store a computer program, wherein the computer program, when executed by a processor, implements the audio recognition method as described in any one of claims 1 to 7.