Audio recognition method, apparatus, electronic device, and computer program product

By employing multi-level feature map reconstruction and data augmentation techniques, the problem of lost location information in deep learning audio recognition was solved, improving the accuracy of audio recognition and model performance, especially in the accuracy of recognizing specific parts of a song.

CN115240704BActive Publication Date: 2026-03-31BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-13
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing deep learning-based audio recognition technologies ignore the location information of audio data during feature extraction, resulting in insufficient recognition accuracy.

Method used

A multi-level feature map extraction technique is adopted, which combines the next-level and previous-level feature maps to reconstruct the target feature map containing rich semantic information and high-resolution location information. Furthermore, data augmentation techniques are used to enhance the diversity and quantity of training data.

Benefits of technology

It improves the accuracy of audio recognition and model performance, especially in the accuracy of recognizing specific parts of a song, such as the chorus.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115240704B_ABST
    Figure CN115240704B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide an audio recognition method, device, electronic equipment and computer program product. The method can include obtaining a target feature map of audio data based on a multi-level feature map of the audio data. The method can also include determining a feature representation of the audio data based on the target feature map. In addition, the method can further include determining a recognition result of the audio data based on at least the feature representation. By implementing the technical solutions of the present disclosure, the determined feature representation has high-resolution position information, thereby optimizing the model performance and improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this disclosure relate to the field of data processing, and more specifically, to audio recognition methods, apparatus, electronic devices, and computer program products. Background Technology

[0002] The technology for intelligently recognizing audio data such as songs and human voices is crucial in many fields of research. Therefore, deep learning-based audio recognition technology has wide applications in many areas. For example, current deep learning-based audio recognition technologies typically utilize operations such as convolution to extract features, which contain rich high-level semantic information but also neglect other information. There is an urgent need for an audio recognition technology whose extracted features can include more information. Summary of the Invention

[0003] The embodiments disclosed herein provide an audio recognition scheme.

[0004] In a first aspect of this disclosure, an audio recognition method is provided. The method may include obtaining a target feature map of the audio data based on a multi-level feature map of the audio data. The method may further include determining a feature representation of the audio data based on the target feature map. Furthermore, the method may further include determining a recognition result of the audio data based at least on the feature representation.

[0005] In a second aspect of this disclosure, an audio recognition device is provided, which may include: a target feature map acquisition module configured to acquire a target feature map of the audio data based on a multi-level feature map of the audio data; a feature representation determination module configured to determine a feature representation of the audio data based on the target feature map; and a recognition result determination module configured to determine a recognition result of the audio data based at least on the feature representation.

[0006] In a third aspect of this disclosure, an electronic device is provided, comprising: a processor; and a memory coupled to the processor, the memory having instructions stored therein, the instructions causing the electronic device to perform actions when executed by the processor, the actions including: acquiring a target feature map of the audio data based on a multi-level feature map of the audio data; determining a feature representation of the audio data based on the target feature map; and determining a recognition result of the audio data based at least on the feature representation.

[0007] In a fourth aspect of this disclosure, a computer program product is provided, which is tangibly stored on a computer-readable medium and includes machine-executable instructions that, when executed, cause a machine to perform any step of the method according to the first aspect.

[0008] This content section is provided for the purpose of presenting a simplified form of the chosen concepts, which will be further described in the detailed embodiments below. This content section is not intended to identify key or major features of this disclosure, nor is it intended to limit the scope of this disclosure. Attached Figure Description

[0009] The above and other objects, features, and advantages of this disclosure will become more apparent from the accompanying drawings, in which the same or similar reference numerals generally represent the same or similar parts. In the drawings:

[0010] Figure 1 A schematic diagram of an example environment in which several embodiments of the present disclosure can be implemented is shown;

[0011] Figure 2 A schematic diagram of a detailed example environment for training and applying a model according to embodiments of the present disclosure is shown;

[0012] Figure 3 A flowchart of a process for audio recognition according to an embodiment of the present disclosure is shown;

[0013] Figure 4 A schematic diagram of an example environment representing a defined feature according to an embodiment of the present disclosure is shown;

[0014] Figure 5 A schematic diagram of feature diagrams according to an embodiment of the present disclosure is shown;

[0015] Figure 6 A schematic diagram of a multi-level feature map according to an embodiment of the present disclosure is shown;

[0016] Figure 7 A schematic diagram of a model training architecture according to an embodiment of the present disclosure is shown;

[0017] Figure 8 A schematic diagram of an audio recognition device according to an embodiment of the present disclosure is shown; and

[0018] Figure 9 A schematic block diagram of an example device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation

[0019] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0020] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0021] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0022] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0023] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0024] The principles of this disclosure will now be described with reference to several exemplary embodiments shown in the accompanying drawings.

[0025] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0026] In embodiments of this disclosure, the term "data" can refer to real-time data to be identified, such as an audio segment extracted from a song, which can be identified using a trained recognition model. Furthermore, the term "data" can also refer to data containing labeled information, such as model training data. This labeled information can be, for example, pre-labeled classification information. The term "classification" generally refers to the recognition result of an audio segment; for example, a recognition model can determine whether an audio frame belongs to a certain type of audio, such as a chorus. The term "feature representation" generally refers to features extracted from data using at least a portion of the network in a deep neural network.

[0027] As described above, with the continuous development of computer technology, deep neural networks are widely used in all aspects of people's lives. To better perform audio recognition classification tasks, the training process of traditional audio recognition models needs to be optimized. In the training process of traditional audio recognition models, as the model progresses, the resolution of the extracted feature maps gradually decreases. Although the feature maps with reduced resolution carry higher-level semantic information, the sacrifice of resolution causes the feature maps to lose precise location information. It should be understood that the "location information" mentioned in this article mainly refers to the position of an audio segment within an audio sequence, such as the start and end times of that audio segment.

[0028] According to embodiments of this disclosure, a scheme for audio recognition is proposed. When extracting a target feature map for determining feature representation, this scheme utilizes not only the nearest parent feature map of the target feature map, but also feature maps extracted at each or multiple levels. This results in a final target feature map containing both rich semantic information and high-resolution location information, thereby addressing the aforementioned problems and / or other potential issues.

[0029] Furthermore, during model training, the quantity and diversity of training data directly determine the model's performance. For audio recognition training data, insufficient sample size and / or diversity negatively impact the training of audio recognition models. Therefore, subsequent embodiments of this disclosure also provide schemes for enhancing the feature representation determined by the target feature map described above.

[0030] The embodiments of this disclosure will be described in detail below with reference to example scenarios. It should be understood that this is for illustrative purposes only and is not intended to limit the scope of this disclosure in any way.

[0031] Figure 1 A block diagram of an example system 100 for audio recognition according to an embodiment of the present disclosure is shown. It should be understood that... Figure 1 The system 100 shown is merely one example of an embodiment that can be implemented according to this disclosure and is not intended to limit the scope of this disclosure. The embodiments of this disclosure are equally applicable to other systems or architectures.

[0032] like Figure 1 As shown, system 100 may include computing device 120. Computing device 120 may be configured to receive audio data 110 and output a recognition result 130 associated with the audio data 110. In some embodiments, audio data 110 is a spectrogram of the time-domain audio data constant Q transform or other transforms.

[0033] In some embodiments, computing device 120 may acquire audio data 110. In some embodiments, audio data 110 may be an audio segment to be identified. In other embodiments, audio data 110 may include multiple training samples used to train a deep neural network or machine learning model (also referred to as a target model). Audio data 110 may have corresponding annotation information. Such annotation information may be generated by manual annotation, automatic model annotation, or other suitable methods.

[0034] In this disclosure, the target model can be designed to perform audio recognition tasks. Examples of target models include, but are not limited to, various deep neural networks (DNNs), convolutional neural networks (CNNs), support vector machines (SVMs), decision trees, random forest models, and so on. In implementations of this disclosure, the target model may also be referred to as a "recognition model." In the following text, the terms "recognition model," "neural network," "learning model," "learning network," "model," and "network" are used interchangeably.

[0035] In some embodiments, the computing device 120 may include, but is not limited to, personal computers, server computers, handheld or laptop devices, mobile devices (such as mobile phones, personal digital assistants (PDAs), media players, etc.), consumer electronics, minicomputers, mainframe computers, cloud computing resources, etc.

[0036] In some embodiments, the identification result 130 may be set as classification information determined from the audio data 110, such as whether the audio data 110, as an audio segment of a song, belongs to the chorus category. Alternatively or additionally, the identification result 130 may also be set as a prediction result that is corrected or updated during model training (this result is compared with the labeled ground truth result in a subsequent process to determine the loss function).

[0037] It should be understood that the devices and / or units included in system 100 are exemplary only and are not intended to limit the scope of this disclosure. It should be understood that system 100 may also include additional devices and / or units not shown. For example, in some embodiments, the computing device 120 of system 100 may further include a storage unit (not shown) for storing pre-input hyperparameters, etc., and a trained model.

[0038] The following will refer to Figure 2 The training and use of the model in computing device 120 are described.

[0039] Figure 2 A schematic diagram of a detailed example environment 200 according to an embodiment of the present disclosure is shown. Figure 1Similarly, example environment 200 may include computing device 220, audio data 210 input to computing device 220, and recognition results 230 output from computing device 220. The difference is that example environment 200 can generally include model training system 260 and model application system 270. As an example, model training system 260 and / or model application system 270 can be, for example... Figure 1 The computing device 120 shown or such Figure 2 The examples are implemented in the computing device 220 shown. It should be understood that the description of the structure and functionality of the example environment 200 is for illustrative purposes only and is not intended to limit the scope of the topics described herein. The topics described herein may be implemented in different structures and / or functionalities.

[0040] As previously described, the process of processing the input audio data 110 to determine the recognition result 230, such as classification information for audio segments, can be divided into two stages: a model training stage and a model application stage. As an example, in the model training stage, the model training system 260 can use the training dataset 250 to train a recognition model 240 to perform the corresponding function. It should be understood that the training dataset 250 can be a combination of multiple sample data (as input to the recognition model 240) and corresponding labeled supervision information (or "labels," "ground truth"). In the model application stage, the model application system 270 can receive the trained recognition model 240. Thus, the recognition model 240 loaded into the computing device 220 of the model application system 270 can determine the recognition result 230 based on the audio data 210.

[0041] In other embodiments, the recognition model 240 can be constructed as a learning network. In some embodiments, the learning network may include multiple networks, each of which may be a multi-layer neural network composed of a large number of neurons. Through a training process, the corresponding parameters of the neurons in each network can be determined. The parameters of the neurons in these networks are collectively referred to as the parameters of the recognition model 240.

[0042] The training process of the recognition model 240 can be performed iteratively until at least some of the parameters of the recognition model 240 converge or until a predetermined number of iterations is reached, thereby obtaining the final model parameters.

[0043] The technical solutions described above are for illustrative purposes only and are not intended to limit this disclosure. It should be understood that other networks can also be arranged in different ways and with different connections. To more clearly explain the principles of the above solutions, references will be made below. Figure 3 The process of determining the recognition result 130 from the audio data 110 will be described in more detail.

[0044] Figure 3A flowchart of a process 300 for audio recognition according to an embodiment of the present disclosure is shown. In some embodiments, process 300 may be performed at... Figure 1 The computing device 120 and Figure 2 This is implemented in computing device 220. Now refer to... Figure 3 The audio recognition process 300 according to an embodiment of this disclosure is described. For ease of understanding, the specific examples mentioned in the following description are exemplary and are not intended to limit the scope of protection of this disclosure.

[0045] In step 302, the computing device 120 can obtain a target feature map of the audio data 110 based on the multi-level feature map of the audio data 110. Then, in step 304, the computing device 120 can determine the feature representation of the audio data 110 based on the target feature map.

[0046] To clearly describe the process of determining the "feature representation" mentioned in this disclosure, reference is now made to... Figure 4 Describe the feature extraction process. Figure 4 A schematic diagram of an example environment 400 representing a defined feature according to an embodiment of the present disclosure is shown.

[0047] like Figure 4 As shown, the example environment 400 includes audio data 410, a feature extraction network 420, and a feature representation 430. It should be understood that audio data 410 can be audio data 110 or a segment of audio data 110. After the audio data 410 is input into the feature extraction network 420, the feature extraction network 420 performs feature extraction operations on the audio data 410. As an example, the feature extraction network 420 can be as follows: Figure 4 The deep neural network or multi-layer feature extractor shown is illustrated. As shown in the figure, the feature extraction network 420 may contain at least a first-level extractor 421 and a second-level extractor 422. It should be understood that the feature extraction network 420 may also contain more levels of extractors.

[0048] To obtain the target feature map, the computing device 120 can utilize a feature extraction network 420, for example, including at least the first-level extractor 421 and the second-level extractor 422 described above, to obtain a multi-level feature map of the audio data 410. As an example, the first-level extractor 421 and the second-level extractor 422 can be convolutional neural networks; therefore, the first-level extractor 421 can perform a convolution operation on the audio data 410 to obtain a first-level feature map, and the second-level extractor 422 can perform a convolution operation on the first-level feature map to obtain a second-level feature map.

[0049] It should be noted that the convolution operation is essentially a downsampling process. Since the next level feature map in a multi-level feature map is extracted from the previous level feature map, the resolution of the second level feature map is lower than that of the first level feature map.

[0050] Subsequently, the computing device 120 can perform feature reconstruction based at least on the next-level feature map and the previous-level feature map to determine the target feature map, thereby obtaining the feature vector of the audio data 410 in the abstract space, i.e., the feature representation 430. In this way, the resolution of the next-level feature map is improved to the resolution of the previous-level feature map through feature reconstruction, and the feature reconstruction is based at least on the next-level feature map and the previous-level feature map. This includes both the rich semantic information extracted from the next-level feature map and the high resolution from the previous-level feature map, making it easier to locate the position of specific types of audio segments.

[0051] To clearly describe the “feature map” mentioned in this disclosure, reference is now made to… Figure 5 Example format for describing feature maps. Figure 5 A schematic diagram of feature figure 510 according to an embodiment of the present disclosure is shown. (As shown) Figure 5 As shown, feature map 510 can be a set of feature data determined based on audio data 410, where A…I are specific numerical values ​​of the aforementioned feature data. For example, feature map 510 can be a 100×100 matrix. After feature map 510 undergoes convolution operation by the first-level extractor 421, feature map 510 is downsampled to, for example, a 50×50 matrix, and after further convolution operation by the second-level extractor 422, feature map 510 is downsampled to, for example, a 25×25 matrix. For the above feature reconstruction process, feature map 510, as a 25×25 matrix, can be upsampled to, for example, a 50×50 matrix, and then upsampled to, for example, a 100×100 matrix. It should be understood that the feature reconstruction process is not limited to this; for a more detailed description of the feature extraction and feature reconstruction process, please refer to... Figure 6 Describe the architecture for determining the target feature map.

[0052] Figure 6 A schematic diagram of a multi-level feature map 600 according to an embodiment of the present disclosure is shown. Figure 6 As shown, the multi-level feature map 600 includes a first-level feature map 601, a second-level feature map 602, a third-level feature map 603, a feature map 604 generated based on the third-level feature map 603, a feature map 605 generated based on the feature map 604 and the second-level feature map 602, and a feature map 606 generated based on the feature map 605 and the first-level feature map 601.

[0053] exist Figure 6 In the middle, the first-level feature map 601 can be composed of Figure 4 The first-level extractor 421 shown extracts from the audio data 410, and the second-level feature map 602 can be derived from... Figure 4The second-level extractor 422 shown extracts from the first-level feature map 601, and consequently, the third-level feature map 603 may be extracted from the second-level feature map 602. It should be understood that... Figure 6 The multi-level feature map 600 shown can have more levels, and the number of levels is related to the network structure of the model.

[0054] Therefore, during feature reconstruction, the computing device 120 can directly copy the values ​​from the third-level feature map 603 into the feature map 604. Then, the computing device 120 can upsample the feature map 604, that is, expand the feature map 604 into a backup feature map 605. In other words, the computing device 120 can copy the values ​​from the upsampled feature map 604 into the feature map 605, and perform operations such as averaging or other calculations between the values ​​from the second-level feature map 602 (at the same level as feature map 605) and the values ​​from feature map 605, storing the result in feature map 605. Similarly, the computing device 120 can further upsample the feature map 605, that is, expand the feature map 605 into a backup feature map 606, and perform operations such as averaging or other calculations between the values ​​from the first-level feature map 601 (at the same level as feature map 606) and the values ​​from feature map 606, storing the result in feature map 606. At this point, feature map 606 is the target feature map. In this way, the target feature map contains both rich semantic information and high-resolution location information, thereby optimizing model performance.

[0055] Back Figure 3 In step 306, computing device 120 can determine recognition result 130 of audio data 110 based at least on feature representation.

[0056] In some embodiments, audio data 110 is an audio segment of a song. To determine the recognition result 130 of audio data 110, computing device 120 can determine whether the audio segment belongs to the chorus category. Thus, the chorus portion of a song can be automatically identified. It should be understood that this disclosure is not limited to identifying the chorus portion of a song, but can also identify other parts of a song, such as verses, transitional phrases, bridges, etc., and can also identify categorizable parts of other audio data.

[0057] In this way, the feature data determined through the above embodiments contains richer information and has more accurate location information compared with traditional audio recognition modules, thereby improving the performance of the model.

[0058] The above embodiments mainly involve the application of the recognition model 240. The training process of the recognition model 240 will be described in detail below. During the model training process, the audio data 110 can be training data or training dataset. After the trained model determines the recognition result 130, the computing device 120 can further determine the loss function value of the trained recognition model based on the recognition result 130 and the pre-labeled ground truth result of the training data, so as to update the parameters of the recognition model.

[0059] To determine the loss function value of the model, computing device 120 needs to compare the ground truth labels with the real-time generated recognition results. Figure 7 A schematic diagram of a model training architecture 700 according to an embodiment of the present disclosure is shown.

[0060] like Figure 7 As shown, audio data 701 can be input into extraction module 710 to determine the feature representation of audio data 701. Then, the determined feature representation is input into prediction module 720 to determine the prediction result of the feature representation of audio data 701. Thus, loss determination module 730 can determine the loss function value 703 of the model based on the determined result and the ground truth label 702 of audio data 701.

[0061] In some embodiments, to optimize the performance of the (generalization) model, computing device 120 may perform data augmentation on the feature representation determined by extraction module 710. As an example, computing device 120 may utilize... Figure 7 The enhancement module 740 determines the distribution of feature representations corresponding to audio segments that belong to or do not belong to the chorus category, and then determines the sampled feature representations in the distribution as additional feature representations.

[0062] In some embodiments, in order to determine the sampled feature representation as an additional feature representation, the computing device 120 may sample a predetermined number of feature representations from the distribution as additional feature representations. Thus, the computing device 120 may input a feature representation determined by the extraction module 710 and multiple additional feature representations obtained through data augmentation into the fully connected layer of the recognition model to determine the recognition result or prediction result. In this way, the present disclosure can augment more training data at the feature vector level, thereby increasing the amount and diversity of training data.

[0063] It should be understood that the feature representation obtained through data augmentation It can be generated based on the following formula (1):

[0064]

[0065] Where a iy is the feature representation, where i is the i-th row of features in the feature representation determined by the extraction module 710. i This indicates the label category of the i-th frame (such as, chorus). Represents category y i The covariance matrix. λ is the hyperparameter of the model, which can be set to λ > 0, for example.

[0066] It should be understood that when the number of sampled feature representations is large, the computational load for model training will increase significantly. Therefore, the computing device 120 can determine the upper limit of the loss function of the recognition model by setting the number of sampled feature representations to positive infinity, thereby determining the loss function value.

[0067] Specifically, assuming the dataset size is N and the number of sampled feature representations is M, then the number of samples for the augmented training data is N×(M+1). In some embodiments, the cross-entropy loss function can be used to train the module. For a fully connected layer, the weight W corresponding to class c can be represented as w c And the corresponding offset b is represented as b c When M is positive infinity:

[0068]

[0069] Formula (2) is equivalent to the following loss function formula:

[0070]

[0071] in

[0072] Using Jensen's inequality E[logX]≤logE[X], the upper bound of the loss function can be derived. That is, as shown in formula (5):

[0073]

[0074] Ultimately, the upper bound of the loss function It can be derived as the following formula (6):

[0075]

[0076] in

[0077] In this way, the loss function can be determined without spending a lot of computational resources as in formula (1), thus quickly obtaining the loss function value and optimizing model training.

[0078] This disclosure also provides a video recognition device. Specifically, Figure 8A schematic diagram of an audio recognition device 800 according to an embodiment of the present disclosure is shown. Figure 8 As shown, the audio recognition device 800 may include at least a target feature map acquisition module 802, a feature representation determination module 804, and a recognition result determination module 806. The target feature map acquisition module 802 can acquire a target feature map of the audio data based on multi-level feature maps of the audio data. The feature representation determination module 804 can further determine the feature representation of the audio data based on the acquired target feature map. Furthermore, the recognition result determination module 806 can further determine the recognition result of the audio data based at least on the determined feature representation.

[0079] In some embodiments, the target feature map acquisition module 802 may include a multi-level feature map acquisition submodule, which is used to acquire multi-level feature maps of the audio data. It should be understood that a lower-level feature map in the multi-level feature map is extracted from a higher-level feature map. The multi-level feature map acquisition submodule may include a first-level extractor, a second-level extractor, etc. The first-level extractor may perform a convolution operation on the audio data to obtain a first-level feature map, and the second-level extractor may perform a convolution operation on the first-level feature map to obtain a second-level feature map. Furthermore, the target feature map acquisition module 802 may also include a target feature map determination submodule, which is used to perform feature reconstruction based at least on the lower-level and higher-level feature maps to determine the target feature map.

[0080] In some embodiments, the target feature map determination submodule may expand the second-level feature map into a first-level backup feature map when performing feature reconstruction, and determine the target feature map based on the first-level backup feature map and the first-level feature map.

[0081] In some embodiments, the audio data may be training data, and the audio recognition device 800 may further include: a loss function value determination submodule, used to determine the loss function value of the trained recognition model based on the recognition result and the pre-labeled ground truth result of the training data, so as to update the parameters of the recognition model.

[0082] In some embodiments, the audio recognition device 800 may further include: a distribution determination module, configured to determine the distribution of feature representations corresponding to audio segments that belong to or do not belong to the chorus category; and an additional feature representation determination module, configured to determine the sampled feature representations in the distribution as additional feature representations.

[0083] In some embodiments, the additional feature representation determination module may be configured to sample a predetermined number of feature representations in the distribution as additional feature representations.

[0084] In some embodiments, the loss function value determination submodule can be configured to determine the loss function value by setting a predetermined number of upper limits of the loss function of the recognition model to positive infinity.

[0085] In some embodiments, the recognition result determination module 806 may be configured to input the feature representation and additional feature representation into a fully connected layer of the recognition model to determine the recognition result.

[0086] In some embodiments, the audio data is an audio segment of a song, and the identification result determination module 806 may include a classification module for determining whether the audio segment belongs to the chorus category or not.

[0087] Figure 9 A schematic block diagram of an example device 900 that can be used to implement embodiments of the present disclosure is shown. For example, as Figure 1 The computing device 120 shown can be implemented by device 900. As shown, device 900 includes a central processing unit (CPU) 901, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 902 or loaded from storage unit 908 into random access memory (RAM) 903. The RAM 903 can also store various programs and data required for the operation of device 900. The CPU 901, ROM 902, and RAM 903 are interconnected via bus 904. Input / output (I / O) interface 905 is also connected to bus 904.

[0088] Multiple components in device 900 are connected to I / O interface 905, including: input unit 906, such as keyboard, mouse, etc.; output unit 907, such as various types of monitors, speakers, etc.; storage unit 908, such as disk, optical disk, etc.; and communication unit 909, such as network card, modem, wireless transceiver, etc. Communication unit 909 allows device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks. It should be understood that this disclosure can utilize output unit 907 to display real-time dynamic changes in user satisfaction, key factor identification information for group or individual users of satisfaction, optimization strategy information, and strategy implementation effectiveness evaluation information, etc.

[0089] Processing unit 901 may be implemented by one or more processing circuits. Processing unit 901 may be configured to perform the various processes and procedures described above, such as process 300. For example, in some embodiments, process 300 may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program may be loaded and / or installed on device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by CPU 901, one or more steps of process 300 described above may be performed.

[0090] Detailed description of effects

[0091] By implementing the above embodiments, the performance of the trained model can be significantly improved. To verify the model's performance, various test datasets are used to test the performance of the trained model and compare it with several traditional models:

[0092] For the RWC (Real World Computation) dataset, the AUC (Area Under the Curve) score of the CNMF (Convolutional Nonnegative Matrix Factorization) model is 0.526, the AUC score of the SCluster model is 0.533, the AUC score of the Highlighter model is 0.804, the AUC score of the Multi2021 model is 0.819, the AUC score of the DeepChorus model is 0.842, while the AUC score of the trained model disclosed in this paper is 0.906.

[0093] For the SP (salami-pop) dataset, the AUC score of the CNMF model is 0.543, the AUC score of the SCluster model is 0.545, the AUC score of the Highlighter model is 0.703, the AUC score of the Multi2021 model is 0.675, the AUC score of the DeepChorus model is 0.780, while the AUC score of the trained model disclosed in this paper is 0.887.

[0094] For the SL (salami-live) dataset, the AUC score of the CNMF model is 0.478, the AUC score of the SCluster model is 0.551, the AUC score of the Highlighter model is 0.671, the AUC score of the Multi2021 model is 0.633, the AUC score of the DeepChorus model is 0.765, while the AUC score of the trained model disclosed in this paper is 0.831.

[0095] For the DC (Di-Chorus) dataset, the AUC score of the CNMF model is 0.488, the AUC score of the SCluster model is 0.568, the AUC score of the Highlighter model is 0.553, the AUC score of the DeepChorus model is 0.811, while the AUC score of the trained model disclosed in this paper is 0.872.

[0096] Furthermore, through other experiments, the F-score of the model disclosed herein is also higher than that of traditional modules. Therefore, it is evident that the audio recognition module trained according to the embodiments of this disclosure exhibits significantly improved performance compared to traditional models.

[0097] This disclosure can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of this disclosure.

[0098] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0099] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0100] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.

[0101] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0102] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0103] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0104] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0105] According to one or more embodiments of this disclosure, Example 1. An audio recognition method includes: obtaining a target feature map of the audio data based on a multi-level feature map of the audio data; determining a feature representation of the audio data based on the target feature map; and determining a recognition result of the audio data based at least on the feature representation.

[0106] Example 2. According to the method of Example 1, obtaining the target feature map includes: obtaining the multi-level feature map of the audio data, wherein the next-level feature map in the multi-level feature map is extracted from the previous-level feature map; and performing feature reconstruction based at least on the next-level feature map and the previous-level feature map to determine the target feature map.

[0107] Example 3. According to the method of Example 2, the multi-level feature map includes at least: a first-level feature map extracted from the audio data; and a second-level feature map extracted based on the first-level feature map.

[0108] Example 4. According to the method of Example 3, the feature reconstruction includes at least: expanding the second-level feature map into a first-level backup feature map; and determining the target feature map based on the first-level backup feature map and the first-level feature map.

[0109] Example 5. According to the method of Example 1, wherein the audio data is training data, and the method further includes: determining the loss function value of the trained recognition model based on the recognition result and the pre-labeled ground truth result of the training data, so as to update the parameters of the recognition model.

[0110] Example 6. The method according to Example 5 further includes: determining the distribution of feature representations corresponding to audio segments that belong to or do not belong to the chorus category; and determining the sampled feature representations in the distribution as additional feature representations.

[0111] Example 7. The method according to Example 6, wherein determining the sampled feature representation as the additional feature representation comprises: sampling a predetermined number of feature representations in the distribution as the additional feature representations.

[0112] Example 8. The method according to Example 7, wherein determining the loss function value comprises: determining the upper limit of the loss function of the recognition model by setting the predetermined number to positive infinity, thereby determining the loss function value.

[0113] Example 9. The method according to Example 6, wherein determining the recognition result based at least on the feature representation comprises: inputting the feature representation and the additional feature representation into a fully connected layer of the recognition model to determine the recognition result.

[0114] Example 10. According to the method of Example 1, wherein the audio data is an audio segment of a song, and determining the recognition result of the audio data includes: determining that the audio segment belongs to the chorus category; or determining that the audio segment does not belong to the chorus category.

[0115] According to one or more embodiments of this disclosure, Example 11. An audio recognition device includes: a target feature map acquisition module configured to acquire a target feature map of the audio data based on a multi-level feature map of the audio data; a feature representation determination module configured to determine a feature representation of the audio data based on the target feature map; and a recognition result determination module configured to determine a recognition result of the audio data based at least on the feature representation.

[0116] Example 12. The audio recognition apparatus according to Example 11, wherein the target feature map acquisition module includes: a multi-level feature map acquisition submodule configured to acquire the multi-level feature map of the audio data, wherein the next-level feature map in the multi-level feature map is extracted from the previous-level feature map; and a target feature map determination submodule configured to perform feature reconstruction based at least on the next-level feature map and the previous-level feature map to determine the target feature map.

[0117] Example 13. The audio recognition apparatus according to Example 12, wherein the multi-level feature map includes at least: a first-level feature map extracted from the audio data; and a second-level feature map extracted based on the first-level feature map.

[0118] Example 14. The audio recognition apparatus according to Example 13, wherein the target feature map acquisition module can be configured during the feature reconstruction to: expand the second-level feature map into a first-level backup feature map; and determine the target feature map based on the first-level backup feature map and the first-level feature map.

[0119] Example 15. The audio recognition apparatus according to Example 11, wherein the audio data is training data, and the audio recognition apparatus further includes: a loss function value determination submodule, configured to determine the loss function value of the trained recognition model based on the recognition result and a pre-labeled ground truth result of the training data, so as to update the parameters of the recognition model.

[0120] Example 16. The audio recognition device according to Example 15 further includes: a distribution determination module configured to determine the distribution of feature representations corresponding to audio segments belonging to or not belonging to the chorus category; and an additional feature representation determination module configured to determine the sampled feature representations in the distribution as additional feature representations.

[0121] Example 17. The audio recognition apparatus according to Example 16, wherein the additional feature representation determination module is configured to sample a predetermined number of feature representations in the distribution as the additional feature representations.

[0122] Example 18. The audio recognition apparatus according to Example 17, wherein the loss function value determination submodule is configured to determine the loss function value by setting the predetermined number to positive infinity to determine the upper limit of the loss function of the recognition model.

[0123] Example 19. An audio recognition apparatus according to Example 16, wherein the recognition result determination module is configured to input the feature representation and the additional feature representation into a fully connected layer of the recognition model to determine the recognition result.

[0124] Example 20. The audio recognition device according to Example 11, wherein the audio data is an audio segment of a song, and the recognition result determination module includes: a classification module configured to determine whether the audio segment belongs to the chorus category or not.

[0125] According to one or more embodiments of this disclosure, Example 21. An electronic device includes: a processor; and a memory coupled to the processor, the memory having instructions stored therein, the instructions causing the electronic device to perform actions when executed by the processor, the actions including: acquiring a target feature map of the audio data based on a multi-level feature map of the audio data; determining a feature representation of the audio data based on the target feature map; and determining a recognition result of the audio data based at least on the feature representation.

[0126] Example 22. The device according to Example 21, wherein obtaining the target feature map includes: obtaining the multi-level feature map of the audio data, wherein a next-level feature map in the multi-level feature map is extracted from a previous-level feature map; and performing feature reconstruction based at least on the next-level feature map and the previous-level feature map to determine the target feature map.

[0127] Example 23. The device according to Example 22, wherein the multi-level feature map includes at least: a first-level feature map extracted from the audio data; and a second-level feature map extracted based on the first-level feature map.

[0128] Example 24. The device according to Example 23, wherein the feature reconstruction includes at least: expanding the second-level feature map into a first-level backup feature map; and determining the target feature map based on the first-level backup feature map and the first-level feature map.

[0129] Example 25. The device according to Example 21, wherein the audio data is training data, and the method further includes: determining a loss function value of a trained recognition model based on the recognition result and a pre-labeled ground truth result of the training data, so as to update the parameters of the recognition model.

[0130] Example 26. The device according to Example 25 further includes: determining a distribution of feature representations corresponding to audio segments that belong to or do not belong to the chorus category; and determining the sampled feature representations in the distribution as additional feature representations.

[0131] Example 27. The device according to Example 26, wherein determining the sampled feature representation as the additional feature representation includes: sampling a predetermined number of feature representations in the distribution as the additional feature representation.

[0132] Example 28. The device according to Example 27, wherein determining the loss function value comprises: determining the upper limit of the loss function of the recognition model by setting the predetermined number to positive infinity, thereby determining the loss function value.

[0133] Example 29. The device according to Example 26, wherein determining the recognition result based at least on the feature representation comprises: inputting the feature representation and the additional feature representation into a fully connected layer of the recognition model to determine the recognition result.

[0134] Example 30. The device according to Example 21, wherein the audio data is an audio segment of a song, and determining the recognition result of the audio data includes: determining that the audio segment belongs to the chorus category; or determining that the audio segment does not belong to the chorus category.

[0135] According to one or more embodiments of this disclosure, Example 31. A computer program product tangibly stored on a computer-readable medium and comprising machine-executable instructions that, when executed, cause a machine to perform the method according to any one of Examples 1 to 10.

[0136] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, and are not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical applications, or technical improvements to the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for identifying a classification of a segment in a song, comprising: obtaining, based on a multi-level feature map of audio data, a target feature map of the audio data by fusing feature information of each level in the multi-level feature map, wherein the target feature map contains both high-resolution location information and semantic information; determining a feature representation of the audio data based on the target feature map; and determining, based at least on the feature representation, an identification result of whether the audio data belongs to a chorus classification, wherein the method further comprises: determining a distribution of feature representations corresponding to audio segments that belong to the chorus classification or do not belong to the chorus classification, and determining a sampling feature representation in the distribution as an additional feature representation, the feature representation and the additional feature representation being input into a fully connected layer of an identification model to determine the identification result.

2. The method of claim 1, wherein obtaining the target feature map comprises: obtaining the multi-level feature map of the audio data, a next-level feature map in the multi-level feature map being extracted from a previous-level feature map; and performing feature reconstruction based at least on the next-level feature map and the previous-level feature map to determine the target feature map.

3. The method of claim 2, wherein the multi-level feature map comprises at least: a first-level feature map extracted from the audio data; and a second-level feature map extracted from the first-level feature map.

4. The method of claim 3, wherein the feature reconstruction comprises at least: augmenting the second-level feature map into a first-level backup feature map; and determining the target feature map based on the first-level backup feature map and the first-level feature map.

5. The method of claim 1, wherein the audio data is training data, and the method further comprises: determining a loss function value of the trained identification model based on the identification result and a pre-labeled ground truth result of the training data, to update parameters of the identification model.

6. The method of claim 5, wherein determining the sampling feature representation as the additional feature representation comprises: sampling a predetermined number of feature representations in the distribution as the additional feature representation.

7. The method of claim 6, wherein determining the loss function value comprises: determining an upper limit of a loss function of the identification model by setting the predetermined number to positive infinity, to determine the loss function value.

8. The method of claim 1, wherein the audio data is an audio segment of a song, and determining an identification result of the audio data comprises: determining that the audio segment belongs to a chorus classification; or determining that the audio segment does not belong to a chorus classification.

9. An apparatus for identifying a classification of a segment in a song, comprising: a target feature map obtaining module configured to obtain, based on a multi-level feature map of audio data, a target feature map of the audio data by fusing feature information of each level in the multi-level feature map, wherein the target feature map contains both high-resolution location information and semantic information; ​ ​ ​ ​ ​ a feature representation determination module configured to determine a feature representation of the audio data based on the target feature map; an identification result determination module configured to determine an identification result of whether the audio data belongs to a chorus classification based on at least the feature representation; a distribution determination module configured to determine a distribution of feature representations corresponding to audio segments that belong to the chorus classification or do not belong to the chorus classification; and an additional feature representation determination module configured to determine a sampled feature representation in the distribution as an additional feature representation, the feature representation and the additional feature representation being input into a fully connected layer of an identification model to determine the identification result.

10. An electronic device for identifying a classification of a segment in a song, comprising: a processor; and a memory coupled with the processor, the memory having stored therein instructions which, when executed by the processor, cause the electronic device to perform acts comprising: obtaining a target feature map of audio data by fusing feature information of each level in a multi-level feature map of the audio data based on the multi-level feature map, wherein the target feature map contains both high-resolution location information and semantic information; determining a feature representation of the audio data based on the target feature map; and determining an identification result of whether the audio data belongs to a chorus classification based on at least the feature representation, wherein the acts further comprise: determining a distribution of feature representations corresponding to audio segments that belong to the chorus classification or do not belong to the chorus classification, and determining a sampled feature representation in the distribution as an additional feature representation, the feature representation and the additional feature representation being input into a fully connected layer of an identification model to determine the identification result.

11. A computer program product tangibly stored on a computer readable medium and comprising machine executable instructions that, when executed, cause a machine to perform the method of any one of claims 1 to 8. ​

Citation Information

Patent Citations

  • Dynamic face generation method based on facial emotion analysis

    CN114299578A

  • Voice positioning method and device, computer readable medium and electronic equipment

    CN114420097A