Feature Frequency Point Identification Model Training and Audio Fingerprint Identification Method, Device and Product
The training method for a feature frequency point recognition model enhances audio fingerprinting accuracy by filtering noise interference, ensuring accurate identification of audio features.
Patent Information
- Application Number
- CN202211094118.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-08
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-09-08
AI Technical Summary
In the prior art, when there is interference noise in the audio signal, the accuracy of characteristic frequency point recognition is low, making it difficult to accurately identify audio fingerprints.
By obtaining the difference between the reference feature frequency points of the noisy song audio and the predicted feature frequency points, adjusting the parameters of the neural network model, training to obtain the feature frequency point recognition model, eliminating the interference frequency points of the noise signal, and retaining the feature frequency points of the original song signal.
It improves the reliability and accuracy of characteristic frequency points, enhances the recognition accuracy of audio fingerprints, and effectively eliminates interference from noise signals.
Smart Images

Figure CN116092521B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of audio processing, and particularly to a method for training a characteristic frequency point recognition model, an audio fingerprint recognition method, a computer device, and a computer program product. Background Art
[0002] With the continuous development of audio processing technology, more and more audio applications provide audio matching technologies, and audio fingerprints are widely used in the field of audio matching.
[0003] In the related art, each characteristic frequency point with a specified characteristic in an audio signal can be identified, and an audio fingerprint of the audio signal can be obtained based on each characteristic frequency point. However, when there is interference noise in the audio, the audio fingerprint obtained by the above method is often inaccurate, and the recognition accuracy is low. Summary of the Invention
[0004] Based on this, it is necessary to provide a method for training a characteristic frequency point recognition model, an audio fingerprint recognition method, a computer device, and a computer program product for the above technical problems.
[0005] In a first aspect, the present application provides a method for training a characteristic frequency point recognition model. The method includes:
[0006] Obtain a noisy song audio of an original song audio; the noisy song signal of the noisy song audio includes a noise signal and an original song signal of the original song audio;
[0007] Determine a reference characteristic frequency point in the frequency domain of the original song signal;
[0008] Input the noisy song signal into a neural network model to be trained, and obtain a predicted characteristic frequency point associated with the original song signal in the noisy song signal in the frequency domain through the neural network model;
[0009] Adjust the model parameters of the neural network model based on the difference value between the predicted characteristic frequency point and the reference characteristic frequency point until the training end condition is satisfied, and obtain a trained characteristic frequency point recognition model.
[0010] In a second aspect, the present application further provides an audio fingerprint recognition method. The method includes:
[0011] Obtain a target song audio of an audio fingerprint to be recognized;
[0012] Input the audio signal of the target song audio into the trained feature frequency point recognition model to obtain multiple feature frequency points of the audio signal of the target song audio in the frequency domain output by the feature frequency point recognition model, where the feature frequency point recognition model is obtained according to the training method of the feature frequency point recognition model described in any one of the above;
[0013] Determine the audio fingerprint of the target song audio based on the multiple feature frequency points.
[0014] In a third aspect, the present application also provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0015] Obtain the noisy song audio of the original song audio; the noisy song signal of the noisy song audio includes a noise signal and the original song signal of the original song audio;
[0016] Determine the reference feature frequency points in the frequency domain of the original song signal;
[0017] Input the noisy song signal into the neural network model to be trained, and obtain the predicted feature frequency points associated with the original song signal in the noisy song signal in the frequency domain through the neural network model;
[0018] Based on the difference value between the predicted feature frequency points and the reference feature frequency points, adjust the model parameters of the neural network model until the training end condition is met to obtain the trained feature frequency point recognition model.
[0019] In a fourth aspect, the present application also provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0020] Obtain the target song audio of the audio fingerprint to be recognized;
[0021] Input the audio signal of the target song audio into the trained feature frequency point recognition model to obtain multiple feature frequency points of the audio signal of the target song audio in the frequency domain output by the feature frequency point recognition model, where the feature frequency point recognition model is obtained according to the training method of the feature frequency point recognition model described in any one of the above;
[0022] Determine the audio fingerprint of the target song audio based on the multiple feature frequency points.
[0023] In a fifth aspect, the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the following steps are implemented:
[0024] Obtain a noisy song audio of the original song audio; the noisy song signal of the noisy song audio includes a noise signal and the original song signal of the original song audio;
[0025] Determine a reference characteristic frequency point in the frequency domain of the original song signal;
[0026] Input the noisy song signal into a neural network model to be trained, and obtain, through the neural network model, a predicted characteristic frequency point associated with the original song signal in the noisy song signal in the frequency domain;
[0027] Based on the difference value between the predicted characteristic frequency point and the reference characteristic frequency point, adjust the model parameters of the neural network model until the training end condition is satisfied, and obtain a trained characteristic frequency point recognition model.
[0028] In a sixth aspect, the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the following steps are implemented:
[0029] Obtain a target song audio of the audio fingerprint to be recognized;
[0030] Input the audio signal of the target song audio into the trained characteristic frequency point recognition model, and obtain multiple characteristic frequency points in the frequency domain of the audio signal of the target song audio output by the characteristic frequency point recognition model, where the characteristic frequency point recognition model is obtained according to the training method of the characteristic frequency point recognition model described in any one of the above;
[0031] Determine the audio fingerprint of the target song audio based on the multiple characteristic frequency points.
[0032] The training method of the above-mentioned characteristic frequency point recognition model, the audio fingerprint recognition method, the computer device and the computer program product can obtain the noisy song audio of the original song audio. Among them, the noisy song signal of the noisy song audio includes the noise signal and the original song signal of the original song audio. Furthermore, the reference characteristic frequency points in the frequency domain of the original song signal can be determined. The noisy song signal is input into the neural network model to be trained, and the predicted characteristic frequency points associated with the original song signal in the noisy song signal are obtained through the neural network model. Based on the difference value between the predicted characteristic frequency points and the reference characteristic frequency points, the model parameters of the neural network model are adjusted until the training end condition is met, and the trained characteristic frequency point recognition model is obtained. In the solution of this embodiment, during the training process, the neural network model is used to identify the predicted characteristic frequency points from the noisy song signal containing the noise signal and the original song signal, and learn based on the reference characteristic frequency points of the original song signal. The model can eliminate the characteristic frequency points associated with the noise signal in the input song signal, and only retain the characteristic frequency points associated with the original song signal in the noisy song signal, effectively eliminating the interference frequency points generated by the noise signal, increasing the reliability and accuracy of the identified characteristic frequency points, and further improving the recognition accuracy of the audio fingerprint. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 FIG. is a schematic flowchart of a training method for a characteristic frequency point recognition model in an embodiment;
[0034] Figure 2 FIG. is a schematic flowchart of steps for obtaining a noisy song audio in an embodiment;
[0035] Figure 3 FIG. is a schematic flowchart of steps for obtaining spectrum information in an embodiment;
[0036] Figure 4 FIG. is a schematic diagram of the spectrum information of a noisy song audio in an embodiment;
[0037] Figure 5 FIG. is an application environment diagram of an audio fingerprint recognition method in an embodiment;
[0038] Figure 6 FIG. is a schematic flowchart of an audio fingerprint recognition method in an embodiment;
[0039] Figure 7 FIG. is an internal structure diagram of a computer device in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] In order to make the objectives, technical solutions, and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0041] The present application provides a method for training a characteristic frequency point recognition model, and this method can be executed by computer devices such as terminals and servers. Among them, the terminal can be, but is not limited to, various personal computers, laptop computers, and tablet computers; the server can be implemented by an independent server or a server cluster composed of multiple servers. In one embodiment, as Figure 1 shown, a method for training a characteristic frequency point recognition model is provided, and this method may include the following steps:
[0042] S101, obtain the noisy song audio of the original song audio; the noisy song signal of the noisy song audio includes a noise signal and the original song signal of the original song audio.
[0043] In practical applications, the noisy song audio of the original song audio can be obtained.
[0044] Among them, the original song audio can be the song audio obtained from a preset song library, and the audio signal of the original song audio may not include a noise signal of a preset type; while the noisy song audio can be the song audio obtained by adding a noise signal of a preset type to the original song audio. In other words, the audio signal of the noisy song audio includes, in addition to the audio signal of the original song audio, a noise signal. For the convenience of distinction, in this embodiment, the audio signal of the noisy song audio is referred to as the noisy song signal, and the audio signal of the original song audio is referred to as the original song signal.
[0045] S102, determine the reference characteristic frequency points in the frequency domain of the original song signal.
[0046] As an example, the characteristic frequency points can be the frequency points with preset frequency point characteristics, and the image formed by multiple characteristic frequency points in a two-dimensional plane (time-frequency) can also be called a constellation diagram. For the convenience of distinguishing the characteristic frequency points extracted based on the original song signal and the characteristic frequency points extracted based on the noisy song signal, the characteristic frequency points extracted based on the original song signal can be called reference characteristic frequency points.
[0047] The frequency point feature can be the feature that a frequency point has in terms of the energy in the frequency domain. Exemplarily, the preset frequency point feature can be at least one of the following: the energy in the frequency domain exceeds a preset energy threshold, and the change amount of the energy in the frequency domain of a frequency point exceeds a preset change amount threshold. Among them, the change amount of the energy in the frequency domain can be the change amount of the energy in the frequency domain of a frequency point compared to the energy in the frequency domain of one or more adjacent frequency points, and this change amount can be an absolute value or a relative value (such as a percentage). In one example, if the frequency point feature used to screen for characteristic frequency points is that the change amount of the energy in the frequency domain of a frequency point exceeds a preset change amount threshold, then the screened characteristic frequency points can also be called local peak points.
[0048] In practical applications, after obtaining the original song signal of the original song audio, the characteristic frequency points of the original song signal in the frequency domain can be obtained. For example, the relationship between the sound intensity and time in the original song audio can be obtained to get the original song signal in the time domain, and the original song signal can be subjected to time-frequency transformation to obtain the original song signal in the frequency domain. Among them, the original song signal in the frequency domain can be the spectral information (such as a spectrogram) of the original song audio, and this spectral information includes multiple frequency points of the original song audio at different frequencies. Furthermore, each frequency point with a preset frequency point feature can be screened out from the multiple frequency points as the characteristic frequency points of the original song signal in the frequency domain.
[0049] S103, input the noisy song signal into the neural network model to be trained, and obtain the predicted characteristic frequency points associated with the original song signal in the noisy song signal through the neural network model.
[0050] S104, based on the difference value between the predicted characteristic frequency points and the reference characteristic frequency points, adjust the model parameters of the neural network model until the training end condition is met, and obtain the trained characteristic frequency point recognition model.
[0051] In a specific implementation, the characteristic frequency points in the frequency domain of the noisy song signal can be obtained. When obtaining the characteristic frequency points, although in related methods, frequency points with preset frequency point features can be screened out from the noisy song signal in the frequency domain as the characteristic frequency points of the noisy song signal, however, this method extracts the characteristic frequency points without discrimination, that is, it does not consider whether the extracted frequency points with preset frequency point features come from the original song signal or the noise signal. As long as a certain frequency point has the preset frequency point feature, it will be extracted as a characteristic frequency point. It can be understood that if the noise signal also contains frequency points with preset frequency point features, the frequency point extraction method of the related technology will wrongly extract this frequency point and cause interference, so that the finally obtained characteristic frequency points not only include the characteristic frequency points for the original song signal, but also include the characteristic frequency points for the noise signal, resulting in the characteristic frequency points being difficult to accurately reflect the characteristics of the original song signal itself.
[0052] Based on this, the present application can input the noisy song signal into the neural network model to be trained. The neural network model obtains the characteristic frequency points in the frequency domain that are associated with the original song signal part in the noisy song signal. These characteristic frequency points can also be called predicted characteristic frequency points. Multiple predicted characteristic frequency points can form the constellation diagram of the noisy song signal. After obtaining the predicted characteristic frequency points, the model parameters of the neural network model can be adjusted based on the difference value between the predicted characteristic frequency points and the reference characteristic frequency points. When the training end condition is met, a trained characteristic point recognition model can be obtained.
[0053] Specifically, in this embodiment, the reference characteristic frequency points of the original song signal can be used as the learning target of the neural network model, and a noise signal as an interference factor is introduced into the original song signal. The noisy song signal containing both the original song signal and the noise signal is input into the neural network model. Thus, during the training process, the neural network model can learn the relevant information of the noise signal in the noisy song signal and eliminate it. Based on the remaining audio signal, characteristic frequency points are screened to obtain the predicted characteristic frequency points only for the original song signal part in the noisy song signal.
[0054] In this embodiment, the noisy song audio of the original song audio can be obtained. Among them, the noisy song signal of the noisy song audio includes a noise signal and the original song signal of the original song audio. Furthermore, the reference characteristic frequency points in the frequency domain of the original song signal can be determined. The noisy song signal is input into the neural network model to be trained. The neural network model obtains the predicted characteristic frequency points associated with the original song signal in the noisy song signal in the frequency domain. Based on the difference value between the predicted characteristic frequency points and the reference characteristic frequency points, the model parameters of the neural network model are adjusted until the training end condition is met, and a trained characteristic frequency point recognition model is obtained. The solution of this embodiment can identify the predicted characteristic frequency points from the noisy song signal containing the noise signal and the original song signal through the neural network model during the training process and learn based on the reference characteristic frequency points of the original song signal. It can enable the model to eliminate the characteristic frequency points associated with the noise signal in the input song signal and only retain the characteristic frequency points associated with the original song signal in the noisy song signal, effectively eliminating the interference frequency points generated by the noise signal, increasing the reliability and accuracy of the identified characteristic frequency points, and further improving the recognition accuracy of the audio fingerprint.
[0055] In one embodiment, step S104 of adjusting the model parameters of the neural network model based on the difference value between the predicted characteristic frequency points and the reference characteristic frequency points until the training end condition is met to obtain a trained characteristic frequency point recognition model may include the following steps:
[0056] Determine the model loss value based on the difference value between the predicted feature frequency points and the reference feature frequency points; the model loss value is positively correlated with the difference value; adjust the model parameters of the neural network model according to the model loss value until the training end condition is met, and obtain the trained feature frequency point recognition model.
[0057] As an example, the difference value may include at least one of the following: the difference value of the frequency point position, the difference value of the number of frequency points. Among them, the frequency point position may include the position of the frequency point itself and / or the relative position of the frequency points. For example, if the predicted feature frequency points A and B are recognized, the position difference between the predicted feature frequency point A and the reference feature frequency point A' can be obtained, or the difference between the relative position between the predicted feature frequency points A and B and the relative position between the reference feature frequency points A' and B' can be obtained.
[0058] In practical applications, the difference value between the predicted feature frequency points and the reference feature frequency points can be obtained, and the model loss value is determined according to the difference value. The model loss value may have a negative correlation with the difference value between the predicted feature frequency points and the reference feature frequency points. In an example, when determining the model loss value, the mean squared error loss function (MSELoss) can be used to determine the model loss value.
[0059] After obtaining the model loss value, the model parameters of the neural network model can be adjusted according to the determined model loss value until the training end condition is met. For example, when the model loss value is less than the threshold or the number of iterations of the neural network model reaches the preset number, it can be determined that the neural network model has been trained well, and the current neural network model can be used as the trained feature frequency point recognition model.
[0060] In this embodiment, the smaller the difference value between the predicted feature frequency points and the reference frequency points, the smaller the model loss value of the neural network model. By determining the model loss value of the neural network model based on the difference value between the predicted feature frequency points and the reference feature frequency points, and adjusting the model parameters of the neural network model in the direction of reducing the model loss value, it can be ensured that during the model training process, the predicted feature frequency points recognized by the neural network model based on the input noisy song signal are more and more similar to the reference feature frequency points recognized based on the original song signal. Therefore, when the audio signal contains both the song signal and the interfering noise signal, the neural network model can effectively remove the noise signal and the feature frequency points of the noise signal that are irrelevant to the song signal, and correctly retain the feature frequency points of the song signal part.
[0061] In one embodiment, the feature frequency point is a local peak point, such as Figure 2 As shown, in S101, obtaining the noisy song audio of the original song audio may include the following steps:
[0062] S201. Obtain the noise signal of non-stationary noise. Among multiple frequency points in the frequency domain of the noise signal, there is at least one local peak point.
[0063] In practical applications, if the characteristic frequency points to be identified from the original song signal are local peak points, since the frequency domain energy of the noise signal of stationary noise is relatively stable, local peak points of interference usually do not appear on the spectrogram. When the audio signal includes both the noise signal of stationary noise and the song signal, the noise signal of stationary noise does not affect the terminal or server from correctly identifying the local peak points related to the song signal in the audio signal. That is, when the noise signal is the noise signal of stationary noise, using the local peak points as characteristic frequency points can have relatively excellent anti-noise ability.
[0064] However, for the noise signal of non-stationary noise, there is at least one local peak point among multiple frequency points in the frequency domain. That is, at least one frequency point of non-stationary noise can have the same frequency point characteristics as the characteristic frequency points of the song signal, resulting in some frequency points in the noise signal of non-stationary noise being possibly mis-identified as the characteristic frequency points of the song signal, thereby affecting the recognition accuracy of the characteristic frequency points and the audio fingerprint. Based on this, in this embodiment, the noise signal of non-stationary noise can be obtained.
[0065] S202. Perform a fusion process on the noise signal and the original song signal of the original song audio in the song library, and obtain the noisy song audio of the original song audio based on the fusion result.
[0066] After obtaining the noise signal of non-stationary noise, the original song audio can be obtained from the song library, and the noise signal of non-stationary noise is fused with the original song signal of the original song audio, and then the noisy song audio of the original song audio can be obtained according to the signal fusion result.
[0067] In this embodiment, by fusing the noise signal of non-stationary noise with the original song signal of the original song audio in the song library, and obtaining the noisy song audio of the original song audio based on the fusion result, data augmentation capable of adding non-stationary noise can be performed during the training of the neural network model, enabling the model to learn to identify and ignore the local peak points of non-stationary noise during the training process, effectively improving the accuracy of the identified local peak points in the case where the audio signal includes the noise signal of non-stationary noise.
[0068] In one embodiment, S201 to obtain the noise signal of non-stationary noise may include the following steps:
[0069] Obtain various types of pre-set non-stationary noises; randomly determine at least one type of non-stationary noise from the various types of non-stationary noises, and obtain the noise signals of the at least one type of non-stationary noise.
[0070] In practical applications, various different types of non-stationary noises can be obtained in advance. The various types of non-stationary noises can include at least two of the following: speech, transient noise, environmental noise, where the environmental noise can include indoor environmental noise and / or outdoor environmental noise. For example, the environmental noise in a public transportation scenario (such as the noise generated when vehicles, airplanes, high-speed trains and other means of transportation are running or starting and stopping), the environmental noise inside a shopping mall.
[0071] When it is necessary to generate a noisy song audio of the original song audio, at least one type of non-stationary noise can be randomly determined from the various types of non-stationary noises, and the noise signal of the determined non-stationary noise can be obtained. Specifically, for example, the various types of non-stationary noises can be randomly combined to obtain the noise signals of transient noise and outdoor environmental noise.
[0072] In this embodiment, various types of pre-set non-stationary noises can be obtained, at least one type of non-stationary noise can be randomly determined from the various types of non-stationary noises, and the noise signals of the at least one type of non-stationary noise can be obtained, which can increase the randomness and diversity of the noise signals of the non-stationary noises included in the noisy song signal, so that the finally trained feature frequency point recognition model can effectively eliminate the feature frequency points of different types of non-stationary noises and improve the robustness of the feature frequency point recognition model.
[0073] In one embodiment, S202 performs a fusion process on the noise signal and the original song signal of the original song audio in the song library, and obtains the noisy song audio of the original song audio based on the fusion result, which can be implemented through the following steps:
[0074] Randomly add the noise signal of the non-stationary noise to the original song signal of the original song audio to obtain the fused audio signal; obtain the noisy song audio of the original song audio based on the fused audio signal.
[0075] After obtaining the noise signal of the non-stationary noise, the noise signal of the non-stationary noise can be randomly added to the original song signal of the original song audio.
[0076] Specifically, when adding the noise signal of non-stationary noise to the original song signal, it can be processed in the time domain for the noise signal and the original song signal. During the processing, the noise signal of non-stationary noise can be randomly superimposed on the original song signal. For example, the noise signal A of non-stationary noise can be superimposed on the original song signal B to obtain the superimposed audio signal C; or, the noise signal of non-stationary noise can also be inserted into the original song signal, such as inserting the noise signal A of non-stationary noise between the original song signals B1 and B2.
[0077] After randomly adding the noise signal to the original song signal, the fused audio signal can be obtained. The fused audio signal can be an audio signal in the time domain, and then the noisy song audio of the original song audio can be obtained based on the fused audio signal.
[0078] In this embodiment, by randomly adding the noise signal of non-stationary noise to the original song signal of the original song audio to obtain the fused audio signal and obtaining the noisy song audio of the original song audio based on the fused audio signal, the randomness and diversity of the appearance of non-stationary noise in the noisy song audio can be increased, enabling the feature frequency point recognition model to more accurately identify and remove the feature frequency points of the noise signals at different positions in the audio signal.
[0079] In one embodiment, as Figure 3 shown, inputting the noisy song signal into the neural network model to be trained may include the following steps:
[0080] S301, obtaining the frame spectrum information of multiple audio frames of the noisy song audio.
[0081] In practical applications, after obtaining the noisy song audio, multiple audio frames of the noisy song audio can be obtained and time-frequency transformation processing can be performed on each audio frame to convert the audio signal of each audio frame in the time domain into an audio signal in the frequency domain, thereby obtaining the frame spectrum information of each audio frame.
[0082] S302, splicing the frame spectrum information of multiple audio frames according to the respective frame orders of the multiple audio frames to obtain the spectrum information of the noisy song audio.
[0083] As an example, the frame order can be the order of the framed audio frames in the noisy song audio.
[0084] After obtaining the frame spectrum information of multiple audio frames, the respective frame orders of the multiple audio frames can be obtained, and the frame spectrum information of the multiple audio frames can be spliced in sequence according to the frame order, and then the splicing result can be used as the spectrum information of the noisy song audio.
[0085] In a specific implementation, when splicing, multiple audio frames are spliced with the frame order of the audio frames or the time of the audio frames as the horizontal axis and the frequency as the vertical axis. For example, Figure 4 As shown, it shows the spectral information after splicing multiple audio frames.
[0086] S303. Input the spectral information of the noisy song audio into the neural network model to be trained.
[0087] As an example, the neural network model can be a residual network (resnet), such as resnet 18, resnet34, resnet 50, resnet 152, etc. Of course, other types of neural network models can also be selected according to the actual situation.
[0088] After obtaining the spectral information of the noisy song audio, it can be input into the neural network model to be trained.
[0089] In this embodiment, the frame spectral information of multiple audio frames of the noisy song audio can be obtained, and according to the frame order of each of the multiple audio frames, the frame spectral information of the multiple audio frames is spliced to obtain the spectral information of the noisy song audio, and the spectral information of the noisy song audio is input into the neural network model to be trained, so that the neural network model can screen out feature frequency points in combination with the complete spectral information of the noisy song audio, improving the screening accuracy of the feature frequency points.
[0090] The embodiment of the present application also provides an audio fingerprint recognition method, which can be applied to an application environment as shown in Figure 5 In this application environment, it includes a terminal and a server, and the terminal and the server can be communicatively connected through a network. Among them, the server can deploy a feature frequency point recognition model, the terminal can obtain the target song audio of the audio fingerprint to be recognized, and send the target audio to the server, and the server recognizes the audio fingerprint corresponding to the target song audio.
[0091] It can be understood that the above application scenario is only an example and does not constitute a limitation on the audio fingerprint recognition method provided by the embodiment of the present application. For example, the feature frequency point recognition model can also be deployed on the terminal. After the terminal obtains the target audio provided by the user, it can identify the audio fingerprint of the target audio through the audio fingerprint recognition method provided by this embodiment and send the audio fingerprint to the server for related song retrieval.
[0092] Exemplarily, the server in this embodiment can be an independent physical server, or a server cluster composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud servers, cloud databases, cloud storage, and CDN. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart watch, etc., but is not limited thereto.
[0093] In one embodiment, as Figure 6 shown, an audio fingerprint recognition method is provided. Taking the server in Figure 5 as an example, the method may include the following steps:
[0094] S601, Obtain the target song audio of the audio fingerprint to be recognized.
[0095] In specific implementation, the target song audio of the audio fingerprint to be recognized can be obtained. The audio signal of the target song audio may include not only the song signal but also the noise signal. Of course, the target song audio may not include the noise signal.
[0096] Among them, the target audio may be a humming song or a cover song recorded by the user. The user can upload the target song audio through the terminal to trigger the search for user demand songs that meet the preset conditions and are associated with the target song audio.
[0097] S602, Input the audio signal of the target song audio into the trained feature frequency point recognition model to obtain multiple feature frequency points of the audio signal of the target song audio in the frequency domain output by the feature frequency point recognition model.
[0098] After obtaining the target song audio, the audio signal of the target song audio can be input into the trained feature frequency point recognition model, and the feature frequency point recognition model outputs multiple feature frequency points of the audio signal of the target song audio in the frequency domain. Among them, the feature frequency point recognition model in this embodiment can be trained according to the training method of the feature frequency point recognition model in the above embodiment.
[0099] In specific implementation, after obtaining the target song audio, the audio signal of the target song audio in the frequency domain can be obtained. Specifically, after obtaining the target song audio, if the audio format of the target song audio is an mp3 file, it can be converted into a preset format file with a sampling rate of 8KHz, and the format-converted target song audio can be framed according to the preset byte bits. For example, it can be divided into one audio frame every 512 bytes. After obtaining multiple audio frames of the target song audio, the frame spectrum information of each audio frame can be obtained, and the frame spectrum information of multiple audio frames can be spliced according to the frame order. Furthermore, the spliced frame spectrum information can be input into the feature frequency point recognition model.
[0100] S603, Determine the audio fingerprint of the target song audio based on multiple feature frequency points.
[0101] After obtaining multiple feature frequency points of the target song audio, the audio fingerprint of the target song audio can be generated based on multiple feature frequency points.
[0102] In this embodiment, the target song audio of the audio fingerprint to be recognized can be obtained, and the audio signal of the target song audio is input into the trained feature frequency point recognition model to obtain multiple feature frequency points of the audio signal of the target song audio in the frequency domain. The feature frequency point recognition model is obtained according to the training method of the feature frequency point recognition model. Furthermore, the audio fingerprint of the target song audio can be determined based on the multiple feature frequency points. In the solution of this embodiment, since the trained feature frequency point recognition model can accurately eliminate the feature frequency points associated with the noise signal in the input audio signal and only retain the feature frequency points associated with the song signal, the interference caused by the noise signal in the audio signal is effectively eliminated during the process of extracting the feature frequency points, the accuracy of the recognized feature frequency points is improved, and thus the recognition accuracy of the audio fingerprint is improved.
[0103] In one embodiment, S603 determining the audio fingerprint of the target song audio based on multiple feature frequency points may include the following steps:
[0104] Obtain at least one set of adjacent feature frequency points from the multiple feature frequency points; determine the audio fingerprint of each set of adjacent feature frequency points based on the frequency values and sampling times of the feature frequency points in each set of adjacent feature frequency points; generate an audio fingerprint sequence based on the audio fingerprints of each set of adjacent feature frequency points as the audio fingerprint of the target song audio.
[0105] In practical applications, after obtaining multiple feature frequency points, at least one set of adjacent feature frequency points can be obtained from the multiple feature frequency points, that is, a set of adjacent feature frequency points is obtained based on two adjacent feature frequency points. Furthermore, for each set of adjacent feature frequency points, the audio fingerprint of the set of adjacent feature frequency points can be determined according to the frequency value and sampling time of each feature frequency point in the set of adjacent feature frequency points.
[0106] Specifically, the relative position of the adjacent feature frequency points can be determined based on the frequency values and sampling times of each feature frequency point in the set of adjacent feature frequency points, and the audio fingerprint of the set of adjacent feature frequency points can be determined according to the relative position. For example, if the frequency values and sampling times of a set of adjacent feature frequency points M and N are (t1, f1) and (t2, f2) respectively, then Δt between t1 and t2 can be obtained, and an audio fingerprint is generated based on t1, f1, f2, and Δt. Exemplarily, the audio fingerprint can be recorded as (t1, HashCode), where HashCode = (f1, f2, Δt).
[0107] After obtaining the audio fingerprint of each set of adjacent feature frequency points, an audio fingerprint sequence can be generated based on the audio fingerprints of the adjacent feature frequency points of the song, and the audio fingerprint sequence is used as the audio fingerprint of the target song audio.
[0108] In this embodiment, at least one set of adjacent characteristic frequency points among multiple characteristic frequency points can be obtained. Based on the frequency values and sampling times of the characteristic frequency points in each set of adjacent characteristic frequency points, the audio fingerprint of each set of the adjacent characteristic frequency points is determined. An audio fingerprint sequence is generated based on the audio fingerprints of each set of adjacent characteristic frequency points and used as the audio fingerprint of the target song audio. This can ensure the accuracy of the characteristic frequency points and the finally obtained audio fingerprint. That is, only the song signal part in the audio signal is screened for characteristic frequency points, and the influence brought by the interference noise in the target song audio is filtered out. The finally generated audio fingerprint can more accurately reflect the characteristics of the song part in the target song audio, enabling the audio fingerprint to better match the relevant original song. At the same time, since the encoding can be performed according to the relative positions of the adjacent characteristic frequency points, the audio fingerprint of the finally obtained target song audio can have strong stability.
[0109] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps does not have a strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential either, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0110] In one embodiment, a computer device is provided. This computer device can be a server, and its internal structure diagram can be as Figure 7 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the spectrum data of songs. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a training method for a characteristic frequency point recognition model and / or an audio fingerprint recognition method.
[0111] Those skilled in the art can understand, Figure 7The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0112] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:
[0113] Obtain the noisy song audio of the original song audio; the noisy song signal of the noisy song audio includes a noise signal and the original song signal of the original song audio;
[0114] Determine the reference characteristic frequency points in the frequency domain of the original song signal;
[0115] Input the noisy song signal into the neural network model to be trained, and obtain the predicted characteristic frequency points associated with the original song signal in the noisy song signal in the frequency domain through the neural network model;
[0116] Based on the difference value between the predicted characteristic frequency points and the reference characteristic frequency points, adjust the model parameters of the neural network model until the training end condition is met, and obtain the trained characteristic frequency point recognition model.
[0117] In one embodiment, when the processor executes the computer program, it also implements the steps in the above-mentioned other embodiments.
[0118] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:
[0119] Obtain the target song audio of the audio fingerprint to be recognized;
[0120] Input the audio signal of the target song audio into the trained characteristic frequency point recognition model to obtain multiple characteristic frequency points in the frequency domain of the audio signal of the target song audio output by the characteristic frequency point recognition model. The characteristic frequency point recognition model is obtained according to the training method of the characteristic frequency point recognition model described in any one of the above;
[0121] Determine the audio fingerprint of the target song audio based on the multiple characteristic frequency points.
[0122] In one embodiment, when the processor executes the computer program, it also implements the steps in the above-mentioned other embodiments.
[0123] In one embodiment, there is provided a computer program product including a computer program which, when executed by a processor, implements the following steps:
[0124] Obtain a noisy song audio of an original song audio; the noisy song signal of the noisy song audio includes a noise signal and an original song signal of the original song audio;
[0125] Determine a reference characteristic frequency point in the frequency domain of the original song signal;
[0126] Input the noisy song signal into a neural network model to be trained, and obtain, through the neural network model, a predicted characteristic frequency point associated with the original song signal in the noisy song signal in the frequency domain;
[0127] Based on the difference value between the predicted characteristic frequency point and the reference characteristic frequency point, adjust the model parameters of the neural network model until the training end condition is satisfied, and obtain a trained characteristic frequency point recognition model.
[0128] In one embodiment, when the computer program is executed by a processor, it also implements the steps in the above-mentioned other embodiments.
[0129] In one embodiment, there is provided a computer program product including a computer program which, when executed by a processor, implements the following steps:
[0130] Obtain a target song audio of an audio fingerprint to be recognized;
[0131] Input the audio signal of the target song audio into the trained characteristic frequency point recognition model, and obtain multiple characteristic frequency points in the frequency domain of the audio signal of the target song audio output by the characteristic frequency point recognition model, where the characteristic frequency point recognition model is obtained according to the training method of the characteristic frequency point recognition model described in any one of the above;
[0132] Determine the audio fingerprint of the target song audio based on the multiple characteristic frequency points.
[0133] In one embodiment, when the computer program is executed by a processor, it also implements the steps in the above-mentioned other embodiments.
[0134] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0135] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., without limitation.
[0136] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0137] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A training method for a characteristic frequency point recognition model, characterized in that The method includes: Obtaining a noisy song audio of an original song audio; the noisy song signal of the noisy song audio includes a noise signal and an original song signal of the original song audio; Determining a reference feature frequency point in the frequency domain of the original song signal; Inputting the noisy song signal into a neural network model to be trained, and obtaining, through the neural network model, a predicted feature frequency point associated with the original song signal in the noisy song signal in the frequency domain; Determining a model loss value based on a difference value between the predicted feature frequency point and the reference feature frequency point; the model loss value is positively correlated with the difference value; Adjusting model parameters of the neural network model according to the model loss value until a training end condition is satisfied, to obtain a trained feature frequency point recognition model; the training end condition includes that the model loss value is less than a threshold, or the number of iterations of the neural network model reaches a preset number of times.
2. The method according to claim 1, wherein The feature frequency point is a local peak point, and the obtaining of the noisy song audio of the original song audio includes: Obtaining a noise signal of non-stationary noise, where at least one local peak point is included among multiple frequency points in the frequency domain of the noise signal; Performing a fusion process on the noise signal and an original song signal of an original song audio in a song library, and obtaining the noisy song audio of the original song audio based on a fusion result.
3. The method according to claim 2, characterized in that The obtaining of the noise signal of non-stationary noise includes: Obtaining multiple types of pre-set non-stationary noise, where the multiple types of non-stationary noise include at least two of the following: speech, transient noise, environmental noise; Randomly determining at least one type of non-stationary noise from the multiple types of non-stationary noise, and obtaining a noise signal of the at least one type of non-stationary noise.
4. The method according to claim 2, wherein The performing of the fusion process on the noise signal and the original song signal of the original song audio in the song library, and obtaining the noisy song audio of the original song audio based on the fusion result includes: Randomly adding the noise signal of the non-stationary noise to the original song signal of the original song audio to obtain a fused audio signal; Obtaining the noisy song audio of the original song audio based on the fused audio signal.
5. The method according to claim 1, wherein The inputting of the noisy song signal into the neural network model to be trained includes: Obtaining frame spectrum information of multiple audio frames of the noisy song audio; Stitching the frame spectrum information of the multiple audio frames according to the frame order of each of the multiple audio frames to obtain spectrum information of the noisy song audio; Inputting the spectrum information of the noisy song audio into the neural network model to be trained.
6. An audio fingerprint recognition method, characterized in that, The method includes: Obtaining a target song audio of an audio fingerprint to be recognized; Inputting an audio signal of the target song audio into the trained feature frequency point recognition model, to obtain multiple feature frequency points in the frequency domain of the audio signal of the target song audio output by the feature frequency point recognition model, where the feature frequency point recognition model is obtained according to the training method of the feature frequency point recognition model according to any one of claims 1-5; Determining the audio fingerprint of the target song audio based on the multiple feature frequency points.
7. The method according to claim 6, characterized in that, Determining the audio fingerprint of the target song audio based on the multiple characteristic frequency points includes: Obtaining at least one set of adjacent characteristic frequency points among the multiple characteristic frequency points; Determining the audio fingerprint of each set of the adjacent characteristic frequency points based on the frequency values and sampling times of the characteristic frequency points in each set of the adjacent characteristic frequency points; Generating an audio fingerprint sequence based on the audio fingerprints of each set of adjacent characteristic frequency points as the audio fingerprint of the target song audio.
8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Acoustic amplification system howling point detection method based on neural network
CN111526469A
Howling detection method and device, medium and computing equipment
CN114067837A