Accompaniment audio processing method and related device
By inputting the accompaniment audio features into the twin network, obtaining the embedded mean and filtering, the problem of difficulty in evaluating the accompaniment audio quality in the prior art is solved, and the rapid and accurate screening of the accompaniment audio quality is achieved, and the user's music entertainment experience is improved.
Patent Information
- Application Number
- CN202211282228.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-19
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2042-10-19
AI Technical Summary
The prior art is difficult to effectively screen or evaluate the accompaniment audio quality, resulting in a lack of accurate information when selecting accompaniment.
The embedded mean of the accompaniment is obtained by obtaining the audio characteristics of the accompaniment and inputting it into the twin network. The quality category of the accompaniment is determined based on the difference between the embedded mean and the preset value, so that it can be screened or evaluated.
It realizes rapid and accurate screening or evaluation of the accompaniment audio quality, helping users choose high-quality accompaniment and improving the music and entertainment experience.
Smart Images

Figure CN115641876B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to an accompaniment audio processing method and related devices. Background Art
[0002] Currently, there are many methods for screening or evaluating high-quality works of songs. For example, in the five-dimensional scoring method, the breath, rhythm, intonation, skills, and emotions of a song are analyzed to determine whether the song is a high-quality work. However, these screening or evaluation methods are all for analyzing the singing situation of the singer, rather than for analyzing the quality of the accompaniment.
[0003] However, currently, there are more and more music entertainment methods in which users sing by selecting accompaniments. Therefore, how to screen or evaluate accompaniment audio is an urgent problem to be solved. Summary of the Invention
[0004] Embodiments of this application provide an accompaniment audio processing method, device, terminal, and storage medium, which can screen or evaluate accompaniment audio.
[0005] In a first aspect, embodiments of this application provide an accompaniment audio processing method, which includes:
[0006] Obtain the audio features of the accompaniment;
[0007] Input the audio features into a siamese network to obtain the embedding mean of the accompaniment;
[0008] When the difference between the embedding mean of the accompaniment and the first embedding mean is less than or equal to a first preset value, determine that the accompaniment is an accompaniment of the first category;
[0009] The siamese network is trained based on the contrast loss of N sample pairs, and the contrast loss of the N sample pairs is obtained based on the audio features of each sample in each sample pair; the N sample pairs include sample pairs composed of samples of the first category and samples of the second category, and sample pairs composed of samples of the first category;
[0010] The accompaniment quality of the accompaniment or sample of the first category is better than that of the accompaniment or sample of the second category;
[0011] The first embedding mean is the average of the embedding values of each sample in the sample pairs composed of samples of the first category, and the embedding value of each sample is obtained by inputting the audio features of each sample into the siamese network.
[0012] In an optional implementation manner, after determining that the accompaniment is an accompaniment of the first category, the method further includes:
[0013] If the difference between the embedded mean of the accompaniment and the first embedded mean is less than the preset threshold, add the accompaniment to the accompaniment library.
[0014] In an alternative embodiment, the method further includes:
[0015] When the difference between the embedded mean of the accompaniment and the second embedded mean is less than or equal to the second preset value, determine that the accompaniment is of the second category;
[0016] The second embedded mean is the average of the embedded values of the samples in the sample pairs composed of the samples of the second category, and the embedded values of the samples are obtained by inputting the audio features of the samples into the siamese network.
[0017] In an alternative embodiment, after determining that the accompaniment is of the second category, the method further includes:
[0018] If the difference between the embedded mean of the accompaniment and the second embedded mean is less than the preset threshold, delete the accompaniment from the accompaniment library where it is located.
[0019] In an alternative embodiment, obtaining the audio features of the accompaniment includes:
[0020] Obtain the Mel spectrogram features of the accompaniment;
[0021] Input the Mel spectrogram features of the accompaniment into the feature extraction network to obtain the target features of the accompaniment, where the target features are features independent of the quality of the accompaniment;
[0022] Perform feature concatenation on the Mel spectrogram features and the target features of the accompaniment to obtain the audio features of the accompaniment.
[0023] In an alternative embodiment, inputting the Mel spectrogram features of the accompaniment into the feature extraction network to obtain the target features of the accompaniment includes:
[0024] Input the Mel spectrogram features of the accompaniment into the first processing module included in the feature extraction network to obtain the first features; the first features are obtained by concatenating the output features of multiple convolutional layers with different scales and the frequency pooling layers corresponding to each convolutional layer;
[0025] Input the first features into the second processing module included in the feature extraction network to obtain the second features;
[0026] Input the second features into the third processing module included in the feature extraction network to obtain the target features of the accompaniment; the third processing module is used to reduce the dimension of the second features.
[0027] In an alternative embodiment, the method further includes:
[0028] Input the audio features of each sample in each of the N sample pairs into the siamese network to obtain the embedding values of each sample in each sample pair;
[0029] Use the embedding values of each sample in each sample pair to determine the contrast loss value of the N sample pairs;
[0030] When the contrast loss value of the N sample pairs does not meet the stop training condition, update the siamese network according to the contrast loss value of the N sample pairs, and execute again the operation of inputting the audio features of each sample in each of the N sample pairs into the siamese network to obtain the embedding values of each sample in each sample pair until the contrast loss value of the N sample pairs meets the stop training condition.
[0031] In an alternative embodiment, using the embedding values of each sample in each sample pair to determine the contrast loss value of the N sample pairs includes:
[0032] Determine the distance between the embedding values of each sample in each sample pair;
[0033] Based on the distance between the embedding values of each sample in each sample pair, determine the first contrast loss value of the N sample pairs.
[0034] In an alternative embodiment, the method further includes:
[0035] Input the audio features of each sample in the sample pairs composed of samples of the first category and the audio features of each sample in the sample pairs composed of samples of the second category into the trained siamese network respectively to obtain the first embedding mean of the first category and the second embedding mean of the second category.
[0036] In a second aspect, an embodiment of the present application provides an accompaniment audio processing device, the device includes:
[0037] An acquisition unit, configured to acquire the audio features of the accompaniment;
[0038] The acquisition unit is further configured to input the audio features into the siamese network to obtain the embedding mean of the accompaniment;
[0039] A determination unit, configured to determine that the accompaniment is an accompaniment of the first category when the difference between the embedding mean of the accompaniment and the first embedding mean is less than or equal to a first preset value;
[0040] Wherein, the siamese network is trained based on the contrast loss of the N sample pairs, and the contrast loss of the N sample pairs is obtained based on the audio features of each sample in each sample pair; the N sample pairs include sample pairs composed of samples of the first category and samples of the second category, and sample pairs composed of samples of the first category;
[0041] The accompaniment quality of the accompaniments or samples of the first category is better than that of the accompaniments or samples of the second category;
[0042] The first embedding mean is the average value among the embedding values of each sample in the sample pairs composed of the samples of the first category, and the embedding value of each sample is obtained by inputting the audio features of each sample into the siamese network.
[0043] In a third aspect, an embodiment of the present application provides a computer device, including: a processor, a communication interface, and a memory. The processor, the communication interface, and the memory are interconnected. Wherein, the memory stores executable program code, and the processor is configured to call the executable program code to execute the method provided by the embodiment of the present application.
[0044] In a fourth aspect, the present application further provides a computer-readable storage medium, in which instructions are stored. When the instructions are run on a computer, the computer is caused to execute the method provided by the embodiment of the present application.
[0045] In a fifth aspect, an embodiment of the present application further provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method provided by the embodiment of the present application
[0046] It can be seen that by adopting the embodiment of the present application, based on the samples of the first category (i.e., high-quality accompaniment samples) and the trained siamese network, the first embedding mean (i.e., the embedding mean of high-quality accompaniments) can be obtained, so that the accompaniment audio can be screened or evaluated according to the embedding mean of the accompaniment and the first embedding mean. Description of the Drawings
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0048] Figure 1 is a schematic diagram of the process of uploading an original accompaniment provided by an embodiment of the present application;
[0049] Figure 2 is a schematic diagram of the process of an accompaniment audio processing method provided by an embodiment of the present application;
[0050] Figure 3It is a schematic diagram of a feature extraction network provided by an embodiment of the present application;
[0051] Figure 4a It is a schematic diagram of a preprocessing module in the feature extraction network provided by an embodiment of the present application;
[0052] Figure 4b It is a schematic diagram of a middle processing module in the feature extraction network provided by an embodiment of the present application;
[0053] Figure 4c It is a schematic diagram of a postprocessing module in the feature extraction network provided by an embodiment of the present application;
[0054] Figure 5 It is a schematic flow diagram of another method for processing accompaniment audio provided by an embodiment of the present application;
[0055] Figure 6 It is a schematic diagram of a method for processing accompaniment audio provided by an embodiment of the present application;
[0056] Figure 7a It is a schematic structural diagram of a siamese network provided by an embodiment of the present application;
[0057] Figure 7b It is a schematic flow diagram of a method for training a siamese network provided by an embodiment of the present application;
[0058] Figure 8 It is a schematic diagram of an apparatus for processing accompaniment audio provided by an embodiment of the present application;
[0059] Figure 9 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0060] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0061] To facilitate the understanding of the embodiments disclosed in the present application, some concepts related to the embodiments of the present application will be first elaborated. The elaboration of these concepts includes but is not limited to the following content.
[0062] 1. Samples of the first category and samples of the second category
[0063] In the present application, the samples of the first category and the samples of the second category are both samples in each sample pair of the N sample pairs used when training the siamese network.
[0064] Among them, the samples of the first category can also be called positive samples or high-quality samples, including but not limited to the song accompaniments with high click rates and high user praise rates in the accompaniment library, as well as the high-quality original accompaniments selected by professionals through artificial evaluation of the original accompaniments.
[0065] The samples of the second category can also be called negative samples or low-quality samples, including but not limited to the song accompaniments with low click rates and high user negative review rates in the accompaniment library, as well as the low-quality original accompaniments selected by professionals through artificial evaluation of the original accompaniments.
[0066] Among them, the original accompaniments in the original accompaniment library are the accompaniments independently uploaded by users. Optionally, users can upload the original accompaniments through the process as Figure 1 shown.
[0067] 2. The first embedding mean and the second embedding mean
[0068] The first embedding mean can also be called the embedding mean of high-quality accompaniments or the embedding mean of high-quality accompaniments. It is obtained by inputting the audio features of each sample in the sample pairs composed of the samples of the first category into the siamese network.
[0069] The second embedding mean can also be called the embedding mean of low-quality accompaniments or the embedding mean of low-quality accompaniments. It is obtained by inputting the audio features of each sample in the sample pairs composed of the samples of the second category into the siamese network.
[0070] Among them, the embedding mean refers to the average value between the output values of the two fully connected layers of the siamese network.
[0071] 3. The accompaniments of the first category and the accompaniments of the second category
[0072] In this application, the accompaniments of the first category and the accompaniments of the second category are both the results obtained by evaluating the accompaniment audio. Among them, the quality of the accompaniments of the first category is better than that of the accompaniments of the second category. The accompaniments of the first category can also be called high-quality accompaniments or high-quality accompaniments. The accompaniments of the second category can also be called low-quality accompaniments or low-quality accompaniments.
[0073] Currently, there are many methods for screening or evaluating high-quality works of songs. For example, the five-dimensional scoring method and the sound quality analysis method, etc. These two methods both use signal processing methods to analyze the internal connections of high-quality works of audio signals. However, due to the numerous factors affecting the quality of accompaniments and many factors being unable to be quantified (such as rhythm, instrument matching, etc.), it is impossible to analyze the accompaniment audio using signal processing methods only from the audio signal itself to complete the screening or evaluation of the accompaniment audio.
[0074] Based on this, the embodiments of the present application provide an accompaniment audio processing method and related devices to screen or evaluate the accompaniment audio. Among them, in this method, the computer device can obtain the audio features of the accompaniment; input the audio features into the siamese network to obtain the embedding mean of the accompaniment; when the difference between the embedding mean of the accompaniment and the first embedding mean is less than or equal to a preset value, determine that the accompaniment is a first-category accompaniment; among them, the siamese network is trained based on the contrast loss of N sample pairs, and the contrast loss of N sample pairs is obtained based on the audio features of each sample in each sample pair; the N sample pairs include sample pairs composed of samples of the first category and samples of the second category, and sample pairs composed of samples of the first category; the accompaniment quality of the first-category accompaniment or sample is better than that of the second-category accompaniment or sample; the first embedding mean is the average value between the embedding values of each sample in the sample pairs composed of samples of the first category, and the embedding values of each sample are obtained by inputting the audio features of each sample into the siamese network. It can be seen that by adopting the embodiments of the present application, the accompaniment audio can be screened or evaluated.
[0075] It should be noted that the above-mentioned computer device can be a terminal device or a server. The terminal devices mentioned here include but are not limited to smart phones, tablet computers, laptop computers, desktop computers, smart speakers, smart watches, in-vehicle infotainment systems, etc. The server mentioned here can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, etc., but is not limited thereto.
[0076] To facilitate the understanding of the embodiments of the present application, the following will elaborate in detail on the specific implementation manner of the above-mentioned accompaniment audio processing method with the computer device as the execution subject.
[0077] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of an accompaniment audio processing method provided by the embodiments of the present application. As Figure 2 shown, the accompaniment audio processing method may include the following steps:
[0078] S201. Obtain the audio features of the accompaniment.
[0079] In an optional implementation manner, for the computer device to obtain the audio features of the accompaniment, it may include: obtaining the Mel spectrogram features of the accompaniment; inputting the Mel spectrogram features of the accompaniment into the feature extraction network to obtain the target features of the accompaniment, where the target features are features irrelevant to the quality of the accompaniment; performing feature concatenation on the Mel spectrogram features and the target features of the accompaniment to obtain the audio features of the accompaniment. Optionally, the target features include but are not limited to genre features, and the genre features are, for example, the features of classical music, electronic music, etc.
[0080] In this embodiment, the computer device obtains the Mel-spectrum features of the accompaniment, including: performing frame division and windowing processing on the audio signal of the accompaniment to obtain multiple processed audio segments; determining the Mel-spectrum features of each audio segment among the multiple audio segments; and obtaining the Mel-spectrum features of the accompaniment according to the Mel-spectrum features of each audio segment.
[0081] Optionally, when the computer device performs frame division and windowing processing on the audio signal of the accompaniment to obtain multiple processed audio segments, it may include: dividing the audio signal with a sampling rate of a preset frequency into audio segments of a preset length; and performing windowing processing on each audio segment according to preset parameters to obtain multiple processed audio segments. For example, when the computer device performs frame division and windowing processing on the audio signal of the accompaniment, it may divide the audio signal with a sampling rate of AkHz into audio segments with a length of Bs, and add a window length (window length) C to each audio segment, where the window shift is D (i.e., the overlap between adjacent audio segments is D). For example, the sampling rate AkHz may be 16kHz, the length of the audio segment Bs may be 3s, the window length C may be 512, and the window shift D may be 256.
[0082] Optionally, when the computer device determines the Mel-spectrum features of each audio segment among the multiple audio segments, it may determine the Mel-spectrum features of each audio segment according to a preset Mel-spectrum dimension. For example, assuming that the preset Mel-spectrum dimension is 96, the length of each audio segment is 3s, the sampling rate is 16kHz, the window length is 512, and the window shift is 256, then the Mel-spectrum features of each audio segment are ((3×16000) / (512 - 256))×96 = 187×96.
[0083] In this embodiment, the computer device inputs the Mel-spectrum features of the accompaniment into a feature extraction network to obtain the target features of the accompaniment, including: inputting the Mel-spectrum features of the accompaniment into a first processing module included in the feature extraction network to obtain a first feature; the first feature is obtained by concatenating the output features of multiple convolutional layers with different scales and the frequency pooling layers corresponding to each convolutional layer; inputting the first feature into a second processing module included in the feature extraction network to obtain a second feature; inputting the second feature into a third processing module included in the feature extraction network to obtain the target features; the third processing module is used to reduce the dimension of the second feature. Optionally, the Mel-spectrum features of the accompaniment are obtained according to the Mel-spectrum features of each audio segment included in the accompaniment. Optionally, the computer device may input the Mel-spectrum features of each audio segment included in the accompaniment into the feature extraction network to obtain the target features of each audio segment; and obtain the target features of the accompaniment according to the target features of each audio segment.
[0084] Please refer to Figure 3 , Figure 3 which is a schematic diagram of a feature extraction network provided by an embodiment of the present application. AsFigure 3 As shown, the computer device can input the Mel-spectrum features of each audio segment, which are 187×96 dimensions, i.e., 3sec×96 bands (corresponding to Figure 2 201 in Figure 2 ), into the preprocessing module of the feature extraction network (i.e., the aforementioned first processing module, corresponding to Figure 2 202 in Figure 2 ), and obtain the front-end features of 187×561 dimensions (i.e., the aforementioned first feature, corresponding to Figure 2 203 in Figure 2 ); input the front-end features 203 into the middle processing module of the feature extraction network (i.e., the aforementioned second processing module, corresponding to Figure 2 204 in
[0085] ), and obtain the middle-end features of 187×753 dimensions (i.e., the aforementioned second feature, corresponding to Figure 4a ), where the middle processing module includes 64 filters; input the middle-end features 205 into the postprocessing module of the feature extraction network (i.e., the aforementioned third processing module, corresponding to Figure 4a 206 in Figure 2 ), and obtain the target features of 1×50 (corresponding to Figure 4a 207 in Figure 2 ).
[0086] Among them, the first convolutional layer and the second convolutional layer are both used to learn the features in the time dimension of the Mel spectrogram features, and the third convolutional layer, the fourth convolutional layer, and the fifth convolutional layer are all used to learn the features in the frequency domain dimension of the Mel spectrogram features. Among them, the convolutional scale of the first convolutional layer is (7, 38)×204, and its output feature is 187×59×204 dimensions. After this output feature undergoes frequency pooling processing, a feature of 187×204 dimensions can be obtained; the convolutional scale of the second convolutional layer is (7, 67)×204, and its output feature is 187×30×204 dimensions. After this output feature undergoes frequency pooling processing, a feature of 187×204 dimensions can be obtained; the convolutional scale of the third convolutional layer is (32, 1)×51, and its output feature is 187×96×51 dimensions. After this output feature undergoes frequency pooling processing, a feature of 187×51 dimensions can be obtained; the convolutional scale of the fourth convolutional layer is (64, 1)×51, and its output feature is 187×96×51 dimensions. After this output feature undergoes frequency pooling processing, a feature of 187×51 dimensions can be obtained; the convolutional scale of the fifth convolutional layer is (128, 1)×51, and its output feature is 187×96×51 dimensions. After this output feature undergoes frequency pooling processing, a feature of 187×51 dimensions can be obtained. The computer device splices the finally obtained five-way features to obtain a first feature of 187×(204 + 204 + 51 + 51 + 51) = 187×561 dimensions.
[0087] Please refer to Figure 4b , Figure 4b is a schematic diagram of the middle processing module in the feature extraction network provided by the embodiment of the present application, corresponding to Figure 2 204 in Figure 4b As shown in Figure 3 , the middle processing module is composed of a one-dimensional convolutional layer and a residual module, and is mainly used to perform further abstraction processing on the first feature. The computer device can input the 187×561-dimensional first feature obtained based on Figure 3 into this middle processing module, and use a one-dimensional convolutional layer with a convolutional kernel size of 7×7 and 64 convolutional kernels and a residual convolution to process the 187×561-dimensional first feature. Among them, the output feature of each one-dimensional convolutional layer is 187×64 dimensions. The features output by each one-dimensional convolutional layer and the residual module are spliced to obtain a second feature of 187×(561 + 64 + 64 + 64) = 187×753 dimensions, and the second feature is used as the Figure 2 input of the post-processing module 206 in
[0088] Please refer to Figure 4c , Figure 4c is a schematic diagram of the post-processing module in the feature extraction network provided by the embodiment of the present application, corresponding to Figure 2 206 in Figure 4cAs shown in the figure, the post - processing module is mainly composed of pooling layers, including average pooling, max pooling, etc. The post - processing module is mainly used to perform pooling and compression processing on the second feature output by the middle - processing module to obtain a target feature, which is a lower - dimensional feature independent of the accompaniment quality. Figure 5 It can be seen that the 187×753 - dimensional second feature is processed by average global pooling to obtain a 1×753 - dimensional feature after average pooling, and the 187×753 - dimensional second feature is processed by max global pooling to obtain a 1×753 - dimensional feature after max pooling; the 1×753 - dimensional feature after average pooling and the 1×753 - dimensional feature after max pooling are combined to obtain a 1×1506 - dimensional feature after global pooling; the 1×1506 - dimensional feature after global pooling is processed by a deep neural network including 200 neurons to obtain a 1×200 - dimensional feature; the 1×200 - dimensional feature is processed by a deep neural network including 50 neurons to obtain a 1×50 - dimensional target feature.
[0089] In this embodiment, the computer device splices the Mel - spectrum feature of the accompaniment and the target feature to obtain the audio feature of the accompaniment, which may include: expanding the Mel - spectrum feature of each audio segment of the accompaniment to obtain the expanded feature of each audio segment; splicing the expanded feature of each audio segment and the target feature corresponding to this audio segment to obtain the spliced feature; using the spliced feature as the audio feature of the accompaniment. For example, assuming that the Mel - spectrum feature of an audio segment is 187×96 - dimensional and the target feature corresponding to this audio segment is 1×50 - dimensional, the computer device can expand the Mel - spectrum feature of this audio segment from 187×96 - dimensional to 1×17952 - dimensional; splicing the expanded feature with the target feature to obtain the spliced feature, that is, 1×18002 - dimensional; using the 1×18002 - dimensional feature as the audio feature of the accompaniment. In this way, the siamese network can eliminate the features irrelevant to the accompaniment quality during the process of screening or evaluating the accompaniment, thus avoiding the influence of features irrelevant to the accompaniment quality on the judgment of the siamese network.
[0090] S202: Input the audio feature into the siamese network to obtain the embedding mean of the accompaniment; the siamese network is trained based on the contrast loss of N sample pairs, and the contrast loss of N sample pairs is obtained based on the audio features of each sample in each sample pair; N sample pairs include sample pairs composed of samples of the first category and samples of the second category, and sample pairs composed of samples of the first category.
[0091] Among them, the accompaniment quality of the accompaniment or sample of the first category is better than that of the accompaniment or sample of the second category.
[0092] In an alternative embodiment, inputting the audio features into the Siamese network to obtain the embedding mean of the accompaniment may include: inputting the audio features of each audio segment included in the accompaniment into the Siamese network to obtain the embedding mean of each audio segment; and obtaining the embedding mean of the accompaniment according to the embedding means of each audio segment.
[0093] It should be noted that since the two networks in the Siamese network are the same and share parameters, for each audio segment, the embedding values output by the two networks in the Siamese network are the same. Thus, the embedding mean of each audio segment can be equivalent to the embedding value output by any one of the networks in the Siamese network.
[0094] S203. When the difference between the embedding mean of the accompaniment and the first embedding mean is less than or equal to the first preset value, determine that the accompaniment is an accompaniment of the first category; the first embedding mean is the average of the embedding values of the samples in the sample pairs composed of the samples of the first category, and the embedding values of the samples are obtained by inputting the audio features of the samples into the Siamese network.
[0095] In an alternative embodiment, the computer device may also determine that the accompaniment is an accompaniment of the second category when the difference between the embedding mean of the accompaniment and the second embedding mean is less than or equal to the second preset value; the second embedding mean is the average of the embedding values of the samples in the sample pairs composed of the samples of the second category, and the embedding values of the samples are obtained by inputting the audio features of the samples into the Siamese network.
[0096] In the embodiments of the present application, the computer device obtains the audio features of the accompaniment; inputs the audio features into the Siamese network to obtain the embedding mean of the accompaniment; and determines that the accompaniment is an accompaniment of the first category (i.e., high-quality accompaniment) when the difference between the embedding mean of the accompaniment and the first embedding mean is less than or equal to the first preset value. It can be seen that by adopting the embodiments of the present application, based on the Siamese network, the accompaniment audio can be screened or evaluated.
[0097] Please refer to Figure 5 , Figure 5 which is a schematic flowchart of another method for processing accompaniment audio provided by the embodiments of the present application. Compared with Figure 1 the method shown, this method further includes, after determining that the accompaniment is an accompaniment of the first category, adding the accompaniment to the accompaniment library if the difference between the embedding mean of the accompaniment and the first embedding mean is less than the preset threshold; and deleting the accompaniment from the accompaniment library where it is located if the difference between the embedding mean of the accompaniment and the second embedding mean is less than the preset threshold after determining that the accompaniment is an accompaniment of the second category. As Figure 2 shown, the method for processing accompaniment audio may include the following steps:
[0098] S501. Obtain the audio features of the accompaniment.
[0099] S502. Input the audio features into the Siamese network to obtain the embedding mean of the accompaniment. The Siamese network is trained based on the contrast loss of N sample pairs, and the contrast loss of the N sample pairs is obtained based on the audio features of each sample in each sample pair. The N sample pairs include sample pairs composed of samples of the first category and samples of the second category, and sample pairs composed of samples of the first category.
[0100] Among them, the accompaniment quality of the accompaniment or sample of the first category is better than that of the accompaniment or sample of the second category.
[0101] In an optional implementation manner, the specific implementations of steps S501 and S502 can be respectively referred to the descriptions in the above steps S501 and S502, and will not be elaborated here.
[0102] S503a. When the difference between the embedding mean of the accompaniment and the first embedding mean is less than or equal to the first preset value, determine that the accompaniment is the accompaniment of the first category. The first embedding mean is the average value between the embedding values of each sample in the sample pair composed of samples of the first category, and the embedding values of each sample are obtained by inputting the audio features of each sample into the Siamese network.
[0103] S504a. If the difference between the embedding mean of the accompaniment and the first embedding mean is less than the preset threshold, add the accompaniment to the accompaniment library.
[0104] It can be understood that step S504a can be applied to the situation where the accompaniment does not exist in the accompaniment library. For example, the accompaniment is an original accompaniment uploaded by the user independently.
[0105] S503b. When the difference between the embedding mean of the accompaniment and the second embedding mean is less than or equal to the second preset value, determine that the accompaniment is the accompaniment of the second category. The second embedding mean is the average value between the embedding values of each sample in the sample pair composed of samples of the second category, and the embedding values of each sample are obtained by inputting the audio features of each sample into the Siamese network.
[0106] S504b. If the difference between the embedding mean of the accompaniment and the second embedding mean is less than the preset threshold, delete the accompaniment from the accompaniment library where it is located.
[0107] It can be understood that step S504b can be applied to the situation where the accompaniment exists in the accompaniment library.
[0108] In the embodiments of the present application, after the computer device determines that the accompaniment is of the first category, if the accompaniment does not exist in the accompaniment library and the difference between the embedding mean value of the accompaniment and the first embedding mean value is less than a preset threshold, the accompaniment is added to the accompaniment library; after determining that the accompaniment is of the second category, if the accompaniment exists in the accompaniment library and the difference between the embedding mean value of the accompaniment and the second embedding mean value is less than a preset threshold, the accompaniment is deleted from the accompaniment library where it is located. It can be seen that by adopting the embodiments of the present application, based on the siamese network, high-quality accompaniments can be quickly and efficiently screened out from a large number of original accompaniments uploaded by users, and low-quality accompaniments can be screened out from the accompaniment library, so as to replace the low-quality accompaniments in the accompaniment library with the high-quality accompaniments in the original accompaniments, thereby optimizing the accompaniment library and further improving the singing experience of users. In addition, compared with the method of manually screening or evaluating the quality of accompaniments, this method saves time and effort, can avoid the problem of inconsistent judgments caused by personal subjective reasons, and the prediction results are consistent and fair.
[0109] Please refer to Figure 6 , Figure 6 which is a schematic diagram of an accompaniment audio processing method provided by an embodiment of the present application. As Figure 6 shown, the computer device divides the accompaniment audio into audio segments with a length of 3 s; extracts the Mel spectrogram features of each audio segment; inputs the Mel spectrogram features of each audio segment into the feature extraction network to obtain the target features. For example, the Mel spectrogram features of each audio segment are input into the genre classification network to obtain the genre features of each audio segment. The target features and Mel spectrogram features of each audio segment are feature concatenated, and the concatenated features are input into the siamese network to obtain the embedding mean value of each audio segment; based on the embedding values of each audio segment, the embedding mean value of the accompaniment is obtained; based on the embedding mean value of the accompaniment, the first embedding mean value, and the second embedding mean value, the category to which the accompaniment belongs is determined. Among them, when the difference between the embedding mean value of the accompaniment and the first embedding mean value is less than or equal to the first preset value, it is determined that the accompaniment is of the first category (high-quality accompaniment); when the difference between the embedding mean value of the accompaniment and the second embedding mean value is less than or equal to the second preset value, it is determined that the accompaniment is of the second category (low-quality accompaniment). If the accompaniment is of the first category, when the accompaniment does not exist in the accompaniment library and the difference between the embedding mean value of the accompaniment and the first embedding mean value is less than the preset threshold, the accompaniment is added to the accompaniment library; if the accompaniment is of the second category, when the accompaniment exists in the accompaniment library and the difference between the embedding mean value of the accompaniment and the second embedding mean value is less than the preset threshold, the accompaniment is deleted from the accompaniment library where it is located.
[0110] Please refer to Figure 7a , Figure 7a which is a schematic diagram of the structure of a siamese network provided by an embodiment of the present application. As Figure 7aAs shown, the Siamese network includes two convolutional branches. Among them, the two convolutional modules and the two fully connected layers have the same structure and share parameters. The contrast loss value can be calculated based on the embedding values output by the two fully connected layers. When training the Siamese network, the parameters in the convolutional modules and fully connected layers in the Siamese network can be adjusted based on the contrast loss value to obtain the trained Siamese network. The training method of the Siamese network will be elaborated in detail below. Please refer to Figure 7b , Figure 7b which is a schematic flowchart of a method for training a Siamese network provided by an embodiment of the present application. The trained Siamese network can correspond to Figure 1 , Figure 5 and Figure 6 the Siamese network mentioned in
[0111] S701. Obtain training samples, where the training samples include N sample pairs.
[0112] Among them, the N sample pairs include sample pairs composed of samples of the first category and samples of the second category, sample pairs composed of samples of the first category, and sample pairs composed of samples of the second category.
[0113] Optionally, the samples of the first category can also be called positive samples or high-quality samples, which include the accompaniments of songs with high click-through rates and high user praise rates in the accompaniment library, and the high-quality original accompaniments selected by professionals after manually evaluating the original accompaniments in the original accompaniment library.
[0114] Optionally, the samples of the second category can also be called negative samples or low-quality samples. It includes the accompaniments of songs with low click-through rates and high user negative review rates in the accompaniment library, and the low-quality original accompaniments selected by professionals after manually evaluating the original accompaniments in the original accompaniment library.
[0115] S702. Obtain the audio features of each sample in each of the N sample pairs.
[0116] In an optional implementation manner, for the computer device to obtain the audio features of each sample in each of the N sample pairs, it may include: obtaining the Mel spectrogram features of each sample in each of the N sample pairs; inputting the Mel spectrogram features of each sample in each sample pair into a feature extraction network to obtain the target features of each sample in each sample pair, where the target features are features irrelevant to the quality of the accompaniment; and performing feature concatenation on the target features of each sample in each sample pair and the Mel spectrogram features of the sample to obtain the audio features of each sample in each sample pair. In this way, the Siamese network can be made to learn to eliminate features irrelevant to the quality of the accompaniment (such as genre features) during training, thereby avoiding the influence of features irrelevant to the quality of the accompaniment on the judgment of high-quality and low-quality accompaniments by the Siamese network.
[0117] Optionally, for the specific implementation process of this embodiment, reference may be made to the relevant description in the foregoing step S101, and details are not described herein again.
[0118] S703. Input the audio features of each sample in each of the N sample pairs into the siamese network to obtain the embedding values of each sample in each sample pair.
[0119] In an alternative embodiment, when the computer device inputs the audio features of each sample in each of the N sample pairs into the siamese network, it may input multiple audio segments of each sample in each sample pair into the siamese network. In this way, based on the embedding values of each audio segment of each sample in each sample pair output by the siamese network, the embedding mean value of each sample in each sample pair can be obtained; and the embedding mean value of each sample in each sample pair is used as the embedding value of each sample in each sample pair.
[0120] S704. Use the embedding values of each sample in each sample pair to determine the contrast loss value of the N sample pairs.
[0121] In an alternative embodiment, when the computer device uses the embedding values of each sample in each sample pair to determine the contrast loss value of the N sample pairs, it may include: determining the distance between the embedding values of each sample in each sample pair; and determining the contrast loss value of the N sample pairs based on the distance between the embedding values of each sample in each sample pair.
[0122] Optionally, in this embodiment, the distance between the embedding values of each sample in each sample pair may be the Euclidean distance, and the computer device may determine the Euclidean distance between the embedding values of each sample in each sample pair through the following formula (1).
[0123]
[0124] In the above formula (1), n represents the nth sample pair, x n represents the embedding value of the first sample in the nth sample pair; x' n represents the embedding value of the second sample in the nth sample pair; d n represents the Euclidean distance between the embedding values of each sample in the nth sample pair.
[0125] Optionally, in this embodiment, the computer device may determine the contrast loss value of the N sample pairs according to the following formula.
[0126]
[0127] In the above formula (2), N represents the total number of sample pairs for which the loss is currently calculated; n represents the nth sample pair; L represents the contrast loss value of the N sample pairs; d nIt represents the Euclidean distance between the embedding values of each sample in the nth sample pair; y represents whether the labels of each sample in the nth sample pair match. Among them, when each sample in the nth sample pair is a sample of the first category (that is, both samples are high-quality samples), the value of y is 1. When one sample in the nth sample pair is a sample of the first category and the other sample is a sample of the second category (that is, one of the two samples is a high-quality sample and the other is a low-quality sample), the value of y is 0; margin represents an adjustable threshold parameter.
[0128] It can be understood that in the above formula (2), when y = 1, the calculation formula of the contrast loss value is as follows formula (3); when y = 0, the calculation formula of the contrast loss value is as follows formula (4).
[0129]
[0130]
[0131] In the above formulas (3) and (4), the physical meanings of the parameters can be referred to the descriptions of the physical meanings of the parameters in formula (2), and will not be elaborated here.
[0132] It can be seen from formula (3) that when each sample in the sample pair belongs to the same category of samples, the greater the Euclidean distance between the two samples in the feature space, the greater the contrast loss value, indicating that the current network is not good. Therefore, it is necessary to optimize the network parameters of the Siamese network. It can be seen from formula (4) that when each sample in the sample pair belongs to different categories of samples, the smaller the Euclidean distance between the two samples in the feature space, the greater the contrast loss value, indicating that the current network is not good. This loss function can well express the matching degree of the sample pair and can also be well used to train the model for feature extraction.
[0133] It can be understood that during the training of the Siamese network, if each sample in the sample pair belongs to the same category of samples, it is required that the Euclidean distance between the embedding values of the two samples output by the Siamese network is as small as possible; if each sample in the sample pair belongs to different categories of samples, it is required that the Euclidean distance between the embedding values of the two samples output by the Siamese network is as large as possible. In this way, the contrast loss value can be made smaller and smaller, so that the performance of the Siamese network is getting better and better.
[0134] S705. When the contrast loss value of the N sample pairs does not meet the stop training condition, update the Siamese network according to the contrast loss value of the N sample pairs, and then execute step S703 again until the contrast loss value of the N sample pairs meets the stop training condition.
[0135] Optionally, the stop training condition can be that the contrast loss value of the N sample pairs is less than or equal to a preset contrast loss value.
[0136] That is to say, the process of training the Siamese network is essentially a process of continuously updating the Siamese network based on the contrast loss values of N sample pairs. In an alternative embodiment, when the computer device updates the Siamese network according to the contrast loss values of N sample pairs, it may update the network parameters of the Siamese network according to the contrast loss values of N sample pairs.
[0137] For example, assuming that the Siamese network mentioned in step S703 refers to the initialized Siamese network, if the audio features of each sample in each of the N sample pairs are input into the initialized Siamese network and the initial contrast loss values of the N sample pairs determined do not meet the stop training condition, the initialized Siamese network can be updated according to the initial contrast loss values of the N sample pairs to obtain the first Siamese network. The audio features of each sample in each of the N sample pairs are input into the first Siamese network to determine whether the first contrast loss values of the N sample pairs meet the stop training condition. If not, the first Siamese network is updated according to the first contrast loss values of the N sample pairs to obtain the second Siamese network. Then, the audio features of each sample in each of the N sample pairs are input into the second Siamese network again to determine whether the second contrast loss values of the N sample pairs meet the stop training condition. If so, the training is stopped and the trained Siamese network is obtained.
[0138] In an alternative embodiment, after the Siamese network training is completed, the computer device may also input the audio features of each sample in the sample pairs composed of the samples of the first category and the audio features of each sample in the sample pairs composed of the samples of the second category into the trained Siamese network to obtain the first embedding mean of the first category and the second embedding mean of the second category.
[0139] In this embodiment, the computer device may splice the Mel spectrogram features of each sample in the sample pair with the target features of each sample to obtain the spliced features of each sample in the sample pair; and use the spliced features of each sample in the sample pair as the audio features of each sample in the sample pair. Optionally, the computer device also obtains the Mel spectrogram features and target features of each sample in the sample pair. Optionally, the specific process of the computer device obtaining the Mel spectrogram features and target features of each sample in the sample pair can refer to the relevant description in the foregoing step S101 and will not be elaborated here.
[0140] Among them, the first embedding mean is the average value between the embedding values of each sample in multiple sample pairs composed of the samples of the first category; and the second embedding mean is the average value between the embedding values of each sample in multiple sample pairs composed of the samples of the second category.
[0141] It can be seen that in the embodiments of the present application, by training the siamese network, the siamese network can automatically learn the commonalities among high-quality accompaniments (i.e., the Euclidean distance between the sample pairs composed of high-quality accompaniments is less than the first distance threshold), and the differences between high-quality accompaniments and low-quality accompaniments (i.e., the Euclidean distance between the sample pairs composed of high-quality accompaniments and low-quality accompaniments is greater than the second distance threshold). Therefore, by using the trained siamese network, the accompaniments can be screened or evaluated quickly and efficiently.
[0142] Please refer to Figure 8 , Figure 8 which is a schematic diagram of an accompaniment audio processing device provided by an embodiment of the present application. The accompaniment audio processing device described in this embodiment may include the following parts:
[0143] An acquisition unit 801, configured to acquire the audio features of the accompaniment;
[0144] The acquisition unit 801 is further configured to input the audio features into the siamese network to obtain the embedding mean of the accompaniment;
[0145] A determination unit 802, configured to determine that the accompaniment is an accompaniment of the first category when the difference between the embedding mean of the accompaniment and the first embedding mean is less than or equal to a first preset value;
[0146] Wherein, the siamese network is trained based on the contrast loss of N sample pairs, and the contrast loss of the N sample pairs is obtained based on the audio features of each sample in each sample pair; the N sample pairs include sample pairs composed of samples of the first category and samples of the second category, and sample pairs composed of samples of the first category;
[0147] The accompaniment quality of the accompaniment or sample of the first category is better than that of the accompaniment or sample of the second category;
[0148] The first embedding mean is the average value between the embedding values of each sample in the sample pairs composed of samples of the first category, and the embedding value of each sample is obtained by inputting the audio features of each sample into the siamese network.
[0149] In an alternative embodiment, the accompaniment audio processing device may further include a processing unit 803.
[0150] In an alternative embodiment, the processing unit 803 is configured to, after the determination unit 802 determines that the accompaniment is an accompaniment of the first category, add the accompaniment to the accompaniment library if the difference between the embedding mean of the accompaniment and the first embedding mean is less than a preset threshold.
[0151] In an alternative embodiment, the determination unit 802 is further configured to determine that the accompaniment is an accompaniment of the second category when the difference between the embedding mean of the accompaniment and the second embedding mean is less than or equal to a second preset value;
[0152] The second embedding mean is the average value between the embedding values of each sample in the sample pairs composed of samples of the second category, and the embedding value of each sample is obtained by inputting the audio features of each sample into the siamese network.
[0153] In an optional implementation manner, after the determining unit 802 determines that the accompaniment is an accompaniment of the second category, the processing unit 803 is further configured to delete the accompaniment from the accompaniment library where it is located if the difference between the embedding mean of the accompaniment and the second embedding mean is less than a preset threshold.
[0154] In an optional implementation manner, when the obtaining unit 801 is used to obtain the audio features of the accompaniment, it is specifically configured to:
[0155] Obtain the Mel spectrogram features of the accompaniment;
[0156] Input the Mel spectrogram features of the accompaniment into the feature extraction network to obtain the target features of the accompaniment, where the target features are features independent of the quality of the accompaniment;
[0157] Perform feature splicing on the Mel spectrogram features and the target features of the accompaniment to obtain the audio features of the accompaniment.
[0158] In an optional implementation manner, when the obtaining unit 801 is used to input the Mel spectrogram features of the accompaniment into the feature extraction network to obtain the target features of the accompaniment, it is specifically configured to:
[0159] Input the Mel spectrogram features of the accompaniment into the first processing module included in the feature extraction network to obtain the first features; the first features are obtained by splicing the output features of multiple convolutional layers with different scales and the frequency pooling layers corresponding to each convolutional layer;
[0160] Input the first features into the second processing module included in the feature extraction network to obtain the second features;
[0161] Input the second features into the third processing module included in the feature extraction network to obtain the target features of the accompaniment; the third processing module is used to reduce the dimension of the second features.
[0162] In an optional implementation manner, the accompaniment audio processing device may further include a training unit 804.
[0163] In an optional implementation manner, the training unit 804 is used to:
[0164] Input the audio features of each sample in each of the N sample pairs into the siamese network to obtain the embedding values of each sample in each sample pair;
[0165] Use the embedding values of each sample in each sample pair to determine the contrast loss values of the N sample pairs;
[0166] When the contrast loss values of N sample pairs do not meet the stop training condition, update the siamese network according to the contrast loss values of the N sample pairs, and perform again the operation of inputting the audio features of each sample in each sample pair of the N sample pairs into the siamese network to obtain the embedding values of each sample in each sample pair until the contrast loss values of the N sample pairs meet the stop training condition.
[0167] In an alternative embodiment, when the training unit 804 is used to determine the contrast loss values of N sample pairs by using the embedding values of each sample in each sample pair, it is specifically used for:
[0168] Determine the distance between the embedding values of each sample in each sample pair;
[0169] Based on the distance between the embedding values of each sample in each sample pair, determine the contrast loss values of the N sample pairs.
[0170] In an alternative embodiment, the training unit 804 is further used to input the audio features of each sample in the sample pairs composed of samples of the first category and the audio features of each sample in the sample pairs composed of samples of the second category into the trained siamese network respectively to obtain a first embedding mean of the first category and a second embedding mean of the second category.
[0171] It can be understood that the specific implementation of each unit in the accompaniment audio processing device described in the embodiments of the present application and the beneficial effects that can be achieved can refer to the description of the foregoing related embodiments, which will not be repeated here.
[0172] Please refer to Figure 9 , Figure 9 which is a schematic structural diagram of a computer device shown in the embodiments of the present application. The computer device described in the embodiments of the present application includes: a processor 901, a user interface 902, a communication interface 903, and a memory 904. Among them, the processor 901, the user interface 902, the communication interface 903, and the memory 904 can be connected through a bus or other means. In the embodiments of the present application, the connection through a bus is taken as an example.
[0173] Among them, the processor 901 (or CPU (Central Processing Unit)) is the computing core and control core of the computer device. It can parse various instructions in the computer device and process various data of the computer device. For example, the CPU can be used to parse the power-on and power-off instructions sent by the user to the computer device and control the computer device to perform power-on and power-off operations. Another example is that the CPU can transfer various interactive data between the internal structures of the computer device, and so on. The user interface 902 is a medium for realizing the interaction and information exchange between the user and the computer device. Its specific manifestations can include a display screen (Display) for output and a keyboard (Keyboard) for input, etc. It should be noted that the keyboard here can be a physical keyboard, a touch screen virtual keyboard, or a keyboard combining a physical and a touch screen virtual keyboard. The communication interface 903 can optionally include standard wired interfaces and wireless interfaces (such as Wi-Fi, mobile communication interfaces, etc.), and is controlled by the processor 901 to receive and send data. The memory 904 (Memory) is a memory device in the computer device, used to store programs and data. It can be understood that the memory 904 here can include both the built-in memory of the computer device and, of course, the extended memory supported by the computer device. The memory 904 provides a storage space, and this storage space stores the operating system of the computer device, which can include but is not limited to: Android system, iOS system, Windows Phone system, etc. This application does not make any limitations in this regard.
[0174] In the embodiment of this application, the processor 901 is used to extract the audio features of the target audio data, identify the audio categories of rare animals by using a few-shot model, etc., and perform the related operations of the above-mentioned audio recognition method.
[0175] The memory 904 is also used to store the received data and the data generated during the processing process.
[0176] In the embodiment of this application, the processor 901 performs the following operations by running the executable program code in the memory 904:
[0177] Obtain the audio features of the accompaniment;
[0178] Input the audio features into the siamese network to obtain the embedding mean of the accompaniment;
[0179] When the difference between the embedding mean of the accompaniment and the first embedding mean is less than or equal to the first preset value, determine that the accompaniment is the accompaniment of the first category;
[0180] Among them, the Siamese network is trained based on the contrast loss of N sample pairs, and the contrast loss of N sample pairs is obtained based on the audio features of each sample in each sample pair; the N sample pairs include sample pairs composed of samples of the first category and samples of the second category, and sample pairs composed of samples of the first category;
[0181] The accompaniment quality of the accompaniment or samples of the first category is better than that of the accompaniment or samples of the second category;
[0182] The first embedding mean is the average value among the embedding values of each sample in the sample pairs composed of samples of the first category, and the embedding value of each sample is obtained by inputting the audio features of each sample into the Siamese network.
[0183] In an optional implementation manner, after determining that the accompaniment is an accompaniment of the first category, the processor 901 is further configured to add the accompaniment to the accompaniment library if the difference between the embedding mean of the accompaniment and the first embedding mean is less than a preset threshold.
[0184] In an optional implementation manner, the processor 901 is further configured to determine that the accompaniment is an accompaniment of the second category when the difference between the embedding mean of the accompaniment and the second embedding mean is less than or equal to a second preset value;
[0185] The second embedding mean is the average value among the embedding values of each sample in the sample pairs composed of samples of the second category, and the embedding value of each sample is obtained by inputting the audio features of each sample into the Siamese network.
[0186] In an optional implementation manner, after determining that the accompaniment is an accompaniment of the second category, the processor 901 is further configured to delete the accompaniment from the accompaniment library where it is located if the difference between the embedding mean of the accompaniment and the second embedding mean is less than a preset threshold.
[0187] In an optional implementation manner, when the processor 901 is used to obtain the audio features of the accompaniment, it is specifically configured to:
[0188] Obtain the Mel spectrogram features of the accompaniment;
[0189] Input the Mel spectrogram features of the accompaniment into the feature extraction network to obtain the target features of the accompaniment, and the target features are features independent of the quality of the accompaniment;
[0190] Perform feature splicing on the Mel spectrogram features and the target features of the accompaniment to obtain the audio features of the accompaniment.
[0191] In an optional implementation manner, when the processor 901 is used to input the Mel spectrogram features of the accompaniment into the feature extraction network to obtain the target features of the accompaniment, it is specifically configured to:
[0192] Input the Mel-spectrum features of the accompaniment into the first processing module included in the feature extraction network to obtain first features; the first features are obtained by concatenating the output features of multiple convolutional layers with different scales and the frequency pooling layers corresponding to each convolutional layer.
[0193] Input the first features into the second processing module included in the feature extraction network to obtain second features.
[0194] Input the second features into the third processing module included in the feature extraction network to obtain the target features of the accompaniment; the third processing module is used to reduce the dimension of the second features.
[0195] In an alternative embodiment, the processor 901 is further configured to:
[0196] Input the audio features of each sample in each of the N sample pairs into the initialized siamese network to obtain the embedding values of each sample in each sample pair.
[0197] Use the embedding values of each sample in each sample pair to determine the contrast loss value of the N sample pairs.
[0198] When the contrast loss value of the N sample pairs does not meet the stop training condition, update the siamese network according to the contrast loss value of the N sample pairs, and execute again the operation of inputting the audio features of each sample in each of the N sample pairs into the siamese network to obtain the embedding values of each sample in each sample pair, until the contrast loss value of the N sample pairs meets the stop training condition.
[0199] In an alternative embodiment, when the processor 901 is configured to use the embedding values of each sample in each sample pair to determine the contrast loss value of the N sample pairs, it is specifically configured to:
[0200] Determine the distance between the embedding values of each sample in each sample pair.
[0201] Based on the distance between the embedding values of each sample in each sample pair, determine the contrast loss value of the N sample pairs.
[0202] In an alternative embodiment, the processor 901 is further configured to input the audio features of each sample in the sample pairs composed of samples of the first category and the audio features of each sample in the sample pairs composed of samples of the second category into the trained siamese network respectively, to obtain the first embedding mean of the first category and the second embedding mean of the second category.
[0203] In a specific implementation, the processor 901, user interface 902, communication interface 903, and memory 904 described in the embodiments of the present application can implement the implementation manners of the computer device described in the accompaniment audio processing method provided in the embodiments of the present application, and can also implement the implementation manners described in the accompaniment audio processing device provided in the embodiments of the present application, which will not be elaborated herein.
[0204] The embodiments of the present application further provide a computer-readable storage medium storing a computer program, where the computer program includes program instructions, and when the program instructions are executed by a processor, the accompaniment audio processing method provided in the embodiments of the present application is implemented. For the specific implementation manners, reference can be made to the implementation manners provided in the above steps, which will not be elaborated herein.
[0205] The embodiments of the present application further provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method as described in the embodiments of the present application. For the specific implementation manners, reference can be made to the foregoing description, which will not be elaborated herein.
[0206] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, some steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0207] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium, and the storage medium can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0208] The foregoing disclosure is only a part of the embodiments of the present application. Of course, the scope of the rights of the present application cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.
Claims
1. A method for processing accompaniment audio, characterized in that: The method comprises: Get the audio features of the accompaniment; Inputting the audio feature into the twin network to obtain the embedding mean of the accompaniment; wherein the embedding mean refers to the average value between the output values of the two fully connected layers of the twin network; When the difference between the embedding mean of the accompaniment and the first embedding mean is less than or equal to a first preset value, determining that the accompaniment is an accompaniment of the first category; The twin network is obtained by training based on the contrast loss of N sample pairs, and the contrast loss of the N sample pairs is obtained based on the audio features of each sample in each sample pair; the N sample pairs include sample pairs consisting of samples of the first category and samples of the second category, and sample pairs consisting of samples of the first category; The accompaniment quality of the accompaniment or sample of the first category is better than the accompaniment quality of the accompaniment or sample of the second category; The first embedding mean is the average value between the embedding values of each sample in the sample pair consisting of the samples of the first category, and the embedding value of each sample is obtained by inputting the audio features of each sample into the twin network.
2. The method according to claim 1, characterized in that: After determining that the accompaniment is an accompaniment of the first category, the method further includes: If the difference between the embedding mean of the accompaniment and the first embedding mean is smaller than a preset threshold, the accompaniment is added to the accompaniment library.
3. The method according to claim 1, characterized in that The method further comprises: When the difference between the embedding mean of the accompaniment and the second embedding mean is less than or equal to a second preset value, determining that the accompaniment is an accompaniment of the second category; The second embedding mean is the average value between the embedding values of each sample in the sample pair consisting of the samples of the second category, and the embedding value of each sample is obtained by inputting the audio features of each sample into the twin network.
4. The method according to claim 3, characterized in that After determining that the accompaniment is an accompaniment of the second category, the method further includes: If the difference between the embedding mean value of the accompaniment and the second embedding mean value is smaller than a preset threshold, the accompaniment is deleted from the accompaniment library where it is located.
5. The method according to claim 1, characterized in that The step of obtaining the audio features of the accompaniment includes: Obtain the Mel-spectrum features of the accompaniment; Inputting the Mel spectrum feature of the accompaniment into a feature extraction network to obtain a target feature of the accompaniment, wherein the target feature is a feature that is irrelevant to the quality of the accompaniment; The mel spectrum feature of the accompaniment and the target feature are concatenated to obtain the audio feature of the accompaniment.
6. The method according to claim 5, characterized in that The step of inputting the Mel spectrum feature of the accompaniment into a feature extraction network to obtain the target feature of the accompaniment includes: Inputting the Mel spectrum feature of the accompaniment into the first processing module included in the feature extraction network to obtain a first feature; the first feature is obtained by splicing the output features of multiple convolutional layers of different scales and the frequency pooling layer corresponding to each convolutional layer; Inputting the first feature into a second processing module included in the feature extraction network to obtain a second feature; The second feature is input into a third processing module included in the feature extraction network to obtain a target feature of the accompaniment; the third processing module is used to reduce the dimension of the second feature.
7. The method according to claim 1, characterized in that The method further comprises: Input the audio features of each sample in each sample pair of N sample pairs into the twin network to obtain the embedding value of each sample in each sample pair; Determine the contrast loss values of the N sample pairs using the embedding value of each sample in each sample pair; When the contrast loss values of the N sample pairs do not meet the condition for stopping training, the twin network is updated according to the contrast loss values of the N sample pairs, and the operation of inputting the audio features of each sample in each sample pair of the N sample pairs into the twin network to obtain the embedding value of each sample in each sample pair is performed again until the contrast loss values of the N sample pairs meet the condition for stopping training.
8. The method according to claim 7, characterized in that The determining the contrast loss values of the N sample pairs by using the embedding value of each sample in each sample pair includes: Determining the distance between the embedding values of each sample in each sample pair; Based on the distance between the embedding values of each sample in each sample pair, the contrast loss values of the N sample pairs are determined.
9. The method according to claim 7, characterized in that: The method further comprises: The audio features of each sample in the sample pair consisting of the samples of the first category and the audio features of each sample in the sample pair consisting of the samples of the second category are respectively input into the trained twin network to obtain a first embedding mean of the first category and a second embedding mean of the second category.
10. A computer device, characterized in that: include: A processor, a communication interface and a memory, wherein the processor, the communication interface and the memory are connected to each other, wherein the memory stores an executable program code, and the processor is used to call the executable program code to execute the method as described in any one of claims 1-9.
11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor is caused to perform the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Accompaniment audio recommendation method and device and computer readable storage medium
CN108182227A
Accompaniment classification method and apparatus
US20220277040A1