Intelligent transspectral model training methods, transspectral methods, equipment and media

CN117558261BActive Publication Date: 2026-08-14TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-15
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0003]现有的智能转谱算法直接通过频谱特征预测音高和音符边界信息,这种方式对于音符边界点的识别存在较大的优化空间

Benefits of technology

[0044]可见,本申请提供了一种智能转谱模型训练方法,包括:从样本音频中提取第一干声音频,并从所述第一干声音频中提取第一频谱特征,将所述第一频谱特征输入至预先训练完的目标语音识别模型,以通过所述目标语音识别模型输出所述第一干声音频的第一语义特征;若所述第一频谱特征的第一特征维度与所述第一语义特征的第二特征维度不一致,则将所述第一特征维度和所述第二特征维度调整至相同的第一目标特征维度;将所述第一频谱特征和所述第一语义特征输入至待训练智能转谱模型,以通过所述待训练智能转谱模型输出与若干个目标维度对应的若干个概率矩阵,并基于与若干个所述目标维度对应的所述若干个概率矩阵计算训练损失;当所述训练损失不满足预设损失条件时,返回至所述将所述第一频谱特征和所述第一语义特征输入至待训练智能转谱模型的步骤,直至得到所述训练损失满足所述预设损失条件的目标智能转谱模型。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117558261B_ABST
    Figure CN117558261B_ABST
Patent Text Reader

Abstract

This application discloses a training method, transcribing method, device, and medium for an intelligent transcribing model, relating to the field of intelligent transcribing technology. The method includes: inputting the first spectral feature of a first dry audio signal from a sample audio sample into a target speech recognition model to obtain the first semantic feature of the first dry audio signal; adjusting the first spectral feature and the first semantic feature to the same feature dimension; inputting the first spectral feature and the first semantic feature into the intelligent transcribing model to be trained to obtain several probability matrices, and calculating the training loss based on the several probability matrices; when the training loss does not meet the loss condition, returning to the step of inputting the first spectral feature and the first semantic feature into the intelligent transcribing model to be trained to obtain a target intelligent transcribing model where the training loss meets the loss condition. This application, by fusing the semantic and spectral information of the audio, enables the target intelligent transcribing model to better identify the boundary points of notes, thereby improving the accuracy of intelligent transcribing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent spectrum transcoding technology, and in particular to an intelligent spectrum transcoding model training method, transcoding method, device and medium. Background Technology

[0002] Intelligent music transcription technology refers to the process of separating vocals from a song and identifying the pitch and note boundaries within the voice. This technology has wide applications, especially in the current era of AIGC (Artificial Intelligence Generated Content), where it can provide sheet music for vocal synthesis, reducing the cost of manual transcription. Furthermore, intelligent music transcription technology can also provide a basis for pitch scoring in karaoke games.

[0003] Existing intelligent transcribing algorithms directly predict pitch and note boundary information through spectral features, but there is considerable room for improvement in the identification of note boundary points.

[0004] Therefore, there is an urgent need for an intelligent transcribing method to improve the accuracy of note boundary point identification during the intelligent transcribing process. Summary of the Invention

[0005] In view of this, the purpose of this invention is to provide an intelligent spectrum transposition model training method, transposition method, device, and medium, which can improve the accuracy of intelligent spectrum transposition. The specific solution is as follows:

[0006] Firstly, this application discloses an intelligent transspectral model training method, including:

[0007] Extract a first dry audio from the sample audio, extract a first spectral feature from the first dry audio, and input the first spectral feature into a pre-trained target speech recognition model so as to output the first semantic feature of the first dry audio through the target speech recognition model.

[0008] If the first feature dimension of the first spectral feature is inconsistent with the second feature dimension of the first semantic feature, then the first feature dimension and the second feature dimension are adjusted to the same first target feature dimension.

[0009] The first spectral feature and the first semantic feature are input into the intelligent transspectral model to be trained, so that the intelligent transspectral model to be trained outputs several probability matrices corresponding to several target dimensions, and the training loss is calculated based on the several probability matrices corresponding to several target dimensions.

[0010] When the training loss does not meet the preset loss condition, return to the step of inputting the first spectral feature and the first semantic feature into the intelligent transspectral model to be trained, until the target intelligent transspectral model whose training loss meets the preset loss condition is obtained.

[0011] Optionally, the step of inputting the first spectral feature and the first semantic feature into the intelligent transspectral model to be trained, so as to output several probability matrices corresponding to several target dimensions through the intelligent transspectral model to be trained, includes:

[0012] The first spectral feature and the first semantic feature are input into the intelligent transspectral model to be trained, so that the convolutional recurrent neural network in the intelligent transspectral model to be trained outputs a first output feature corresponding to the first spectral feature and a second output feature corresponding to the first semantic feature, respectively.

[0013] The first output feature and the second output feature are concatenated to obtain the concatenated feature, and the concatenated feature is input into the fully connected layer in the intelligent transspectral model to be trained, so that the fully connected layer outputs several probability matrices corresponding to several target dimensions.

[0014] Optionally, the step of inputting the first spectral feature and the first semantic feature into the intelligent transspectral model to be trained, so as to output several probability matrices corresponding to several target dimensions through the intelligent transspectral model to be trained, includes:

[0015] The first spectral feature and the first semantic feature are input into the intelligent transspectral model to be trained, so that the intelligent transspectral model to be trained outputs four probability matrices corresponding to the four dimensions;

[0016] The four probability matrices each include the classification results of several frames of the first dry audio in the four dimensions. The classification result in the first dimension is a binary classification result of whether each frame of the first dry audio is a note start point or note end point. The classification result in the second dimension is a binary classification result of whether each frame of the first dry audio contains human voice. The classification result in the third dimension is a five-class classification result of human voice octave. The classification result in the fourth dimension is a twelve-class classification result of human voice pitch.

[0017] Optionally, the step of inputting the first spectral feature and the first semantic feature into the intelligent transspectral model to be trained, so as to output several probability matrices corresponding to several target dimensions through the intelligent transspectral model to be trained, includes:

[0018] The first spectral feature and the first semantic feature are input into the intelligent transspectral model to be trained, so that the intelligent transspectral model to be trained outputs three probability matrices corresponding to the three dimensions;

[0019] The three probability matrices each include classification results of several frames of the first dry audio in the three dimensions. The classification result in the first dimension is a binary classification result of whether each frame of the first dry audio is a note start point or note end point. The classification result in the second dimension is a six-class classification result of human voice octave. The classification result in the third dimension is a thirteen-class classification result of human voice pitch. One of the six-class classification results and one of the thirteen-class classification results is a classification result of whether each frame of the first dry audio contains human voice.

[0020] Optionally, the first spectral feature may include Mel spectral features or CQT spectral features.

[0021] Secondly, this application discloses a spectrum transfer method, which includes a target intelligent spectrum transfer model trained based on the aforementioned intelligent spectrum transfer model training method, comprising:

[0022] Extract a second dry audio from the audio to be transcribed, and extract a second spectral feature from the second dry audio. Input the second spectral feature into a pre-trained target speech recognition model, so as to output the second semantic feature of the second dry audio through the target speech recognition model.

[0023] If the third feature dimension of the second spectral feature is inconsistent with the fourth feature dimension of the second semantic feature, then the third feature dimension and the fourth feature dimension are adjusted to the same second target feature dimension.

[0024] The second spectral feature and the second semantic feature are input into the target intelligent transspectral model, so as to complete the intelligent transspectral transposition of the audio to be transposed through the probability matrices corresponding to several target dimensions output by the target intelligent transspectral model.

[0025] Optionally, the step of inputting the second spectral feature and the second semantic feature into the target intelligent transspectral model, so as to complete the intelligent transspectral transposition of the audio to be transspectralized through the probability matrices corresponding to several target dimensions output by the target intelligent transspectral model, includes:

[0026] The second spectral feature and the second semantic feature are input into the target intelligent transspectral model, so as to determine the probability that the second dry audio in each frame is a note start point or note end point through the first probability matrix output by the target intelligent transspectral model corresponding to the first dimension;

[0027] The second dry audio of the first target frame that conforms to the first target discrimination rule is determined as the note start point or the note end point;

[0028] The first target discrimination rule includes: the probability that the second dry audio in the first few frames and the second dry audio in the last few frames of the first target frame is the starting point or the ending point of the note is less than the probability that the first target frame is the starting point or the ending point of the note, and the probability that the first target frame is the starting point or the ending point of the note is greater than a first preset probability.

[0029] Optionally, the step of inputting the second spectral feature and the second semantic feature into the target intelligent transspectral model, so as to complete the intelligent transspectral transposition of the audio to be transspectralized through the probability matrices corresponding to several target dimensions output by the target intelligent transspectral model, includes:

[0030] The second spectral feature and the second semantic feature are input into the target intelligent transcribing model, so as to determine the probability that the second dry audio in each frame between the note start point and the note end point contains human voice through the second probability matrix corresponding to the second dimension output by the target intelligent transcribing model;

[0031] The second target frame and the second dry audio that meet the second target discrimination rule are determined as the second dry audio containing human voice; wherein, the second target discrimination rule includes: the probability that the second dry audio of the second target frame contains human voice is greater than a second preset probability.

[0032] Thirdly, this application discloses an intelligent transspectral model training device, comprising:

[0033] The sample audio feature extraction module is used to extract a first dry audio from the sample audio, extract a first spectral feature from the first dry audio, and input the first spectral feature into a pre-trained target speech recognition model so as to output the first semantic feature of the first dry audio through the target speech recognition model.

[0034] The sample audio feature dimension adjustment module is used to adjust the first feature dimension and the second feature dimension to the same first target feature dimension if the first feature dimension of the first spectral feature is inconsistent with the second feature dimension of the first semantic feature.

[0035] The model training module is used to input the first spectral feature and the first semantic feature into the intelligent transspectral model to be trained, so as to output several probability matrices corresponding to several target dimensions through the intelligent transspectral model to be trained, and to calculate the training loss based on the several probability matrices corresponding to several target dimensions; when the training loss does not meet the preset loss condition, the module returns to the step of inputting the first spectral feature and the first semantic feature into the intelligent transspectral model to be trained, until a target intelligent transspectral model is obtained in which the training loss meets the preset loss condition.

[0036] Fourthly, this application discloses a spectrum transfer device, which includes a target intelligent spectrum transfer model trained based on the aforementioned intelligent spectrum transfer model training method, comprising:

[0037] The audio feature extraction module for transcribing audio is used to extract a second dry audio from the audio to be transcribing audio, and to extract a second spectral feature from the second dry audio. The second spectral feature is then input into a pre-trained target speech recognition model so that the target speech recognition model can output a second semantic feature of the second dry audio.

[0038] The audio feature dimension adjustment module to be converted to spectrum is used to adjust the third feature dimension and the fourth feature dimension to the same second target feature dimension if the third feature dimension of the second spectral feature is inconsistent with the fourth feature dimension of the second semantic feature.

[0039] The intelligent transspectral module is used to input the second spectral feature and the second semantic feature into the target intelligent transspectral model, so as to complete the intelligent transspectral transposition of the audio to be transposed through the probability matrices corresponding to several target dimensions output by the target intelligent transspectral model.

[0040] Fifthly, this application discloses an electronic device, comprising:

[0041] Memory, used to store computer programs;

[0042] A processor for executing the computer program to implement the aforementioned disclosed method.

[0043] Sixthly, this application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned disclosed method.

[0044] As can be seen, this application provides a method for training an intelligent transspectral model, comprising: extracting a first dry audio from sample audio, extracting a first spectral feature from the first dry audio, inputting the first spectral feature into a pre-trained target speech recognition model, so as to output a first semantic feature of the first dry audio through the target speech recognition model; if the first feature dimension of the first spectral feature is inconsistent with the second feature dimension of the first semantic feature, then adjusting the first feature dimension and the second feature dimension to the same first target feature dimension; inputting the first spectral feature and the first semantic feature into the intelligent transspectral model to be trained, so as to output a plurality of probability matrices corresponding to a plurality of target dimensions through the intelligent transspectral model to be trained, and calculating a training loss based on the plurality of probability matrices corresponding to the plurality of target dimensions; when the training loss does not meet a preset loss condition, returning to the step of inputting the first spectral feature and the first semantic feature into the intelligent transspectral model to be trained, until a target intelligent transspectral model is obtained in which the training loss meets the preset loss condition.

[0045] The beneficial effects of this application are as follows: This application integrates first semantic features and first spectral features during intelligent music transcription. Since note boundary points are often pitch or word transition points, the model in this application can fully utilize the first semantic features and first spectral features to predict note boundary information, thereby improving the accuracy of intelligent music transcription. Furthermore, to prevent inconsistencies between the first feature dimension of the first spectral features and the second feature dimension of the first semantic features, which could prevent the first semantic features and first spectral features from being learned equally by the model, this application adjusts the first spectral features and first semantic features to the same feature dimension when their feature dimensions are mismatched. This allows the first spectral features and first semantic features to be learned equally by the model, further improving the accuracy of intelligent music transcription. Attached Figure Description

[0046] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0047] Figure 1 This is a flowchart of a traditional intelligent spectrum conversion algorithm;

[0048] Figure 2 This is a flowchart of a training method for an intelligent transspectral model disclosed in this application;

[0049] Figure 3 This is a flowchart of a specific intelligent transspectral model training method disclosed in this application;

[0050] Figure 4 This is a flowchart of a spectrum transduction method disclosed in this application;

[0051] Figure 5 This is a schematic diagram of a decoding method disclosed in this application;

[0052] Figure 6 This is a schematic diagram of the structure of an intelligent spectrum conversion model training device disclosed in this application;

[0053] Figure 7 This is a schematic diagram of the structure of a spectrum conversion device disclosed in this application;

[0054] Figure 8 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation

[0055] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0056] Existing intelligent transcribing algorithms directly predict pitch and note boundary information through spectral features, see [link to relevant documentation]. Figure 1 As shown, it can be seen that traditional transcribing methods do not fully utilize the semantic information of lyrics, resulting in inaccurate identification of note boundary points. Therefore, there is still considerable room for improvement in the current identification of note boundary points.

[0057] Therefore, this application proposes an intelligent spectrum transfer model training scheme to improve the accuracy of intelligent spectrum transfer.

[0058] This application discloses an intelligent transspectral model training method. (See also...) Figure 2 As shown, the method includes:

[0059] Step S11: Extract the first dry audio from the sample audio, extract the first spectral feature from the first dry audio, and input the first spectral feature into the pre-trained target speech recognition model so as to output the first semantic feature of the first dry audio through the target speech recognition model.

[0060] In this embodiment, the sample audio can be sample audio downloaded from the network or sample audio acquired in real time by the microphone of the terminal device; no specific limitation is made here.

[0061] Furthermore, a first dry audio signal is extracted from the sample audio, and a first spectral feature is extracted from the first dry audio signal. The first spectral feature includes, but is not limited to, Mel spectral features and CQT (Constant Q Transformation) spectral features.

[0062] After extracting the first spectral feature, the first spectral feature is input into a pre-trained target speech recognition model to output the first semantic feature of the first dry audio file. In one specific embodiment, the target speech recognition model can be the Whisper speech recognition model, which is a speech recognition model trained on a large amount of data in more than 120 languages, and therefore can support multilingual semantic information.

[0063] Step S12: If the first feature dimension of the first spectral feature is inconsistent with the second feature dimension of the first semantic feature, then adjust the first feature dimension and the second feature dimension to the same first target feature dimension.

[0064] It should be noted that, in order to prevent the first feature dimension of the first spectral feature from being inconsistent with the second feature dimension of the first semantic feature, which would result in the first semantic feature and the first spectral feature not being able to be learned equally by the model, this embodiment adjusts the first spectral feature and the first semantic feature to the same first target feature dimension when the feature dimensions of the first spectral feature and the first semantic feature do not match. In this way, the first semantic feature and the first spectral feature can be learned equally by the model.

[0065] For example, assuming the first spectral feature is 39-dimensional and the first semantic feature is 1024-dimensional, this embodiment uses PCA (Principal Component Analysis) dimensionality reduction to map the 1024-dimensional first semantic feature to 39-dimensional. PCA dimensionality reduction can remove redundant information in the first semantic feature and retain useful information.

[0066] Step S13: Input the first spectral feature and the first semantic feature into the intelligent transspectral model to be trained, so as to output several probability matrices corresponding to several target dimensions through the intelligent transspectral model to be trained, and calculate the training loss based on the several probability matrices corresponding to several target dimensions.

[0067] In this embodiment, after the first spectral feature and the first semantic feature are input into the intelligent transspectral model to be trained, the convolutional recurrent neural network in the intelligent transspectral model to be trained outputs a first output feature corresponding to the first spectral feature and a second output feature corresponding to the first semantic feature, respectively. Further, the first output feature and the second output feature are concatenated to obtain a concatenated feature, and the concatenated feature is input into the fully connected layer in the intelligent transspectral model to be trained. Then, the fully connected layer outputs several probability matrices corresponding to several target dimensions.

[0068] In other words, the intelligent transspectral model to be trained in this embodiment is a dual-input CRNN network (Convolutional Recurrent Neural Network). One input is the first spectral feature, and the output is the first output feature. The other input is the first semantic feature, and the output is the second output feature. Furthermore, the output features of the two networks are concatenated to obtain concatenated features, and the concatenated features are input into the fully connected layer in the intelligent transspectral model to be trained, so that the fully connected layer outputs several probability matrices corresponding to several target dimensions.

[0069] In a first specific implementation, the first spectral feature and the first semantic feature are input into the intelligent transcribing model to be trained, so that the intelligent transcribing model to be trained outputs four probability matrices corresponding to four dimensions; wherein, the four probability matrices respectively include the classification results of several frames of the first dry audio in the four dimensions, the classification result in the first dimension is a binary classification result of whether each frame of the first dry audio is a note start point or note end point, the classification result in the second dimension is a binary classification result of whether each frame of the first dry audio contains human voice, the classification result in the third dimension is a five-class classification result of human voice octave, and the classification result in the fourth dimension is a twelve-class classification result of human voice pitch.

[0070] In summary, this embodiment achieves intelligent musical notation through four-dimensional prediction output. Specifically, this embodiment determines the starting and ending points of a note using a first-dimensional probability matrix, determines whether a human voice is present between the starting and ending points using a second-dimensional probability matrix, and if a human voice is present, determines the octave value of the human voice using a third-dimensional probability matrix, and the pitch of the human voice using a fourth-dimensional probability matrix. Conversely, if a human voice is not present, the starting and ending points of the note are discarded.

[0071] In a second specific implementation, the first spectral feature and the first semantic feature are input into the intelligent transcribing model to be trained, so that the intelligent transcribing model to be trained outputs three probability matrices corresponding to the three dimensions; wherein, the three probability matrices respectively include the classification results of several frames of the first dry audio in the three dimensions, the classification result in the first dimension is a binary classification result of whether each frame of the first dry audio is a note start point or note end point, the classification result in the second dimension is a six-class classification result of human octave, and the classification result in the third dimension is a thirteen-class classification result of human pitch; one of the six classification results and one of the thirteen classification results is the classification result of whether each frame of the first dry audio contains human voice.

[0072] In summary, this embodiment can achieve intelligent transcribing through three-dimensional prediction output. Specifically, this embodiment determines the note start point and note end point using the probability matrix in the first dimension. It then determines whether a human voice is present between the note start point and the note end point using one of the classification results from the probability matrices in the second and third dimensions. If a human voice is present, the octave of the human voice is determined using other classification results from the probability matrix in the third dimension, and the pitch of the human voice is determined using other classification results from the probability matrix in the fourth dimension. In one specific implementation, if one of the classification results in the second and third dimensions indicating whether a human voice is present in the first dry audio is -1, it is determined that no human voice is present between the note start point and the note end point, and the note start point and the note end point are discarded. Conversely, if the classification results in the second and third dimensions are not -1, it is determined that a human voice is present between the note start point and the note end point. The octave of the human voice is determined using the other five classification results in the second dimension, and the pitch of the human voice is determined using the other twelve classification results in the third dimension.

[0073] Furthermore, this embodiment determines the training loss based on the above-mentioned transspectral results.

[0074] Step S14: When the training loss does not meet the preset loss condition, return to the step of inputting the first spectral feature and the first semantic feature into the intelligent transspectral model to be trained, until the target intelligent transspectral model whose training loss meets the preset loss condition is obtained.

[0075] It is understood that when the training loss does not meet the preset loss condition, the process returns to the step of inputting the first spectral feature and the first semantic feature into the intelligent transspectral model to be trained, until the target intelligent transspectral model whose training loss meets the preset loss condition is obtained.

[0076] The training method for the above-mentioned intelligent spectrum transposition model is illustrated below with an example. (See [link to example]). Figure 3 As shown:

[0077] (1) Extract the dry audio from the sample song audio to obtain the dry audio.

[0078] (2) Extract 39-dimensional mel spectrum features from the dry audio to obtain a spectral feature matrix (x, 39), where x is the current audio frame number and 39 is the feature dimension. Input the extracted mel spectrum features into the speech recognition model whisper, and obtain 1024-dimensional semantic features, i.e., a (x, 1024)-dimensional matrix, at the encoder output of the whisper model.

[0079] (3) Since the initial semantic features and spectral features have significantly different dimensions, directly inputting them into the network for training would cause the network to favor the semantic features and ignore the original spectral features. In order to ensure that both can be learned by the network equally, this embodiment introduces PCA dimensionality reduction to map the 1024-dimensional semantic features to 39 dimensions.

[0080] (4) The 39-dimensional mel spectrum features and the 39-dimensional semantic features are respectively input into the two inputs of the CRNN network. In this embodiment, the CNN in the CRNN network can identify the transition points and pitch, while the RNN can learn the temporal information.

[0081] (5) The two output features are concatenated at the output of the CRNN, and the concatenated features are then fed into a fully connected layer for prediction. Figure 3 In the example, the output of a fully connected layer has four dimensions: the first dimension is 2-class classification, which determines whether the current frame is the start or end point of a note; the second dimension is 2-class classification, which determines whether there is a human voice between the start and end points of a note; the third dimension is octave 5-class classification; and the fourth dimension is pitch 13-class classification.

[0082] (6) After training the model until it converges, save the model.

[0083] As can be seen, this application provides a method for training an intelligent transspectral model, comprising: extracting a first dry audio from sample audio, extracting a first spectral feature from the first dry audio, inputting the first spectral feature into a pre-trained target speech recognition model, so as to output a first semantic feature of the first dry audio through the target speech recognition model; if the first feature dimension of the first spectral feature is inconsistent with the second feature dimension of the first semantic feature, then adjusting the first feature dimension and the second feature dimension to the same first target feature dimension; inputting the first spectral feature and the first semantic feature into the intelligent transspectral model to be trained, so as to output a plurality of probability matrices corresponding to a plurality of target dimensions through the intelligent transspectral model to be trained, and calculating a training loss based on the plurality of probability matrices corresponding to the plurality of target dimensions; when the training loss does not meet a preset loss condition, returning to the step of inputting the first spectral feature and the first semantic feature into the intelligent transspectral model to be trained, until a target intelligent transspectral model is obtained in which the training loss meets the preset loss condition.

[0084] The beneficial effects of this application are as follows: This application integrates first semantic features and first spectral features during intelligent music transcription. Since note boundary points are often pitch or word transition points, the model in this application can fully utilize the first semantic features and first spectral features to predict note boundary information, thereby improving the accuracy of intelligent music transcription. Furthermore, to prevent inconsistencies between the first feature dimension of the first spectral features and the second feature dimension of the first semantic features, which could prevent the first semantic features and first spectral features from being learned equally by the model, this application adjusts the first spectral features and first semantic features to the same feature dimension when their feature dimensions are mismatched. This allows the first spectral features and first semantic features to be learned equally by the model, further improving the accuracy of intelligent music transcription.

[0085] This application discloses a spectrum transfer method based on a target intelligent spectrum transfer model trained using the aforementioned intelligent spectrum transfer model training method. See also... Figure 4 As shown, it specifically includes:

[0086] Step S21: Extract the second dry audio from the audio to be converted to spectrum, and extract the second spectral features from the second dry audio. Input the second spectral features into the pre-trained target speech recognition model, so as to output the second semantic features of the second dry audio through the target speech recognition model.

[0087] Step S22: If the third feature dimension of the second spectral feature is inconsistent with the fourth feature dimension of the second semantic feature, then the third feature dimension and the fourth feature dimension are adjusted to the same second target feature dimension.

[0088] Step S23: Input the second spectral feature and the second semantic feature into the target intelligent transspectral model, so as to complete the intelligent transspectral transposition of the audio to be transposed through the probability matrices output by the target intelligent transspectral model that correspond to several target dimensions.

[0089] In this embodiment, after the audio to be transcribed is separated into dry audio, a second dry audio is obtained. Further, the mel spectrum features of the second dry audio are extracted and input into the Whisper speech recognition model to extract the semantic features of the second dry audio. After dimensionality reduction of the semantic features, they are input into the target transcribing model to output several probability matrices corresponding to several target dimensions. As can be seen from the above analysis, in a specific implementation, the several probability matrices can be four probability matrices. After obtaining the four probability matrices, the four probability matrices are decoded to obtain the final transcribing result, that is, the pitch sequence.

[0090] It should be noted that, in the process of decoding the probability matrix, to avoid inaccurate local predictions caused by directly taking the maximum probability value, this embodiment uses a maximum value window decoding method to determine the note start point and note end point from the probability matrix of the first dimension. Specifically, the second dry audio of the first target frame that conforms to the first target discrimination rule is determined as the note start point or the note end point; wherein, the first target discrimination rule includes: the probability that the second dry audio of the first several frames and the second dry audio of the last several frames of the first target frame is the note start point or the note end point is less than the probability that the first target frame is the note start point or the note end point, and the probability that the first target frame is the note start point or the note end point is greater than a first preset probability. For example, as shown... Figure 5 As shown, a point is considered the starting point of a note if and only if the probabilities of the three points before and after it are all smaller than the current point, and the probability of the current point itself is greater than 0.7. Figure 5 The first target frame with a probability of 0.8 corresponds to the second dry audio as the audio start point. Similarly, the note end point is determined in the same way. In addition, if no note end point is matched before the next note start point, the next note start point is taken as the note end point corresponding to the current note start point.

[0091] Furthermore, this embodiment determines the probability that the second dry audio frame between the note start point and the note end point contains human voice through a probability matrix in the second dimension. Specifically, this embodiment determines the second dry audio frame of the second target frame that conforms to the second target discrimination rule as the second dry audio containing human voice; wherein, the second target discrimination rule includes: the probability that the second dry audio frame of the second target frame contains human voice is greater than the second preset probability.

[0092] For example, if the second preset probability is set to 0.8, then the current frame with a probability of not less than 0.8 is determined as the second target frame. Further, if the second target frame exists, then the octave value of the human voice is determined according to the probability matrix in the third dimension, the pitch of the human voice is determined according to the probability matrix in the fourth dimension, and finally the pitch sequence is output. The pitch sequence includes the decoded note start point, note end point, corresponding octave value of the human voice, and pitch value.

[0093] It should be noted that intelligent transcribing technology has a wide range of applications. However, many current software transcribing methods have low accuracy and cannot be used directly. Existing transcribing algorithms also have room for optimization. For song synthesis technology based on musical scores, inaccurate boundary points in the musical score can lead to unsatisfactory synthesized vocal effects. This application proposes an intelligent transcribing method based on the fusion of multilingual semantic features and spectral features. A multilingual speech recognition model is introduced into the transcribing model to assist in boundary point prediction. Furthermore, to prevent the introduction of the problem of dimensional mismatch between semantic features and spectral features, this application uses PCA dimensionality reduction to map semantic features to the same dimension as spectral features. In this way, the accuracy of boundary recognition in the transcribing algorithm can be effectively improved, providing higher-quality intelligent pitch sequences for song synthesis based on musical scores and karaoke scoring.

[0094] Accordingly, this application also discloses an intelligent transspectral model training device, see [link to relevant documentation]. Figure 6 As shown, the device includes:

[0095] The sample audio feature extraction module 11 is used to extract a first dry audio from the sample audio, extract a first spectral feature from the first dry audio, and input the first spectral feature into a pre-trained target speech recognition model so as to output the first semantic feature of the first dry audio through the target speech recognition model.

[0096] The sample audio feature dimension adjustment module 12 is used to adjust the first feature dimension and the second feature dimension to the same first target feature dimension if the first feature dimension of the first spectral feature is inconsistent with the second feature dimension of the first semantic feature.

[0097] The model training module 13 is used to input the first spectral feature and the first semantic feature into the intelligent transspectral model to be trained, so as to output a plurality of probability matrices corresponding to a plurality of target dimensions through the intelligent transspectral model to be trained, and to calculate the training loss based on the plurality of probability matrices corresponding to the plurality of target dimensions; when the training loss does not meet the preset loss condition, the process returns to the step of inputting the first spectral feature and the first semantic feature into the intelligent transspectral model to be trained, until a target intelligent transspectral model is obtained in which the training loss meets the preset loss condition.

[0098] For more detailed information on the working process of each of the above modules, please refer to the relevant content disclosed in the foregoing embodiments, which will not be repeated here.

[0099] As can be seen, this application provides a method for training an intelligent transspectral model, comprising: extracting a first dry audio from sample audio, extracting a first spectral feature from the first dry audio, inputting the first spectral feature into a pre-trained target speech recognition model, so as to output a first semantic feature of the first dry audio through the target speech recognition model; if the first feature dimension of the first spectral feature is inconsistent with the second feature dimension of the first semantic feature, then adjusting the first feature dimension and the second feature dimension to the same first target feature dimension; inputting the first spectral feature and the first semantic feature into the intelligent transspectral model to be trained, so as to output a plurality of probability matrices corresponding to a plurality of target dimensions through the intelligent transspectral model to be trained, and calculating a training loss based on the plurality of probability matrices corresponding to the plurality of target dimensions; when the training loss does not meet a preset loss condition, returning to the step of inputting the first spectral feature and the first semantic feature into the intelligent transspectral model to be trained, until a target intelligent transspectral model is obtained in which the training loss meets the preset loss condition.

[0100] The beneficial effects of this application are as follows: This application integrates first semantic features and first spectral features during intelligent music transcription. Since note boundary points are often pitch or word transition points, the model in this application can fully utilize the first semantic features and first spectral features to predict note boundary information, thereby improving the accuracy of intelligent music transcription. Furthermore, to prevent inconsistencies between the first feature dimension of the first spectral features and the second feature dimension of the first semantic features, which could prevent the first semantic features and first spectral features from being learned equally by the model, this application adjusts the first spectral features and first semantic features to the same feature dimension when their feature dimensions are mismatched. This allows the first spectral features and first semantic features to be learned equally by the model, further improving the accuracy of intelligent music transcription.

[0101] Accordingly, this application also discloses a spectrum conversion device, see [link to relevant documentation]. Figure 7 As shown, the device includes:

[0102] The audio feature extraction module 21 is used to extract a second dry audio from the audio to be converted to the spectrum, extract a second spectral feature from the second dry audio, and input the second spectral feature into a pre-trained target speech recognition model so as to output the second semantic feature of the second dry audio through the target speech recognition model.

[0103] The audio feature dimension adjustment module 22 is used to adjust the third feature dimension and the fourth feature dimension to the same second target feature dimension if the third feature dimension of the second spectral feature is inconsistent with the fourth feature dimension of the second semantic feature.

[0104] The intelligent transspectral module 23 is used to input the second spectral feature and the second semantic feature into the target intelligent transspectral model, so as to complete the intelligent transspectral transposition of the audio to be transposed through the probability matrices corresponding to several target dimensions output by the target intelligent transspectral model.

[0105] For more detailed information on the working process of each of the above modules, please refer to the relevant content disclosed in the foregoing embodiments, which will not be repeated here.

[0106] As can be seen, this application proposes a transspectral method, comprising: extracting a second dry audio from the audio to be transspectralized, extracting a second spectral feature from the second dry audio, inputting the second spectral feature into a pre-trained target speech recognition model, so as to output a second semantic feature of the second dry audio through the target speech recognition model; if the third feature dimension of the second spectral feature is inconsistent with the fourth feature dimension of the second semantic feature, then adjusting the third feature dimension and the fourth feature dimension to the same second target feature dimension; inputting the second spectral feature and the second semantic feature into the target intelligent transspectral model, so as to complete the intelligent transspectralization of the audio to be transspectralized through a plurality of probability matrices corresponding to a plurality of target dimensions output by the target intelligent transspectral model.

[0107] The beneficial effects of this application are as follows: This application integrates second semantic features and second spectral features during intelligent music transcription. Since note boundary points are often pitch or word transition points, the model in this application can fully utilize the second semantic features and second spectral features to predict note boundary information, thereby improving the accuracy of intelligent music transcription. Furthermore, to prevent inconsistencies between the second feature dimensions of the second spectral features and the second feature dimensions of the second semantic features, which could prevent the second semantic features and the spectral features from being learned equally by the model, this application adjusts the second spectral features and the second semantic features to have the same feature dimensions when their feature dimensions are mismatched. This allows the second spectral features and the second semantic features to be learned equally by the model, further improving the accuracy of intelligent music transcription.

[0108] Furthermore, embodiments of this application also provide an electronic device. Figure 8 This is a structural diagram of an electronic device 30 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.

[0109] Figure 8 This is a schematic diagram of the structure of an electronic device 30 provided in an embodiment of this application. Specifically, the electronic device 30 may include: at least one processor 31, at least one memory 32, a display screen 33, an input / output interface 34, a communication interface 35, a power supply 36, and a communication bus 37. The memory 32 stores a computer program, which is loaded and executed by the processor 31 to implement the relevant steps in the intelligent transspectral model training method and transspectral method disclosed in any of the foregoing embodiments. Alternatively, the electronic device 30 in this embodiment may specifically be an electronic computer.

[0110] In this embodiment, the power supply 36 is used to provide operating voltage for each hardware device on the electronic device 30; the communication interface 35 can create a data transmission channel between the electronic device 30 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 34 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0111] Furthermore, the memory 32, as a carrier for resource storage, can be a read-only memory, random access memory, disk, or optical disk, etc. The resources stored thereon can include computer programs 321, and the storage method can be temporary storage or permanent storage. The computer programs 321, in addition to including computer programs capable of performing the intelligent transspectral model training method and transspectral method executed by the electronic device 30 as disclosed in any of the foregoing embodiments, may further include computer programs capable of performing other specific tasks.

[0112] Furthermore, embodiments of this application also disclose a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned intelligent transspectral model training method and transspectral method.

[0113] For the specific steps of this method, please refer to the relevant content disclosed in the foregoing embodiments, which will not be repeated here.

[0114] The various embodiments in this application are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. For the same or similar parts between the various embodiments, refer to each other. As for the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and relevant parts can be referred to in the method section.

[0115] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0116] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0117] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0118] The above provides a detailed description of the intelligent transspectral model training method, transspectral method, device, and storage medium provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for training an intelligent transspectral model, characterized in that, include: Extract a first dry audio from the sample audio, extract a first spectral feature from the first dry audio, and input the first spectral feature into a pre-trained target speech recognition model so as to output the first semantic feature of the first dry audio through the target speech recognition model. If the first feature dimension of the first spectral feature is inconsistent with the second feature dimension of the first semantic feature, then the first feature dimension and the second feature dimension are adjusted to the same first target feature dimension. The first spectral feature and the first semantic feature are input into the intelligent transspectral model to be trained, so that the intelligent transspectral model to be trained outputs several probability matrices corresponding to several target dimensions, and the training loss is calculated based on the several probability matrices corresponding to several target dimensions. When the training loss does not meet the preset loss condition, return to the step of inputting the first spectral feature and the first semantic feature into the intelligent transspectral model to be trained, until the target intelligent transspectral model whose training loss meets the preset loss condition is obtained; Wherein, the first feature dimension represents the feature dimension of the first spectral feature; the second feature dimension represents the feature dimension of the first semantic feature; and the preset loss condition represents the condition for determining whether to stop model training after calculating the training loss based on the plurality of probability matrices output by the intelligent transspectral model to be trained.

2. The intelligent transspectral model training method according to claim 1, characterized in that, The step of inputting the first spectral feature and the first semantic feature into the intelligent transspectral model to be trained, so as to output several probability matrices corresponding to several target dimensions through the intelligent transspectral model to be trained, includes: The first spectral feature and the first semantic feature are input into the intelligent transspectral model to be trained, so that the convolutional recurrent neural network in the intelligent transspectral model to be trained outputs a first output feature corresponding to the first spectral feature and a second output feature corresponding to the first semantic feature, respectively. The first output feature and the second output feature are concatenated to obtain the concatenated feature, and the concatenated feature is input into the fully connected layer in the intelligent transspectral model to be trained, so that the fully connected layer outputs several probability matrices corresponding to several target dimensions.

3. The intelligent transspectral model training method according to claim 1, characterized in that, The step of inputting the first spectral feature and the first semantic feature into the intelligent transspectral model to be trained, so as to output several probability matrices corresponding to several target dimensions through the intelligent transspectral model to be trained, includes: The first spectral feature and the first semantic feature are input into the intelligent transspectral model to be trained, so that the intelligent transspectral model to be trained outputs four probability matrices corresponding to the four dimensions; The four probability matrices each include the classification results of several frames of the first dry audio in the four dimensions. The classification result in the first dimension is a binary classification result of whether each frame of the first dry audio is a note start point or note end point. The classification result in the second dimension is a binary classification result of whether each frame of the first dry audio contains human voice. The classification result in the third dimension is a five-class classification result of human voice octave. The classification result in the fourth dimension is a twelve-class classification result of human voice pitch.

4. The intelligent transspectral model training method according to claim 1, characterized in that, The step of inputting the first spectral feature and the first semantic feature into the intelligent transspectral model to be trained, so as to output several probability matrices corresponding to several target dimensions through the intelligent transspectral model to be trained, includes: The first spectral feature and the first semantic feature are input into the intelligent transspectral model to be trained, so that the intelligent transspectral model to be trained outputs three probability matrices corresponding to the three dimensions; The three probability matrices each include classification results of several frames of the first dry audio in the three dimensions. The classification result in the first dimension is a binary classification result of whether each frame of the first dry audio is a note start point or note end point. The classification result in the second dimension is a six-class classification result of human voice octave. The classification result in the third dimension is a thirteen-class classification result of human voice pitch. One of the six-class classification results and one of the thirteen-class classification results is a classification result of whether each frame of the first dry audio contains human voice.

5. The intelligent transspectral model training method according to any one of claims 1 to 4, characterized in that, The first spectral feature includes Mel spectral features or CQT spectral features.

6. A spectral transcoding method, characterized in that, The target intelligent spectrum transfer model trained based on the intelligent spectrum transfer model training method as described in any one of claims 1 to 5 includes: Extract a second dry audio from the audio to be transcribed, and extract a second spectral feature from the second dry audio. Input the second spectral feature into a pre-trained target speech recognition model, so as to output the second semantic feature of the second dry audio through the target speech recognition model. If the third feature dimension of the second spectral feature is inconsistent with the fourth feature dimension of the second semantic feature, then the third feature dimension and the fourth feature dimension are adjusted to the same second target feature dimension. The second spectral feature and the second semantic feature are input into the target intelligent transspectral model, so as to complete the intelligent transspectral transposition of the audio to be transposed through the probability matrices corresponding to several target dimensions output by the target intelligent transspectral model.

7. The transspectral method according to claim 6, characterized in that, The step of inputting the second spectral feature and the second semantic feature into the target intelligent transspectral model, so as to complete the intelligent transspectral transposition of the audio to be transposed through the probability matrices corresponding to several target dimensions output by the target intelligent transspectral model, includes: The second spectral feature and the second semantic feature are input into the target intelligent transspectral model, so as to determine the probability that the second dry audio in each frame is a note start point or note end point through the first probability matrix output by the target intelligent transspectral model corresponding to the first dimension; The second dry audio of the first target frame that conforms to the first target discrimination rule is determined as the note start point or the note end point; The first target discrimination rule includes: the probability that the second dry audio in the first few frames and the second dry audio in the last few frames of the first target frame is the starting point or the ending point of the note is less than the probability that the first target frame is the starting point or the ending point of the note, and the probability that the first target frame is the starting point or the ending point of the note is greater than a first preset probability.

8. The transspectral method according to claim 7, characterized in that, The step of inputting the second spectral feature and the second semantic feature into the target intelligent transspectral model, so as to complete the intelligent transspectral transposition of the audio to be transposed through the probability matrices corresponding to several target dimensions output by the target intelligent transspectral model, includes: The second spectral feature and the second semantic feature are input into the target intelligent transcribing model, so as to determine the probability that the second dry audio between the note start point and the note end point contains human voice through the second probability matrix corresponding to the second dimension output by the target intelligent transcribing model; The second target frame and the second dry audio that meet the second target discrimination rule are determined as the second dry audio containing human voice; wherein, the second target discrimination rule includes: the probability that the second dry audio of the second target frame contains human voice is greater than a second preset probability.

9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the method as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, Used to store a computer program; wherein, when the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Audio synthesis method, terminal equipment and computer readable storage medium

    CN114038474A

  • Basic frequency sequence identification model training method, fundamental frequency sequence identification method, equipment and product

    CN115510911A