Music structure analysis method, terminal device, and computer-readable storage medium

By fusing multidimensional feature information from audio signals with a training model, the problem of detecting choruses in lyricsless music was solved, achieving high-precision audio attribute recognition.

CN119785822BActive Publication Date: 2026-01-13NIO TECH ANHUI CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411902265.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2026-01-13
Estimated Expiration
2044-12-20

AI Technical Summary

Technical Problem

Current technology cannot accurately detect the audio properties of lyricsless music, especially the chorus.

Method used

By acquiring multi-dimensional feature information of audio signals, including rhythm features, emotion features, and frequency domain features, and performing feature fusion processing, the trained model is used to detect audio attributes. The model parameters are then optimized by combining a multi-task loss function to achieve accurate detection of audio attributes.

Benefits of technology

Without relying on lyrics, it significantly improves the accuracy and efficiency of audio attribute detection, especially the ability to recognize the chorus.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119785822B_ABST
    Figure CN119785822B_ABST
Patent Text Reader

Abstract

The application is suitable for the field of audio processing technology, and particularly relates to a music structure analysis method, a terminal device and a computer readable storage medium. The method comprises the following steps: obtaining first multi-dimensional feature information of a first audio signal corresponding to a first time frame; wherein the first multi-dimensional feature information comprises feature information of different dimensions, and the feature information of different dimensions represents different audio semantic features; performing feature fusion processing according to the first multi-dimensional feature information to obtain first fusion feature information; and detecting an audio attribute of the first audio signal according to the first fusion feature information; wherein the audio attribute is used to represent whether the first audio signal is a refrain or a non-refrain. Through the above method, the semantic features of the audio signal can be learned in a deep and more sufficient manner, the audio attribute of the audio signal is detected by using the fusion feature information, and the detection accuracy of the audio attribute can be effectively improved without relying on lyrics texts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of audio processing technology, and in particular relates to a music structure analysis method, terminal device and computer-readable storage medium. Background Technology

[0002] With the development of artificial intelligence (AI) technology, its applications are becoming increasingly widespread. For example, in the music field, AI technology can be used to automatically detect the chorus. The chorus is typically the most prominent part of a song, often possessing high repetition, emotional intensity, and auditory appeal. Chorus detection involves both low-level signal processing and advanced pattern recognition techniques.

[0003] Current chorus detection techniques often rely on the lyrics text information corresponding to the audio, that is, using the lyrics text information as feature information to help detect audio attributes (chorus or non-chorus). However, for some music without lyrics, existing technologies cannot accurately detect the audio attributes of the music. Summary of the Invention

[0004] This application provides a music structure analysis method, terminal device, and computer-readable storage medium, which can effectively improve the detection accuracy of audio attributes without relying on lyrics text.

[0005] In a first aspect, embodiments of this application provide a method for music structure analysis, including:

[0006] Obtain the first multidimensional feature information of the first audio signal corresponding to the first time frame; wherein, the first multidimensional feature information includes feature information of different dimensions, and the feature information of different dimensions represents different audio semantic features;

[0007] Based on the first multidimensional feature information, feature fusion processing is performed to obtain the first fused feature information;

[0008] The audio attributes of the first audio signal are detected based on the first fusion feature information; wherein, the audio attributes are used to indicate whether the first audio signal is a chorus or not.

[0009] In this embodiment, by fusing the multidimensional feature information of the audio signal, the algorithm can learn the semantic features of the audio signal more deeply and fully; by using the fused feature information to detect the audio attributes of the audio signal, the detection accuracy of audio attributes can be effectively improved without relying on the lyrics text.

[0010] In one possible implementation of the first aspect, the first multidimensional feature information includes first feature information, second feature information, and third feature information; the first feature information is the rhythmic feature of the first audio signal, the second feature information is the emotional feature of the first audio signal, and the third feature information is the frequency domain feature of the first audio signal.

[0011] The step of performing feature fusion processing based on the first multidimensional feature information to obtain the first fused feature information includes:

[0012] The fourth feature information is obtained by performing feature fusion processing based on the second feature information and the third feature information;

[0013] The first fused feature information is obtained by performing feature fusion processing based on the fourth feature information and the first feature information.

[0014] In this embodiment, the third feature information is the frequency domain feature of the audio signal. This is equivalent to superimposing the rhythm and emotion features of the audio signal on the frequency domain feature as a basis. In this way, different audio semantic features such as audio spectrum, audio rhythm and audio emotion can be fused to obtain deep audio feature information, providing reliable data basis for the subsequent detection process.

[0015] In one possible implementation of the first aspect, the step of performing feature fusion processing based on the second feature information and the third feature information to obtain the fourth feature information includes:

[0016] A first vector is obtained by performing a linear transformation based on the second feature information;

[0017] A second vector is obtained by performing a linear transformation based on the third feature information;

[0018] The fourth feature information is obtained by mapping the probability distributions of the first vector and the second vector.

[0019] The above method can linearly combine multidimensional feature information and map it into a probability distribution, providing reliable data for subsequent detection tasks.

[0020] In one possible implementation of the first aspect, obtaining the first multidimensional feature information of the first audio signal corresponding to the first time frame includes:

[0021] The first audio signal is input into the trained first model, and the first feature information is output; wherein, the first model is used to identify the rhythmic features of the input audio signal;

[0022] The first audio signal is input into the trained second model, which outputs second feature information; wherein, the second model is used to identify the emotional features of the input audio signal;

[0023] The frequency domain characteristics of the first audio signal are identified to obtain the third feature information.

[0024] In this embodiment, a first model and a second model are pre-trained. Since the detection accuracy of the first model and the second model after training meets the standard, they can accurately detect the feature information of the audio signal in the subsequent detection process, which helps to improve the accuracy and efficiency of detection.

[0025] In one possible implementation of the first aspect, detecting the audio attributes of the first audio signal based on the first fused feature information includes:

[0026] Based on the first fused feature information, feature extraction processing is performed to obtain the first target feature information;

[0027] The segment category of the first audio signal is detected based on the first target feature information to obtain a first detection value; wherein, the first detection value is used to represent the segment category to which the first audio signal belongs;

[0028] Based on the first target feature information, the segment boundaries of the first audio signal are detected to obtain a second detection value; wherein, the second detection value is used to represent the probability that the first audio signal undergoes a segment transition;

[0029] Based on the first target feature information, the first detection value, and the second detection value, the audio attributes of the first audio signal are detected.

[0030] By using the above methods, the model can accurately identify paragraph categories and paragraph boundaries. By combining the detection values ​​of different detection tasks to detect the audio attributes of the audio signal, it is equivalent to considering different semantic features in the audio signal, which helps to improve the model's detection accuracy for the chorus.

[0031] In one possible implementation of the first aspect, detecting the audio attributes of the first audio signal based on the first target feature information, the first detection value, and the second detection value includes:

[0032] An observation sequence is generated based on the first target feature information, the first detection value, and the second detection value;

[0033] The predicted probability for each audio attribute is calculated based on the observation sequence; wherein the predicted probability represents the probability of the state of the second audio signal changing to the state of the first audio signal belonging to the audio attribute, the second audio signal includes the audio signal corresponding to the second time frame, and the second time frame is at least one time frame within a preset time window before the first time frame.

[0034] The audio attribute corresponding to the maximum value in the calculated predicted probabilities is determined as the audio attribute of the first audio signal.

[0035] In the above implementation, the audio attributes of the audio signal in the current time frame are detected based on the state transition of the audio signal between consecutive time frames. This, combined with the feature information of the audio signals from historical time frames, facilitates the extraction of semantic relationships between the audio signals of consecutive time frames, thereby improving detection accuracy. Furthermore, a preset time window can be set according to requirements, increasing the adaptability of the music structure analysis method.

[0036] In one possible implementation of the first aspect, detecting the audio attributes of the first audio signal based on the first fused feature information includes:

[0037] Obtain the trained detection model, which is used to detect the audio attributes of the first audio signal;

[0038] The first fused feature information is input into the trained detection model to obtain the audio attributes of the first audio signal.

[0039] In the above implementation, the audio attributes of the audio signal are detected by the trained detection model, which helps to improve the detection efficiency. In addition, since the detection accuracy of the trained detection model meets the standard, the detection accuracy is improved by using the trained detection model.

[0040] In one possible implementation of the first aspect, obtaining the trained detection model includes:

[0041] Obtain the second multidimensional feature information of the sample signal; the audio properties of the sample signal are known.

[0042] The second multidimensional feature information is input into the detection model, and the audio attributes of the sample signal are output; wherein, the detection model includes a feature fusion module and a first network; the feature fusion module is used to perform feature fusion processing based on the input multidimensional feature information to obtain fused feature information; the first network is used to perform feature extraction processing based on the fused feature information output by the feature fusion module to obtain target feature information;

[0043] The loss value of the detection model is calculated based on the audio properties of the sample signal and the multi-task loss function; wherein, the multi-task loss function includes multiple different loss functions;

[0044] The detection model parameters are updated based on the loss value to obtain the updated detection model;

[0045] Continue training the updated detection model until the detection accuracy of the detection model reaches a preset threshold, and obtain the trained detection model.

[0046] In this embodiment, the multi-task loss function shares model parameters, which can reduce the number of parameters required for training and reduce computational costs, thereby improving training efficiency and the model's generalization ability. In addition, since it includes multiple different loss functions, the model can optimize multiple detection tasks simultaneously, which is beneficial to the model's training efficiency. Furthermore, by learning multiple detection tasks, the model can better understand the internal structure of the data, which is beneficial to enhancing the model's detection accuracy.

[0047] In one possible implementation of the first aspect, the multi-task loss function includes a first function, a second function, and a third function;

[0048] The first function is used to calculate the degree of difference between the audio attributes output by the detection model and the true audio attributes of the sample signal;

[0049] The second function is used to calculate the probability of the sample signal undergoing a segment switching;

[0050] The third function is used to calculate the similarity between the feature representations of the sample signal at different time resolutions.

[0051] The first function enables the model to learn the difference between chorus and non-chorus sections, thus giving it the ability to distinguish between them. The second function enables the model to learn the difference between sections with and without transformations, thus enabling it to distinguish whether a section has undergone a transformation. Since chorus sections often exhibit repetition across multiple time scales, the third function allows the model to capture feature representations at different time resolutions. For example, low-frequency features capture long-term dependencies, while high-frequency features capture short-term changes, thereby improving the model's accuracy in detecting chorus sections. The aforementioned multi-task loss function, combining section classification, section transformation, and feature representations at different time resolutions, enables the model to accurately identify section categories and section boundaries, thereby improving the model's accuracy in detecting chorus sections.

[0052] In a second aspect, embodiments of this application provide a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the music structure analysis method as described in any one of the first aspects above.

[0053] Thirdly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the music structure analysis method as described in any one of the first aspects above.

[0054] Fourthly, embodiments of this application provide a computer program product that, when run on a terminal device, causes the terminal device to execute the music structure analysis method described in any one of the first aspects.

[0055] It is understood that the beneficial effects of the second to fourth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0056] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0057] Figure 1 This is a schematic diagram of the detection system provided in an embodiment of this application;

[0058] Figure 2 This is a schematic diagram of a detection system provided in another embodiment of this application;

[0059] Figure 3 This is a flowchart illustrating the music structure analysis method provided in an embodiment of this application;

[0060] Figure 4 This is a schematic diagram of the structure of the terminal device provided in the embodiments of this application. Detailed Implementation

[0061] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0062] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0063] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0064] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrases "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0065] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0066] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized.

[0067] With the development of artificial intelligence (AI) technology, its applications are becoming increasingly widespread. For example, in the music field, AI technology can be used to automatically detect the chorus. The chorus is typically the most prominent part of a song, often possessing high repetition, emotional intensity, and auditory appeal. Chorus detection involves both low-level signal processing and advanced pattern recognition techniques.

[0068] Current chorus detection techniques often rely on the lyrics text information corresponding to the audio, that is, using the lyrics text information as feature information to help detect audio attributes (chorus or non-chorus). However, for some music without lyrics, existing technologies cannot accurately detect the audio attributes of the music.

[0069] Based on this, this application provides a music structure analysis method. In this application, by fusing the multi-dimensional feature information of the audio signal, the algorithm can learn the semantic features of the audio signal more deeply and fully; by using the fused feature information to detect the audio attributes of the audio signal, the detection accuracy of audio attributes can be effectively improved without relying on lyrics.

[0070] First, the detection system involved in the embodiments of this application is introduced. See also Figure 1 This is a schematic diagram of the detection system provided in the embodiments of this application.

[0071] As an example rather than a limitation, such as Figure 1 As shown, the detection system may include a first model 11, a second model 12, an audio processing module 13, a detection model 14, and a post-processing module 15.

[0072] The first model 11 is used to identify the rhythmic features of the input audio signal. Optionally, the first model 11 can be a neural network model.

[0073] The second model 12 is used to identify the emotional features of the input audio signal. Optionally, the second model 12 can be a neural network model.

[0074] In one embodiment, a first model 11 and a second model 12 can be pre-trained.

[0075] Taking the training of the first model 11 as an example, sample signals are acquired, each corresponding to real rhythm information; the sample signals are input into the first model, and the predicted rhythm information of the sample signals is output; the similarity between the predicted rhythm information output by the model and the real rhythm information is calculated as the loss value; the model parameters of the first model are adjusted according to the loss value; this process is repeated until the first model reaches the preset accuracy, thus obtaining the trained first model. In practical applications, the part of the network used for feature extraction in the trained first model can be used to detect the rhythm features of audio signals.

[0076] Taking the training of the second model 12 as an example, sample signals are acquired, each corresponding to a real sentiment label; the sample signals are input into the second model, and the predicted sentiment label of the sample signal is output; the similarity between the predicted sentiment label output by the model and the real sentiment label is calculated as the loss value; the model parameters of the second model are adjusted according to the loss value; this process is repeated until the second model reaches the preset accuracy, thus obtaining the trained second model. In practical applications, the part of the network used for feature extraction in the trained second model can be used to detect the sentiment features of audio signals.

[0077] In this embodiment, a first model and a second model are pre-trained. Since the detection accuracy of the first model and the second model after training meets the standard, they can accurately detect the feature information of the audio signal in the subsequent detection process, which helps to improve the accuracy and efficiency of detection.

[0078] The audio processing module 13 is used to identify the frequency domain characteristics of the input audio signal. Optionally, the audio processing module 13 may encapsulate audio processing algorithms, such as time-frequency transformation algorithms, Mel filtering algorithms, etc., and perform audio processing on the input audio signal through the encapsulated audio processing algorithms.

[0079] The detection model 14 is used to detect the audio attributes of the audio signal based on the feature information output by the first model 11, the second model 12 and the audio processing module 13 respectively.

[0080] In one embodiment, see Figure 2 This is a schematic diagram of a detection system provided in another embodiment of this application. It is intended as an example and not a limitation. Figure 2 As shown, the detection model 14 may include a feature fusion module 141, a first network 142, a first detection head 143, and a second detection head 144.

[0081] The feature fusion module 141 is used to perform feature fusion processing on the feature information output by the first model 11, the second model 12, and the audio processing module 13 to obtain fused feature information. Optionally, the feature fusion module 141 may encapsulate a feature fusion algorithm based on a cross-attention mechanism.

[0082] The first network 142 is used to perform feature extraction processing based on the fused feature information output by the feature fusion module 141 to obtain target feature information.

[0083] The first detection head 143 is used to detect the segment category of the audio signal based on the target feature information output by the first network 142.

[0084] The second detection head 144 is used to detect segment boundaries of the audio signal based on the target feature information output by the first network 142.

[0085] In one embodiment, the detection model 14 can be pre-trained. The training process may include:

[0086] Obtain the second multidimensional feature information of the sample signal; the audio properties of the sample signal are known.

[0087] The second multidimensional feature information is input into the detection model, and the audio properties of the sample signal are output.

[0088] The loss value of the detection model is calculated based on the audio properties of the sample signal and the multi-task loss function; where the multi-task loss function includes multiple different loss functions.

[0089] The model parameters of the detection model are updated based on the loss value to obtain the updated detection model;

[0090] Continue training the updated detection model until the detection accuracy of the detection model reaches the preset threshold, and obtain the trained detection model.

[0091] In this embodiment, the multi-task loss function shares model parameters, which can reduce the number of parameters required for training and reduce computational costs, thereby improving training efficiency and the model's generalization ability. In addition, since it includes multiple different loss functions, the model can optimize multiple detection tasks simultaneously, which is beneficial to the model's training efficiency. Furthermore, by learning multiple detection tasks, the model can better understand the internal structure of the data, which is beneficial to enhancing the model's detection accuracy.

[0092] In one implementation, the multi-task loss function includes a first function, a second function, and a third function.

[0093] The first function is used to calculate the degree of difference between the audio attributes output by the detection model and the true audio attributes of the sample signal.

[0094] Optionally, the first function can use the cross-entropy loss function to constrain the chorus and non-chorus parts.

[0095] For example, the cross-entropy loss function as the first function can be:

[0096]

[0097] Among them, y i and These represent the real labels and the predicted labels output by the detection model, respectively. N represents the number of categories of audio attributes. In this embodiment, N = 2 (including chorus and non-chorus).

[0098] The first function enables the model to learn the difference between the chorus and non-chorus, thus giving the model the ability to distinguish between the chorus and non-chorus.

[0099] The second function is used to calculate the probability of a segment switching occurring in the sample signal.

[0100] Optionally, the second function can use the cross-entropy loss function to identify the probability of a segment switching occurring in each time frame.

[0101] For example, the cross-entropy loss function as the second function can be:

[0102]

[0103] Among them, a i and These represent the true label and the predicted label output by the detection model, respectively. N represents the transformation category, which can be set to 2 (including transformed and non-transformed categories) in this embodiment.

[0104] The second function enables the model to learn the difference between paragraph transformation and non-transformation, thus giving the model the ability to distinguish whether a paragraph has undergone transformation.

[0105] The third function is used to calculate the similarity between the feature representations of the sample signals at different time resolutions.

[0106] Optionally, the third function can employ a cosine similarity function to identify the degree of similarity between feature representations at different time resolutions.

[0107] For example, the cosine similarity function as the third function can be:

[0108]

[0109] Among them, U i and V i These are two feature vectors corresponding to different time frames at the i-th resolution (such as the target feature information output by the first network), K represents the number of resolutions used, y represents whether the two feature vectors belong to the same segment (1 for yes, 0 for no), cos() is the cosine similarity function, and margin is the control threshold.

[0110] Since the chorus usually has repetition at multiple time scales, the third function enables the model to capture feature representations at different time resolutions. For example, low-frequency features capture long-term dependencies, while high-frequency features capture short-term changes, thereby improving the model's accuracy in detecting the chorus.

[0111] The aforementioned multi-task loss function combines paragraph classification, paragraph transformation, and feature representation at different time resolutions, enabling the model to accurately identify paragraph categories and paragraph boundaries, thereby improving the model's detection accuracy for the chorus.

[0112] The post-processing module 15 is used to detect the audio attributes of the audio signal based on the detection values ​​output by the first detection head 143, the detection values ​​output by the second detection head 144, and the target feature information output by the first network 142. Optionally, the post-processing module 15 may encapsulate a post-processing method based on a hidden Markov model.

[0113] It should be noted that the models or modules in the above detection system are merely examples. In actual applications, more or fewer models / modules can be set up to achieve the same technical solution. This application does not impose specific limitations on this.

[0114] The method provided in this application can be applied to smart devices, which may include terminal devices, driving devices, smart vehicles, robots, and other similar devices. Driving devices may be, for example, vehicles. Terminal devices may be servers, computers, mobile phones, tablets, or other devices capable of audio attribute analysis.

[0115] The music structure analysis method provided in this application is described below. See also... Figure 3 This is a flowchart illustrating the music structure analysis method provided in this application embodiment. It is intended as an example and not a limitation. The method may include the following steps:

[0116] S301, Obtain the first multidimensional feature information of the first audio signal corresponding to the first time frame.

[0117] In the first multidimensional feature information, the feature information of different dimensions represents different audio semantic features.

[0118] It should be noted that the first audio signal in the embodiments of this application may refer to the signal corresponding to a time frame in a certain piece of music. The first audio signal may be a discrete signal or may include multiple continuous signals.

[0119] In one embodiment, the first multidimensional feature information includes first feature information, second feature information, and third feature information. The first feature information is the rhythmic feature of the first audio signal, the second feature information is the emotional feature of the first audio signal, and the third feature information is the frequency domain feature of the first audio signal.

[0120] based on Figure 1 The detection system shown, in one implementation, may include:

[0121] The first audio signal is input into the first model, and the first feature information is output.

[0122] The first audio signal is input into the second model, and the second feature information is output.

[0123] The first audio signal is input into the audio processing module to identify the frequency domain characteristics of the first audio signal and obtain the third feature information.

[0124] In this embodiment, a first model and a second model are pre-trained. Since the detection accuracy of the first model and the second model after training meets the standard, they can accurately detect the feature information of the audio signal in the subsequent detection process, which helps to improve the accuracy and efficiency of detection.

[0125] In this embodiment, the first model and the second model may be obtained by the smart device from other devices (such as the cloud), or the smart device may have both a first model and a second model. Alternatively, the smart device may send the first audio signal to a device with the first model and / or the second model, and then the device with the first model and / or the second model may use the first model and / or the second model to obtain the first feature information and / or the second feature information and send it to the smart device. This embodiment does not limit the scope of the application.

[0126] Optionally, the process by which the audio processing module identifies the frequency domain characteristics of the first audio signal may include:

[0127] The first audio signal is subjected to time-frequency transformation to obtain the first frequency domain signal corresponding to the first audio signal;

[0128] The first frequency domain signal is subjected to Mel filtering to obtain the third feature information.

[0129] In the Mel filtering process, the Mel filter bank is applied to the power spectrum of the signal, the energy in each filter is added together, and the logarithm of the energy of all filter banks is subjected to discrete cosine transform (DCT) to finally extract parameters that reflect speech features.

[0130] Signals processed by Mel filtering can better reflect the characteristics of the human auditory system, thereby improving the efficiency and accuracy of audio detection.

[0131] S302, perform feature fusion processing based on the first multidimensional feature information to obtain the first fused feature information.

[0132] In one embodiment, S302 may include:

[0133] The second and third feature information are fused to obtain the fourth feature information.

[0134] The fourth feature information and the first feature information are fused together to obtain the first fused feature information.

[0135] Understandably, in practical applications, another embodiment is to first perform feature fusion processing on the first feature information and the third feature information to obtain the fifth feature information; then perform feature fusion processing on the fifth feature information and the first feature information to obtain the first fused feature information.

[0136] In this embodiment, the third feature information is the frequency domain feature of the audio signal. This is equivalent to superimposing the rhythm and emotion features of the audio signal on the frequency domain feature as a basis. In this way, different audio semantic features such as audio spectrum, audio rhythm and audio emotion can be fused to obtain deep audio feature information, providing reliable data basis for the subsequent detection process.

[0137] Taking the fusion of second and third feature information as an example, in one implementation, the feature fusion processing method may include concatenating the second and third feature information into a feature vector.

[0138] In another implementation, the feature fusion process may also include: weighted summation of corresponding elements in the second and third feature information.

[0139] In another implementation, the feature fusion process can include:

[0140] A first vector is obtained by performing a linear transformation based on the second feature information;

[0141] The second vector is obtained by performing a linear transformation based on the third feature information.

[0142] The fourth feature information is obtained by mapping the probability distribution based on the first and second vectors.

[0143] For example, feature fusion can be performed using the following formula:

[0144]

[0145] in, This is the second feature information. For the third feature information, W Q and W K Here are the linear transformation parameters, Q is the first vector, K and V are the second vectors, and softmax() is the function used to map the probability distribution.

[0146] It should be noted that in practical applications, the sigmoid function can also be used instead of the sofmax function. This application does not specifically limit the function used to implement the mapping of probability distributions.

[0147] It is understandable that the process of fusing the fourth feature information and the first feature information is the same as the process of fusing the second feature information and the third feature information described above. For details, please refer to the process of fusing the second feature information and the third feature information described above, which will not be repeated here.

[0148] based on Figure 2 In the detection system shown, one implementation of S302 involves inputting the first multi-dimensional feature information into the feature fusion module 141 and outputting the first fused feature information. It is understood that in this implementation, the feature fusion module 141 encapsulates an algorithm for implementing the aforementioned feature fusion processing method.

[0149] The above method can linearly combine multidimensional feature information and map it into a probability distribution. Compared with the simple superposition fusion method, this method can better reflect the correlation between feature information and provide reliable data basis for subsequent detection tasks.

[0150] S303, detect the audio attributes of the first audio signal based on the first fusion feature information.

[0151] In this embodiment of the application, the audio attributes may include two categories, which are used to indicate whether the first audio signal is a chorus or not.

[0152] In one embodiment, S303 may include: detecting the audio properties of the first audio signal based on the first fusion feature information.

[0153] In another embodiment, S303 may include:

[0154] Feature extraction is performed based on the first fused feature information to obtain the first target feature information;

[0155] The segment category of the first audio signal is detected based on the first target feature information to obtain a first detection value; wherein, the first detection value is used to represent the segment category to which the first audio signal belongs;

[0156] Based on the first target feature information, the segment boundaries of the first audio signal are detected to obtain a second detection value; wherein, the second detection value is used to represent the probability that the first audio signal undergoes a segment transition;

[0157] The audio attributes of the first audio signal are detected based on the first target feature information, the first detection value, and the second detection value.

[0158] Compared with directly detecting audio attributes based on the first fusion information, the above method enables the model to accurately identify the paragraph category and paragraph boundary point. By combining the detection values ​​of different detection tasks to detect the audio attributes of the audio signal, it is equivalent to considering different semantic features in the audio signal, which helps to improve the detection accuracy of the model for the chorus.

[0159] based on Figure 2The detection system shown can be implemented in one way as follows: inputting the first fused feature information into the first network 142 and outputting the first target feature information; inputting the first target feature information into the first detection head 143 and outputting the first detection value; inputting the first target feature information into the second detection head 144 and outputting the second detection value; inputting the first target feature information, the first detection value, and the second detection value into the post-processing module 15 and outputting the audio attributes of the first audio signal.

[0160] Optionally, the first network 142 can use an embedding layer. In practical applications, the structure of the embedding layer can be set according to requirements to achieve different functions.

[0161] For example, the embedding layer can transform high-dimensional sparse data into a low-dimensional dense vector representation through matrix multiplication. By performing feature extraction processing on the first fused feature information through such an embedding layer, the dimensionality of the first target feature information is reduced, which enables the detection head to better understand the input data and helps reduce the subsequent computational load.

[0162] For example, the embedding layer can extract high-dimensional features from low-dimensional data. By performing feature extraction processing on the first fused feature information through such an embedding layer, the obtained first target feature information can express richer audio features, thereby enabling the detection head to better capture the semantic information of the audio and improve detection accuracy.

[0163] In one embodiment, the processing steps of the post-processing module 15 may include:

[0164] An observation sequence is generated based on the first target feature information, the first detection value, and the second detection value;

[0165] The predicted probability for each audio attribute is calculated based on the observation sequence. The predicted probability represents the probability of the state of the second audio signal changing to the state of the first audio signal belonging to the audio attribute. The second audio signal includes the audio signal corresponding to the second time frame. The second time frame is at least one time frame within a preset time window before the first time frame.

[0166] The audio attribute corresponding to the maximum value in the calculated predicted probabilities is determined as the audio attribute of the first audio signal.

[0167] In the above implementation, the audio attributes of the audio signal in the current time frame are detected based on the state transition of the audio signal between consecutive time frames. This, combined with the feature information of the audio signals from historical time frames, facilitates the extraction of semantic relationships between the audio signals of consecutive time frames, thereby improving detection accuracy. Furthermore, a preset time window can be set according to requirements, increasing the adaptability of the music structure analysis method.

[0168] Optionally, the post-processing module 15 can detect the audio properties of the first audio signal using a post-processing method based on a hidden Markov model.

[0169] For example, the recursive formula for the forward algorithm of a Hidden Markov Model is:

[0170]

[0171] Where, α t (i) = P(o1,o2,…,p) t ,q t =S i |λ) represents the hidden state S at time t. i And the characteristic sequence o1, o2, ..., o was observed. t The probability of that. t This represents the observation sequence corresponding to time t. The preset parameters of the Hidden Markov Model include A and B, where A = {a...} ij} is the state transition probability matrix, B = {b j (o t )} is the observation probability matrix.

[0172] In this embodiment, six hidden layer states can be set, including before the chorus, during the chorus, after the chorus, before the chorus, during the chorus, and after the chorus. That is, M = 6 in the above formula.

[0173] The above formula can be used to calculate the probability value corresponding to each hidden state, that is, the prediction probability corresponding to each audio attribute.

[0174] Because Hidden Markov Models (HMMs) can abstract decision problems into a series of state and action transitions, and by establishing a state transition probability matrix, complex decision problems can be simplified, using HMM-based post-processing algorithms can simplify the complexity of music structure analysis methods. Furthermore, since HMMs focus on the changes between preceding and following states, using HMM-based post-processing algorithms can help uncover the semantic relationships between audio signals from preceding and following time frames, thereby improving detection accuracy.

[0175] In this embodiment, by fusing the multidimensional feature information of the audio signal, the algorithm can learn the semantic features of the audio signal more deeply and fully; by using the fused feature information to detect the audio attributes of the audio signal, the detection accuracy of audio attributes can be effectively improved without relying on the lyrics text.

[0176] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0177] Figure 4 This is a schematic diagram of the structure of the terminal device provided in the embodiments of this application. For example... Figure 4 As shown, the terminal device 4 in this embodiment includes: at least one processor 40 ( Figure 4 (Only one is shown) a processor, a memory 41, and a computer program 42 stored in the memory 41 and executable on the at least one processor 40, which, when executing the computer program 42, implements the steps in any of the above-described embodiments of the music structure analysis method.

[0178] The terminal device may be a desktop computer, laptop, handheld computer, or cloud server, etc. This terminal device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that... Figure 4 This is merely an example of terminal device 4 and does not constitute a limitation on terminal device 4. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0179] The processor 40 can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0180] In some embodiments, the memory 41 may be an internal storage unit of the terminal device 4, such as a hard disk or memory of the terminal device 4. In other embodiments, the memory 41 may be an external storage device of the terminal device 4, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the terminal device 4. Furthermore, the memory 41 may include both internal and external storage units of the terminal device 4. The memory 41 is used to store the operating system, applications, bootloader, data, and other programs, such as the program code of the computer program. The memory 41 can also be used to temporarily store data that has been output or will be output.

[0181] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the steps in the above-described method embodiments.

[0182] This application provides a computer program product that, when run on a terminal device, enables the terminal device to implement the steps described in the various method embodiments.

[0183] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a device / terminal equipment, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0184] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0185] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0186] In the embodiments provided in this application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0187] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0188] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for analyzing musical structure, characterized in that, include: Obtain the first multidimensional feature information of the first audio signal corresponding to the first time frame; wherein, the first multidimensional feature information includes feature information of different dimensions, and the feature information of different dimensions represents different audio semantic features; Based on the first multidimensional feature information, feature fusion processing is performed to obtain the first fused feature information; The audio attributes of the first audio signal are detected based on the first fusion feature information; wherein, the audio attributes are used to indicate whether the first audio signal is a chorus or not. The step of detecting the audio attributes of the first audio signal based on the first fusion feature information includes: Based on the first fused feature information, feature extraction processing is performed to obtain the first target feature information; The segment category of the first audio signal is detected based on the first target feature information to obtain a first detection value; wherein, the first detection value is used to represent the segment category to which the first audio signal belongs; Based on the first target feature information, the segment boundaries of the first audio signal are detected to obtain a second detection value; wherein, the second detection value is used to represent the probability that the first audio signal undergoes a segment transition; Based on the first target feature information, the first detection value, and the second detection value, the audio attributes of the first audio signal are detected; The step of detecting the audio attributes of the first audio signal based on the first target feature information, the first detection value, and the second detection value includes: An observation sequence is generated based on the first target feature information, the first detection value, and the second detection value; The predicted probability for each audio attribute is calculated based on the observation sequence; wherein the predicted probability represents the probability of the state of the second audio signal changing to the state of the first audio signal belonging to the audio attribute, the second audio signal includes the audio signal corresponding to the second time frame, and the second time frame is at least one time frame within a preset time window before the first time frame. The audio attribute corresponding to the maximum value in the calculated predicted probabilities is determined as the audio attribute of the first audio signal.

2. The music structure analysis method as described in claim 1, characterized in that, The first multidimensional feature information includes first feature information, second feature information, and third feature information; the first feature information is the rhythmic feature of the first audio signal, the second feature information is the emotional feature of the first audio signal, and the third feature information is the frequency domain feature of the first audio signal. The step of performing feature fusion processing based on the first multidimensional feature information to obtain the first fused feature information includes: The fourth feature information is obtained by performing feature fusion processing based on the second feature information and the third feature information; The first fused feature information is obtained by performing feature fusion processing based on the fourth feature information and the first feature information.

3. The music structure analysis method as described in claim 2, characterized in that, The step of performing feature fusion processing based on the second feature information and the third feature information to obtain the fourth feature information includes: A first vector is obtained by performing a linear transformation based on the second feature information; A second vector is obtained by performing a linear transformation based on the third feature information; The fourth feature information is obtained by mapping the probability distributions of the first vector and the second vector.

4. The music structure analysis method as described in claim 2 or 3, characterized in that, The acquisition of the first multidimensional feature information of the first audio signal corresponding to the first time frame includes: The first audio signal is input into the trained first model, and the first feature information is output; wherein, the first model is used to identify the rhythmic features of the input audio signal; The first audio signal is input into the trained second model, which outputs second feature information; wherein, the second model is used to identify the emotional features of the input audio signal; The frequency domain characteristics of the first audio signal are identified to obtain the third feature information.

5. The music structure analysis method as described in claim 1, characterized in that, The step of detecting the audio attributes of the first audio signal based on the first fusion feature information includes: Obtain the trained detection model, which is used to detect the audio attributes of the first audio signal; The first fused feature information is input into the trained detection model to obtain the audio attributes of the first audio signal.

6. The music structure analysis method as described in claim 5, characterized in that, The process of obtaining the trained detection model includes: Obtain the second multidimensional feature information of the sample signal; the audio properties of the sample signal are known. The second multidimensional feature information is input into the detection model, and the audio attributes of the sample signal are output; wherein, the detection model includes a feature fusion module and a first network; the feature fusion module is used to perform feature fusion processing based on the input multidimensional feature information to obtain fused feature information; the first network is used to perform feature extraction processing based on the fused feature information output by the feature fusion module to obtain target feature information; The loss value of the detection model is calculated based on the audio properties of the sample signal and the multi-task loss function; wherein, the multi-task loss function includes multiple different loss functions; The detection model parameters are updated based on the loss value to obtain the updated detection model; Continue training the updated detection model until the detection accuracy of the detection model reaches a preset threshold, and obtain the trained detection model.

7. The music structure analysis method as described in claim 6, characterized in that, The multi-task loss function includes a first function, a second function, and a third function; The first function is used to calculate the degree of difference between the audio attributes output by the detection model and the true audio attributes of the sample signal; The second function is used to calculate the probability of the sample signal undergoing a segment switching; The third function is used to calculate the similarity between the feature representations of the sample signal at different time resolutions.

8. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 7.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Musical tone signal classification method based on multi-feature fusion and feature selection

    CN107133643A

  • Chinese song emotion classification method based on multi-modal fusion

    CN110674339A