Song semantic information indexing method and its device, equipment, medium, product

By encoding and feature extraction of song audio data, and using the feature extraction model of shared network and branch network, the problem of insufficient accuracy and efficiency of cover version song recognition in the prior art is solved, and efficient identification and retrieval of massive music is achieved.

CN114817621BActive Publication Date: 2025-06-20GUANGZHOU KUGOU COMP TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111491602.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-08
Publication Date
2025-06-20
Estimated Expiration
2041-12-08

AI Technical Summary

Technical Problem

It is difficult for the prior art to accurately identify cover versions of songs. The traditional method has insufficient recognition accuracy and efficiency, and cannot be applied to the search of massive music.

Method used

A song semantic information index method is adopted to encode the song audio data, and multi-level feature extraction is performed using the shared network of the feature extraction model and multiple branch networks to obtain a high-dimensional index vector of deep semantic information.

Benefits of technology

It realizes accurate identification of cover version songs, improves recognition accuracy and efficiency, and can be suitable for the search and retrieval of massive music.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114817621B_ABST
    Figure CN114817621B_ABST
Patent Text Reader

Abstract

The present application discloses a method and apparatus, device, medium, and product for indexing song semantic information. The method includes: encoding audio information in song audio data to obtain corresponding encoded information; sequentially performing multi-level feature extraction on the encoded information by using a plurality of convolutional blocks in a shared network of a feature extraction model that has been trained to a convergent state to obtain intermediate feature information that extracts the deep semantic information of the song audio data; using a global branch network of the feature extraction model to extract global significant features from the intermediate feature information to obtain a global output feature vector; using a local branch network of the feature extraction model to respectively extract semantic local features from the intermediate feature information by equally dividing by channels to obtain a channel output feature vector; and splicing the global output feature vector and the channel feature vector into a high-dimensional index vector. The present application can achieve representation learning of the deep semantic information of song audio data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of music information retrieval, and in particular, to a method for indexing song semantic information, as well as a corresponding device, computer device, computer-readable storage medium, and computer program product. Background Art

[0002] With the popularity of short videos, live broadcasts, and radio stations, the number of cover songs is increasing, and the scenarios requiring music recognition are becoming more and more complex. Compared with the original version, the cover version may have differences or even be completely different in music components such as timbre, fundamental frequency, rhythm, speed, harmony, lyrics, singing style, and overall structure. Therefore, cover song recognition is a very challenging research task.

[0003] There are various existing technologies related to cover song recognition, and each existing technology has certain deficiencies to some extent. For example: (1) Traditional landmark-based song recognition technologies can only recognize songs of the same source version and cannot recognize the above-mentioned cover versions with certain differential information; (2) Traditional melody matching-based humming recognition technologies can only recognize clean a cappella / humming and cannot recognize the above-mentioned cover versions with background accompaniment; (3) Traditional cover song recognition technology solutions mainly extract audio features such as Pitch Class Profile (PCP), and then use algorithms such as dynamic programming to calculate the similarity distance between songs. Due to the diversity of cover versions, the above solutions can only be applied to cover song solutions with minor adaptations, with low accurate recognition rate and slow recognition speed, and cannot be applied to the search for a large amount of music.

[0004] Therefore, the existing technical solutions for song recognition lack general adaptability, have low recognition accuracy, and low recognition efficiency. It is necessary to explore more effective technical solutions. Summary of the Invention

[0005] The primary objective of this application is to solve at least one of the above problems and provide a method for indexing song semantic information, as well as a corresponding device, computer device, computer-readable storage medium, and computer program product.

[0006] To meet the various objectives of this application, the following technical solutions are adopted:

[0007] A method for indexing song semantic information provided to meet one of the objectives of this application includes the following steps:

[0008] Encode the audio information in the song audio data to obtain corresponding encoded information;

[0009] Use multiple convolutional blocks in the shared network of the feature extraction model that has been trained to convergence to sequentially perform multi-level feature extraction on the encoded information, obtaining intermediate feature information that extracts the deep semantic information of the song audio data;

[0010] Use the global branch network of the feature extraction model to extract global significant features from the intermediate feature information, obtaining a global output feature vector;

[0011] Use the local branch network of the feature extraction model to respectively extract semantic local features from the intermediate feature information by equally dividing along channels, obtaining a channel output feature vector;

[0012] Concatenate the global output feature vector and the channel feature vector into a high-dimensional index vector.

[0013] In a further embodiment, the step of using the global branch network of the feature extraction model to extract global significant features from the intermediate feature information and obtaining a global output feature vector includes the following steps:

[0014] Perform multi-level feature extraction on the intermediate feature information through multiple convolutional blocks of the global branch network in sequence to obtain deep feature information;

[0015] Perform a maximum pooling operation on the deep feature information to obtain a global output feature vector that extracts the global significant features of the song audio data.

[0016] In a further embodiment, the step of using the local branch network of the feature extraction model to respectively extract semantic local features from the intermediate feature information by equally dividing along channels and obtaining a channel output feature vector includes the following steps:

[0017] Perform multi-level feature extraction on the intermediate feature information through multiple convolutional blocks of the local branch network in sequence to obtain deep feature information;

[0018] Equally divide the deep feature information along the channel direction to obtain multiple equally divided feature information;

[0019] Perform an average pooling operation on each equally divided feature information respectively to obtain multiple equally divided feature vectors, and concatenate all the equally divided feature vectors into the local output feature vector.

[0020] In an alternative embodiment, in the step of encoding the audio information in the song audio data, the audio information is any one of the time-frequency spectrum information, Mel spectrum information, CQT filter information, pitch contour information, and Chroma feature information of the song audio data.

[0021] In a specific embodiment, in the shared network, at least one of the convolutional blocks applies an attention module to extract key information from the song audio data, and the attention module is a spatial attention module or a channel attention module.

[0022] In a further embodiment, the convolutional block is used to perform the following steps:

[0023] Perform a convolutional transformation on the information input thereto to obtain transformed feature information;

[0024] Perform instance normalization and batch normalization on the transformed feature information respectively and then combine them into concatenated feature information, and activate and output the concatenated feature information;

[0025] Obtain residual information after performing multiple convolutional operations and batch normalization on the concatenated feature information activated and output;

[0026] Superimpose the residual information on the information input thereto and activate and output.

[0027] In an extended embodiment, the training process of the feature extraction model includes the following steps of iterative training:

[0028] Call a training sample from the training set, and determine the encoded information of the audio information of the training sample. The training sample is pre-collected song audio data, and the song audio data is a complete song or a segment thereof;

[0029] Input the encoded information into the feature extraction model to train it to obtain corresponding output feature vectors;

[0030] Perform classification prediction on each of the output feature vectors respectively to map out corresponding classification labels;

[0031] Calculate the loss value of the feature extraction model using the supervision label corresponding to the training sample and the classification label, and perform gradient update on the feature extraction model according to the loss value;

[0032] Judge whether the loss value reaches a preset threshold. When it does not reach the preset threshold, call the next training sample in the training set to continue iterative training of the feature extraction model until the loss value reaches the preset threshold.

[0033] In an extended embodiment, after the step of fusing the global output feature vector and the channel feature vector into a high-dimensional index vector, the following steps are included:

[0034] Respond to a query request for query audio data, and call the feature extraction model to extract its corresponding high-dimensional index vector as a query vector;

[0035] Calculate the similarity between the query vector and multiple high-dimensional index vectors in a preset song feature library to obtain a similarity data sequence. The high-dimensional index vectors stored in the song feature library are obtained by the feature extraction model extracting corresponding song audio data in a preset song library.

[0036] Determine the song audio data corresponding to the maximum similarity exceeding a preset threshold in the similarity data sequence as the similar song of the query audio data.

[0037] In a specific embodiment, before the step of calculating the similarity between the query vector and multiple high-dimensional index vectors in a preset song feature library, the following steps are included:

[0038] Calculate the similarity between the query vector and the high-dimensional index vectors of each song audio data in a non-melody feature library to obtain a corresponding similarity sequence. The high-dimensional index vectors in the non-melody feature library are obtained by the feature extraction model extracting each song audio data with non-melody information.

[0039] Compare whether the similarity of each song audio data in the similarity sequence is lower than a preset threshold. When the similarity of all song audio data is lower than this preset threshold, continue the subsequent steps; otherwise, terminate the subsequent steps.

[0040] In an extended embodiment, after the step of concatenating the global output feature vector and the channel feature vector into a high-dimensional index vector, the following steps are included:

[0041] Obtain the query audio data to be compared.

[0042] Call the feature extraction model to determine the high-dimensional index vector corresponding to the query audio data.

[0043] Calculate the similarity between the high-dimensional index vector of the query audio data and the high-dimensional index vector of the song audio data.

[0044] Judge whether the similarity exceeds a preset threshold. When it exceeds the preset threshold, determine that the query audio data and the song audio data form a similar song.

[0045] A song semantic information indexing device provided to meet one of the purposes of the present application includes: an encoding processing module, a shared extraction module, a global extraction module, a local extraction module, and an indexing processing module. Among them, the encoding processing module is used to encode the audio information in the song audio data to obtain corresponding encoded information; the shared extraction module is used to sequentially perform multi-level feature extraction on the encoded information by using multiple convolutional blocks in the shared network of the feature extraction model that has been trained to a convergent state, to obtain intermediate feature information that extracts the deep semantic information of the song audio data; the global extraction module is used to extract global significant features from the intermediate feature information by using the global branch network of the feature extraction model to obtain a global output feature vector; the local extraction module is used to extract semantic local features from the intermediate feature information by equally dividing it by channel by using the local branch network of the feature extraction model to obtain a channel output feature vector; the indexing processing module is used to splice the global output feature vector and the channel feature vector into a high-dimensional index vector.

[0046] A computer device provided to meet one of the purposes of the present application includes a central processing unit and a memory. The central processing unit is used to call and run a computer program stored in the memory to execute the steps of the song semantic information indexing method described in the present application.

[0047] A computer-readable storage medium provided to meet another purpose of the present application stores a computer program implemented based on the song semantic information indexing method in the form of computer-readable instructions. When the computer program is called and run by a computer, it executes the steps included in the method.

[0048] A computer program product provided to meet another purpose of the present application includes a computer program / instructions. When the computer program / instructions are executed by a processor, they implement the steps of the method described in any embodiment of the present application.

[0049] Compared with the prior art, the advantages of the present application are as follows:

[0050] First, the present application encodes the audio information of the song audio data to obtain corresponding encoded information to obtain style-invariant features of the song audio data. Then, an intermediate feature information is extracted from the encoded information through a shared network. Based on the intermediate feature information, through multiple branch networks, deep semantic information of the song audio data is extracted from different angles to obtain corresponding output feature information. Finally, these output feature information are used as the high-dimensional index vector corresponding to the song audio data to complete the end-to-end representation learning of the song audio data.

[0051] Secondly, since the feature extraction model adopted by the present application uses a combination of a shared network and multiple branch networks to realize multi-angle feature extraction of deep semantic information of song audio data, the obtained high-dimensional index vector can be made more representative. Specifically, not only the deep semantic representation of the global significant features of the song audio data is realized through the global branch network, but also the deep semantic representation of the local significant features of the song audio data is realized through the local branch network, thereby realizing more effective indexing of the corresponding song audio data. On this basis, downstream processing such as retrieval, query, and matching of song audio data can be performed, which can obtain more accurate and efficient matching effects, and can be used in various application scenarios such as cover recognition, song identification, humming recognition, and song infringement determination.

[0052] In addition, this application fuses the output feature vectors obtained by multiple branch networks into a single high-dimensional index vector, which has a wide range of uses and is flexible in usage. It can achieve relatively obvious scale results when processing the representation learning of massive song audio data. It can be deployed in the background of the online music service platform to realize a standardized interface, and then serve the needs of various different application scenarios, provide comprehensive and multi-purpose open services, and enhance the platform's economic advantages in music information retrieval. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0054] Figure 1 A flowchart of a typical embodiment of the song semantic information indexing method of the present application;

[0055] Figure 2 This is a principle block diagram of the network architecture of a feature extraction model in one embodiment of the present application;

[0056] Figure 3 A schematic diagram of the working process of the residual convolution block used in the feature extraction model of the present application;

[0057] Figure 4 A flowchart of the process of training the feature extraction model of the present application;

[0058] Figure 5 This is a principle block diagram of the classification model to which the feature extraction model of this application is connected during the training phase;

[0059] Figure 6 A schematic diagram of a flow chart of an implementation of the feature extraction model of the present application for implementing similar song matching;

[0060] Figure 7Schematic diagram of another implementation mode in which the feature extraction model of the present application is used to implement similar song matching. In this implementation mode, a melody-free feature library is referenced for filtering;

[0061] Figure 8 Schematic diagram of yet another implementation mode in which the feature extraction model of the present application is used to implement similar song matching. In this implementation mode, it is required for song infringement comparison;

[0062] Figure 9 Principle block diagram of the song semantic information indexing device of the present application;

[0063] Figure 10 Schematic diagram of the structure of a computer device adopted by the present application. Specific implementation mode

[0064] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below with reference to the drawings are exemplary and are only used to explain the present application and cannot be construed as a limitation to the present application.

[0065] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present application means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.

[0066] Those skilled in the art of the present technology can understand that unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.

[0067] Those skilled in the art can understand that the "client", "terminal", and "terminal device" used herein include both devices with a wireless signal receiver that only has the ability to receive and no ability to transmit, and devices with both receiving and transmitting hardware that can conduct two-way communication on a two-way communication link. Such devices can include: cellular or other communication devices such as personal computers and tablet computers, which have a single-line display or a multi-line display or a cellular or other communication device without a multi-line display; PCS (Personal Communications Service), which can combine voice, data processing, fax, and / or data communication capabilities; PDA (Personal Digital Assistant), which can include a radio frequency receiver, a pager, Internet / intranet access, a web browser, a notepad, a calendar, and / or a GPS (Global Positioning System) receiver; conventional laptop and / or palm computers or other devices, which are conventional laptop and / or palm computers or other devices with and / or including a radio frequency receiver. The "client", "terminal", and "terminal device" used herein can be portable, transportable, installed in a vehicle (air, sea, and / or land), or suitable for and / or configured to run locally, and / or run in a distributed manner at any other location on the earth and / or in space. The "client", "terminal", and "terminal device" used herein can also be a communication terminal, an Internet access terminal, a music / video playback terminal, for example, it can be a PDA, a MID (Mobile Internet Device), and / or a mobile phone with music / video playback function, or can also be devices such as a smart TV and a set-top box.

[0068] The hardware referred to by names such as "server", "client", and "service node" in this application is essentially an electronic device with the equivalent capabilities of a personal computer, and is a hardware device with the necessary components disclosed by the von Neumann principle, including a central processing unit (including an arithmetic unit and a controller), a memory, an input device, and an output device. The computer program is stored in its memory, and the central processing unit loads the program stored in the external memory into the memory for execution, executes the instructions in the program, and interacts with the input / output devices to complete specific functions.

[0069] It should be noted that the concept of "server" in this application can similarly be extended to apply to server clusters. According to the network deployment principles understood by those skilled in the art, the servers should be logically divided. Physically, these servers can either be independent of each other but can be invoked through interfaces, or integrated into a single physical computer or a set of computer clusters. Those skilled in the art should understand this flexibility and should not be restricted by it in the implementation of the network deployment method of this application.

[0070] One or several technical features of this application, unless explicitly specified, can either be deployed on the server and accessed by the client remotely invoking the online service interface provided by the server, or directly deployed and run on the client for access.

[0071] The neural network models cited or possibly cited in this application, unless explicitly specified, can either be deployed on a remote server and remotely invoked by the client, or deployed on a client capable of handling the device and directly invoked. In some embodiments, when it runs on the client, its corresponding intelligence can be obtained through transfer learning to reduce the requirements for the client's hardware operation resources and avoid excessive occupation of the client's hardware operation resources.

[0072] All kinds of data involved in this application, unless explicitly specified, can either be remotely stored on the server or stored on the local terminal device, as long as it is suitable for being invoked by the technical solution of this application.

[0073] Those skilled in the art should be aware that although the various methods of this application are described based on the same concept and thus show commonality with each other, unless otherwise specified, these methods can be executed independently. Similarly, for the various embodiments disclosed in this application, they are all proposed based on the same inventive concept. Therefore, for concepts with the same expression, as well as concepts that are only appropriately transformed for convenience although the concept expressions are different, they should be equivalently understood.

[0074] For the various embodiments to be disclosed in this application, unless explicitly pointed out that there is a mutually exclusive relationship between them, the relevant technical features involved in each embodiment can be cross-combined to flexibly construct new embodiments, as long as this combination does not deviate from the creative spirit of this application and can meet the requirements in the prior art or solve certain deficiencies in the prior art. Those skilled in the art should be aware of this flexibility.

[0075] A method for indexing song semantic information in this application can be programmed as a computer program product and deployed to run on the server. Thereby, the client can access the interface opened after the computer program product runs in the form of a web program or an application program, and achieve human-computer interaction with the process of the computer program product through the graphical user interface.

[0076] See also Figure 1 The song semantic information indexing method of the present application, in its typical embodiment, comprises the following steps:

[0077] Step S1100: Encode the audio information in the song audio data to obtain corresponding encoding information:

[0078] The song audio data may be audio data in any format such as MP3, WMA, M4A, WAV, etc., or may be audio data obtained by separating audio from various video files. Song audio data generally includes multiple voice data packets in the time domain. The song audio data may come from various songs pre-stored in the music library of an online music service platform, or may be songs or a cappella audio submitted by users in real time. In this regard, the present application may flexibly respond and process according to specific tasks to serve different types of needs. For example, when constructing a song feature library of song audio data in the music library to serve the needs of song recognition, humming recognition, and cover recognition, it is necessary to extract features from the song audio data in the music library one by one; for example, when it is necessary to retrieve, query, and match songs for song recognition, humming recognition, cover recognition, and song infringement determination, it is necessary to obtain the song audio data submitted by the client for feature extraction, etc. On this basis, various transformations are usually performed on the voice data packets to encode the audio information in the song audio data as required by the present application, thereby obtaining the corresponding encoding information.

[0079] The audio information is mainly used to describe the relevant information of the style-invariant features in the song audio data, and can be of various types, including but not limited to the time-frequency spectrum information, Mel-spectrum information, CQT filter information, pitch contour information, Chroma feature information, etc. extracted from the voice data packet of the song audio data. Such information can be encoded using a corresponding algorithm to obtain the corresponding type of encoding information. In the present application, any of the above types of encoding information can be used to implement feature extraction in the present application. In practice, it is recommended to encode the optimal CQT filter information measured to obtain the encoding information.

[0080] Those skilled in the art understand that the above various audio information can all be encoded using corresponding algorithms. During the encoding process, it is necessary to first perform conventional processing such as pre-emphasis, framing, and windowing on the song audio data, and then perform time-domain or frequency-domain analysis, that is, to achieve speech signal analysis. The purpose of pre-emphasis is to enhance the high-frequency part of the speech signal and make the spectrum smooth; generally, pre-emphasis is achieved through a first-order high-pass filter. Before analyzing the speech signal, it is also necessary to frame it. Usually, the length of each frame of the speech signal is set to 20 ms. Considering the frame shift factor, there is an overlap of 10 ms between adjacent frames. To achieve framing, windowing operations can be performed on the speech signal. Different window selections will affect the results of speech signal analysis. More commonly, the window function corresponding to the Hamming window (Hamm) is used to perform the windowing operation.

[0081] On the basis of completing the preprocessing required for speech signal analysis of the song audio data, time-domain and frequency-domain analysis can be further performed on it to achieve encoding and obtain corresponding encoded information:

[0082] For the time-frequency spectrum information described above, by performing pre-emphasis, framing, windowing, and short-time Fourier transform (STFT) on the speech data of each speech data packet in the time domain to transform it into the frequency domain, data corresponding to the spectrogram can be obtained, thereby constituting the time-frequency spectrum information described above.

[0083] The Mel spectrum information described above can be obtained by filtering the time-frequency spectrum information using a Mel-scale filter bank. Similarly, by taking the logarithm and performing a DCT transform on the Mel spectrum information, the corresponding Mel cepstrum information can also be obtained. It can be understood that the Mel spectrum information and its Mel cepstrum information can better describe the style-invariant features in the song, such as pitch, intonation, timbre, etc.

[0084] For the CQT filtering information, since in music, all tones are composed of the 12 equal temperament of several octaves, that is, the twelve-tone equal temperament, corresponding to the twelve semitones on one octave of a piano. The frequency ratio between adjacent semitones is 2 1 / 12. Obviously, for two octaves of the same pitch, the frequency of the high octave is twice that of the low octave. Therefore, in music, sounds are distributed exponentially, but the audio spectra obtained by Fourier transform are distributed linearly. The frequency points of the two cannot correspond one to one, which will cause errors in the estimated values ​​of some scale frequencies. Therefore, the CQT time-frequency transform algorithm can be used to replace the Fourier transform method for speech analysis. CQT, Constant Q Transform, refers to a filter group with a center frequency distributed according to an exponential law, different filter bandwidths, but a center frequency to bandwidth ratio of constant Q. It is different from Fourier transform in that the horizontal axis frequency of its spectrum is not linear, but based on log2, and the filter window length can be changed according to the different spectrum line frequencies to obtain better performance. Since CQT has the same distribution as the scale frequency, the amplitude value of the music signal at each note frequency can be directly obtained by calculating the CQT spectrum of the music signal, which is more perfect for music signal processing. Therefore, this embodiment recommends using this information to perform corresponding encoding to obtain corresponding encoding information as the input of the neural network model of this application.

[0085] The pitch contour information, including PCP (Pitch Class Profile) and HPCP (Harmonic Pitch Class Profile), is intended to extract the corresponding pitch sequence from the song audio data, convert it into a melody contour sequence after regularization, merging, and segmentation, and then convert the standard pitch difference generated by the standard pitch into the corresponding feature representation. The coding information constructed based on the pitch contour information has good robustness to environmental noise.

[0086] The Chroma feature information is a general term for the chroma vector and the chroma spectrum. The chroma vector is a vector containing 12 elements, which represent the energy in 12 pitches within a period of time (such as 1 frame). The energy of the same pitch in different octaves is accumulated, and the chroma spectrum is a sequence of chroma vectors. Specifically, the voice data packet of the song audio data is transformed from the time domain to the frequency domain by short-time Fourier transform, and then some noise reduction processing is performed, and then tuning is performed; the absolute time is converted into a frame according to the length of the selected window, and the energy of each pitch in each frame is recorded to form a pitch spectrum; on the basis of the pitch spectrum, the energy (in loudness) of the notes of the same time, the same pitch, and different octaves is superimposed on the elements of the pitch in the chroma vector to form a chroma spectrum. The data corresponding to the chroma spectrum is the Chroma feature information.

[0087] Any of the above specific audio information can be used as the input of the feature extraction model of this application. For the convenience of processing by the feature extraction model, the audio information can be converted into corresponding encoded information according to a certain preset format. For example, the audio information corresponding to each voice packet is organized into a row vector, and the row vectors of each voice data packet of the song audio data are organized together by time sequence to obtain a two-dimensional matrix as the encoded information. And so on, it can be preset to adapt to the feature extraction model and can be flexibly implemented by those skilled in the art.

[0088] Step S1200: Use multiple convolutional blocks in the shared network of the feature extraction model that has been trained to a convergent state to sequentially perform multi-level feature extraction on the encoded information, and obtain intermediate feature information that extracts the deep semantic information of the song audio data.

[0089] In order to perform feature extraction on the corresponding song audio data based on the encoded information, this application proposes a new feature extraction model based on the neural network model architecture. By pre-training the model to a convergent state, it enables the model to learn the ability to extract the deep semantic information of the song audio data according to the encoded information and thus obtain the corresponding output feature vector, completing the representation learning of the style-invariant features of the song audio data for querying, retrieval, and matching.

[0090] As Figure 2 shown in the principle block diagram, the feature extraction model is composed of a shared network and two branch networks. Among them, the shared network includes multiple convolutional blocks for gradually extracting the deep semantic information of the encoded information to obtain intermediate feature information; the two branch networks respectively perform the extraction of different types of deep semantic information based on the intermediate feature information to obtain the corresponding output feature information. Between the two branch networks, each contains a part of the same structure, which contains multiple convolutional blocks for gradually extracting deep semantic information. After the output of the last convolutional block, different processing can be performed according to the different functions of each branch network.

[0091] The convolution block described above can be implemented using convolutional layers based on CNN and RNN. Preferably, a convolution block based on the principle of residual convolution is used. In order to play the role of context sorting to extract key information from the song audio data, an attention mechanism can be applied in any one of the convolution blocks, and a corresponding attention module is added. Preferably, the attention module described above is applied in the last convolution block of the shared network, specifically a Spatial Attention Module (SAM) or a Channel Attention Module (CAM). In an enhanced embodiment, instance normalization operation (IN) and batch normalization (BN) operations are applied in the convolution block to divide the information input therein into two parts. One part undergoes instance normalization operation to learn style-invariant features, and the other part performs batch normalization operation to achieve normalization. Therefore, the so-called IBN architecture is applied. By applying this architecture, music attribute invariant features with highly diverse styles of song audio data, such as notes, rhythms, timbres, etc., can be learned while retaining version information.

[0092] Accordingly, it is not difficult to understand that the feature extraction model is applicable to different application scenarios. By enabling two branch networks and training them to a convergent state using a preselected training set first, the corresponding feature extraction ability can be obtained, so as to be suitable for performing tasks corresponding to the application scenarios and extracting the output feature information corresponding to the song audio data from the encoded information of the song audio data input therein. The training process of the feature extraction model will be given in the exemplary embodiments of this application and will not be elaborated here for the time being.

[0093] In this step, in the architecture as Figure 2 shown, after the encoded information is subjected to feature extraction step by step through multiple convolution blocks of the shared network, especially after key information extraction through the last convolution block therein, intermediate feature information containing the key information of the encoded information can be obtained. This intermediate feature information is divided into multiple paths and output to the two branch networks for further extraction of deep semantic information from different angles.

[0094] Step S1300: Use the global branch network of the feature extraction model to extract global significant features from the intermediate feature information to obtain a global output feature vector:

[0095] As described above, Figure 2 in the architecture shown, the intermediate feature information output by the shared network is respectively fed into the two branch networks for further feature extraction processing. In this step, the global branch network receives the input of the intermediate feature information

[0096] According toFigure 2 In the architecture shown, each branch network belonging to the same structural part includes two convolutional blocks. After the two convolutional blocks sequentially perform feature extraction on the feature information output to them, the extracted output can be processed differently according to the specific structure of different branch networks.

[0097] For the global branch network, after performing hierarchical feature extraction on the intermediate feature information through two convolutional blocks with the same structure as other branch networks, the output of the last convolutional block is used to extract the global significant feature information through a max-pooling operation. Thus, a global output feature vector is correspondingly output. According to this architecture, the global output feature vector can capture the overall significant features of the song audio data, improving the recognition ability of the model.

[0098] In an embodiment of a further modification, this step S1300 can be implemented as including the following specific steps:

[0099] Step S1310: Sequentially perform multi-level feature extraction on the intermediate feature information through multiple convolutional blocks of the global branch network to obtain deep feature information:

[0100] Specifically, two of the above-mentioned convolutional blocks are used, and the two convolutional blocks can adopt exactly the same network architecture and both apply the IBN block, so that the deep feature information obtained after feature extraction can learn the style-invariant features in the song audio data.

[0101] Step S1320: Perform a max-pooling operation on the deep feature information to obtain a global output feature vector that extracts the global significant features of the song audio data:

[0102] Based on the deep feature information, a max-pooling (MaxPooling) operation is directly performed on it through a pooling layer to extract the significant features from a global perspective.

[0103] In this application, the output feature information output by each branch network is normalized to be represented as an output feature vector. For example, the global output feature vector of the global branch network can be normalized to a 512-dimensional feature vector.

[0104] Step S1400: Use the local branch network of the feature extraction model to separately extract semantic local features from the intermediate feature information by channel equal segmentation to obtain channel output feature vectors:

[0105] Figure 2In the given local branch network, after the intermediate feature information is successively subjected to feature extraction through two convolutional blocks with the same structure as the global branch network, the output of the last convolutional block (number of channels * number of frequency bands * number of frames) is divided into four equal parts along the channel dimension and then output. Then, average pooling is respectively performed on each equal part to obtain the channel output feature information corresponding to the four parts of the channels. Then, they are re-concatenated into a channel output feature vector, and its exemplary dimension is 2048 dimensions. In this process, by performing average pooling respectively to capture local features of each channel, for audio with large adaptation differences and a large amount of information being submerged by strong noise or other interfering sounds, a feature representation can be established from a few local significant common features.

[0106] In an embodiment of further modification, this step S1400 can be implemented as including the following specific steps:

[0107] Step S1410: Successively perform multi-level feature extraction on the intermediate feature information through multiple convolutional blocks of the local branch network to obtain deep feature information:

[0108] Specifically, two of the convolutional blocks are used. The two convolutional blocks can adopt exactly the same network architecture, and both apply the IBN block, so that the deep feature information obtained after feature extraction can learn the style-invariant features in the song audio data.

[0109] Step S1420: Perform equal division on the deep feature information along the channel direction to obtain multiple equal division feature information:

[0110] It is not difficult to understand that during the process of each convolutional block extracting the feature information corresponding to the encoded information, the output feature information is represented as a matrix structure of "number of channels * number of frequency bands * number of frames", so it contains multiple channels. In this embodiment, for the deep feature information output by the last convolutional block in the local branch network, along the channel direction, this matrix structure is equally divided into four equal parts according to the number of channels, that is, four equal division feature information, so as to perform feature extraction on each part of the channels.

[0111] In an embodiment equivalent to this embodiment, instead of along the channel direction, the feature information output by the last convolutional block can also be equally divided into, for example, four equal parts along the frequency band direction. Thus, the frequency band feature information of each equal part has a significant effect on resisting the frequency band selective attenuation in a harsh pickup environment, balancing the contributions of high and low frequency information in feature composition, resisting the addition or deletion of content in a fixed range of frequency bands (such as adding or reducing a kind of drum sound) or strong interference in a fixed frequency band range.

[0112] Step S1430: Respectively perform average pooling operations on each equal division feature information to obtain multiple equal division feature vectors, and concatenate all the equal division feature vectors into the local output feature vector:

[0113] In order to capture the salient features of each partial channel, each equally divided feature information is average pooled through the pooling layer to obtain multiple equally divided feature vectors. On this basis, all equally divided feature vectors are simply concatenated into a single local output feature vector.

[0114] Step S1500: concatenate the global output feature vector and the channel feature vector into a high-dimensional index vector:

[0115] The output feature vectors obtained by each branch network can be finally converted into a high-dimensional index vector for storage or direct use. The high-dimensional index vector is a high-dimensional vector used to index the corresponding song audio data. Since each branch network has normalized its output feature information into an output feature vector, in this case, depending on the specific purpose of the feature extraction model, the high-dimensional index vector can be processed flexibly. For example, for application requirements that are only for storage and standby and separate calls, each output feature vector can be used as a plurality of corresponding high-dimensional index vectors, which are dispersedly stored in the song feature library, so that the high-dimensional index vectors output by different branch networks can be called on demand for retrieval, query, and matching. For another example, for specific tasks such as song recognition by listening to songs, cover recognition, song recognition by humming, and infringement comparison, all output feature vectors output by all constructed branch networks can be orderly spliced ​​according to the needs of the specific tasks, so as to obtain a single high-dimensional index vector, which can be stored or used for matching immediately. In this embodiment, it is preferred to further simply concatenate multiple output feature vectors into the same high-dimensional index vector. At this point, the representation learning of the song audio data is realized through the high-dimensional index vector, which represents the global significant feature information of the song audio data and also represents its local significant feature information.

[0116] It should be noted that step S1300 and step S1400 are processed in parallel, and the two do not need to rely on each other's data output and input, and parallel tasks can be implemented to improve operating efficiency.

[0117] According to the principles disclosed above in this typical embodiment, a song feature library can be prepared for part or all of the songs in the music library of the online music service platform according to the process of this embodiment. By applying the steps of this embodiment to the song audio data of each corresponding song or its fragments in the music library, the high-dimensional index vectors corresponding to the audio data of each song can be obtained. These high-dimensional index vectors are associated with the corresponding songs and stored to construct a song feature library. Subsequently, the high-dimensional index vector corresponding to any song can be directly called from the song feature library for retrieval, query, matching and other operations.

[0118] Similarly, another application method can be derived, that is, applying the above steps, extracting the corresponding high-dimensional index vectors from the audio data of two songs or their fragments, and then comparing the high-dimensional index vectors of the two songs for similarity, examining the data distance between the two songs, and using a preset threshold to determine whether the two songs are similar. If they are similar, they can be determined to be the same, otherwise they are different. This can be used to determine song infringement or simple matching between two songs.

[0119] In addition to the above application methods, there may be many different uses for mining and utilization based on the high-dimensional index vectors obtained in this application, which can be flexibly applied by technical personnel in this field according to the principles disclosed here without affecting the creative embodiment of this application.

[0120] According to the introduction of this typical embodiment, it can be understood that the implementation of this application has a wealth of advantages, including but not limited to the following aspects:

[0121] First, the application uses the audio information of the song audio data to encode the corresponding encoding information to obtain the style-invariant features of the song audio data, and then extracts intermediate feature information from the encoding information through a shared network. Based on the intermediate feature information, the application uses multiple branch networks to extract the deep semantic information of the song audio data from different angles to obtain the corresponding output feature information. Finally, the output feature information is used as the high-dimensional index vector corresponding to the song audio data to complete the end-to-end representation learning of the song audio data.

[0122] Secondly, since the feature extraction model adopted by the present application uses a combination of a shared network and multiple branch networks to realize multi-angle feature extraction of deep semantic information of song audio data, the obtained high-dimensional index vector can be made more representative. Specifically, not only the deep semantic representation of the global significant features of the song audio data is realized through the global branch network, but also the deep semantic representation of the local significant features of the song audio data is realized through the local branch network, thereby realizing more effective indexing of the corresponding song audio data. On this basis, downstream processing such as retrieval, query, and matching of song audio data can be performed, which can obtain more accurate and efficient matching effects, and can be used in various application scenarios such as cover recognition, song identification, humming recognition, and song infringement determination.

[0123] In addition, this application fuses the output feature vectors obtained by multiple branch networks into a single high-dimensional index vector, which has a wide range of uses and is flexible in usage. It can achieve relatively obvious scale results when processing the representation learning of massive song audio data. It can be deployed in the background of the online music service platform to realize a standardized interface, and then serve the needs of various different application scenarios, provide comprehensive and multi-purpose open services, and enhance the platform's economic advantages in music information retrieval.

[0124] See also Figure 3 In a further embodiment, the convolution block is implemented based on residual convolution and is used to perform the following steps:

[0125] Step S2100: Perform convolution transformation on the input information to obtain transformation feature information:

[0126] For any convolution block in the feature extraction model of the present application, each convolution block first performs a convolution operation on the information input therein, whether it is the encoded information or the intermediate feature information output by the previous convolution block, to obtain the corresponding transformed feature information.

[0127] Step S2200: performing instance normalization and batch normalization processing on the transformed feature information respectively, and then combining them into spliced ​​feature information, and activating and outputting the spliced ​​feature information:

[0128] After the first convolution, an instance batch normalization layer (IN) is applied to process the transformed feature information. The transformed feature information is divided into two paths, and a batch normalization block (BN) is used to batch normalize half of the channels, while the instance normalization layer is applied to the other channels to perform instance normalization, which enables the corresponding convolution block to capture the style-invariant features of the song audio data, thereby achieving better utilization of the song representation with diverse styles in a single data. After different normalization processes, the two parts of the channels can be spliced ​​into the same spliced ​​feature information for activation output.

[0129] Step S2300: The spliced ​​feature information of the activated output is subjected to multiple convolution operations and batch normalization to obtain residual information:

[0130] The spliced ​​feature information of the activated output is further convolved through multiple convolutional layers to further extract features. Each such convolutional layer is followed by a batch normalization layer for normalization and output. The last convolutional layer is implemented with a 1*1 convolution kernel to avoid the attenuation of the representation learning ability of the entire feature extraction model after normalization of multiple instances of multiple convolutional blocks. Therefore, the feature information finally output is the residual information in the residual convolution process.

[0131] Step S2400: superimpose the residual information onto the input information to activate the output:

[0132] Finally, according to the residual convolution principle, referring to the transformed feature information obtained by the first convolution, superimposing the residual information on it and then activating the output, the intermediate feature information output after the residual convolution operation of the current convolution block can be obtained.

[0133] In this embodiment, a convolutional block required for constructing the feature extraction model of the present application is constructed by applying residual convolution combined with instance batch normalization operation. The residual convolution network in it is improved based on the basic model of the Resnet series, and at the same time, the IBN architecture is superimposed. The feature extraction model constructed in this way is easier to train and can achieve a more accurate feature extraction effect, and is particularly suitable for feature extraction of song audio data.

[0134] Please refer to Figure 4 , in an extended embodiment, the training process of the feature extraction model includes the following steps of iterative training:

[0135] Step S3100: Call a training sample from the training set, and determine the encoding information of the audio information of the training sample. The training sample is pre-collected song audio data, and the song audio data is a complete song or a segment thereof:

[0136] Those skilled in the art can understand that to adapt to different downstream tasks, different training sets for training the feature extraction model can be constructed. Each training set contains a sufficient number of training samples, and each training sample is prepared with a corresponding supervision label.

[0137] The training samples can be pre-collected by those skilled in the art. Each training sample is a song audio data. To adapt to different downstream tasks, these song audio data can be complete songs, song MIDI melody segments, songs with accompaniment, a cappella vocal parts of songs, song segments without melody parts, song segments with melody parts, and so on. Different singing versions of the same song can be combined into the same classification, that is, corresponding to the same supervision label, to enhance the generalization ability of model classification. When the duration of the song audio data in the training sample is too long, it can also be segmented into multiple song segments according to a certain duration as multiple training samples associated with the same supervision label for training. When segmenting the song, the time stamps of the lyrics of the song can be referred to for implementation, so that the song segments are segmented based on one or more complete lyrics.

[0138] For the training samples in the training set, for the convenience of model training, the encoding information corresponding to the song audio data can be prepared in advance, or the corresponding encoding information can be obtained in real time when each song audio data is called for training the feature extraction model. As for the specific encoding principle, it can be processed by referring to the corresponding process disclosed in the previous text of the present application.

[0139] Step S3200: Input the encoding information into the feature extraction model for training to obtain corresponding output feature vectors:

[0140] During the training process of a training sample, the corresponding encoded information of the training sample is output to the feature extraction model for feature extraction. For the principle of feature extraction, please refer to the descriptions of the principles of the feature extraction model in the foregoing embodiments, which will not be elaborated here. In this process, the feature extraction model realizes the representation learning of the training sample, and obtains the corresponding output feature vectors of each branch network.

[0141] Step S3300: Perform classification prediction on each of the output feature vectors respectively, so as to map out the corresponding classification labels:

[0142] In this application, the training task of the feature extraction model is understood as a classification task. Therefore, by connecting the output feature vectors of each path of the feature extraction model to the corresponding prepared classification models, examining the classification results of each classification model, and supervising with the corresponding supervision labels, the training of the model can be implemented. Based on this principle, in the training stage, when implementing the training for the feature extraction model implemented in any embodiment of this application, a classification model is connected to the output end of each output feature vector of each branch network of it.

[0143] The classification model adopted is as Figure 5 shown. A batch normalization layer is used to perform batch normalization operation on the output feature vectors, and then through a fully connected layer, the output feature vectors are mapped to the classification space. The classification probabilities corresponding to each classification label are calculated through the classification function, so as to determine that the one with the largest classification probability is the classification label corresponding to the training sample.

[0144] The classifier in the classification model can be constructed by using a multi-classifier implemented by the Softmax function, or can be constructed by using a multi-classifier implemented by the AM-Softmax function that can enhance the intra-class compactness and expand the inter-class sparsity. Obviously, the latter has better classification advantages.

[0145] Step S3400: Calculate the loss value of the feature extraction model by using the supervision label corresponding to the training sample and the classification label, and perform gradient update on the feature extraction model according to the loss value:

[0146] In the classification model, the batch normalization layer is adopted to achieve the balance of the triplet loss and the cross-entropy classification loss. Subsequently, the triplet loss can be calculated for the batch normalization layer, and the cross-entropy classification loss can be calculated for the fully connected layer. By synthesizing these two losses, the optimization of the output feature vectors can be realized.

[0147] Accordingly, after the corresponding classification label of a training sample is predicted, the loss value between the supervision label and the classification label can be calculated according to the corresponding supervision label, and then the feature extraction model can be updated by gradient according to this loss value to correct the weight parameters of each link of the entire model, prompting the model to converge.

[0148] Since there are multiple branch networks, there are multiple outputs of output feature vectors, and correspondingly, multiple classification models need to be connected. Therefore, when calculating the loss value, a weighted method can be adopted, that is, the triplet loss and the classification loss in each classification model are first weighted and summed to obtain the loss value corresponding to each output feature vector, and then the loss values corresponding to each output feature vector are weighted and summed again to obtain the final loss value, and the entire feature extraction model can be updated by gradient according to this loss value.

[0149] Step S3500: Determine whether the loss value reaches a preset threshold. When it does not reach the preset threshold, call the next training sample in the training set to continue the iterative training of the feature extraction model until the feature extraction model exceeds the preset threshold.

[0150] For the loss value calculated for each training sample, determine whether it infinitely approaches 0 or whether it reaches the preset threshold. When these judgment conditions are met, it can be determined that the feature extraction model has been trained to the convergence state. Accordingly, the training of the model can be terminated and the feature extraction model can be put into the production stage, such as for feature extraction of songs in the song library or serving other downstream tasks, etc. If the convergence state is not reached, the next training sample in the training set can be continuously called to continue the iterative training of the feature extraction model until the feature extraction model is trained to the convergence state.

[0151] This embodiment discloses the training principle and process of the feature extraction model of the present application. It can be seen from this embodiment that by training the feature extraction model with the prepared training set, the feature extraction model can learn the ability to extract the corresponding output feature vector from the encoded information of the song audio data, realizing effective representation learning of the deep semantic information of the song audio data. Moreover, the output feature vectors in multiple aspects of the same song audio data can be jointly trained, with higher training efficiency and richer model functions. When it is put into production, the deep semantic information in multiple aspects of the same song audio data can be quickly obtained.

[0152] The classification model of this embodiment uses a multi-classifier with a batch normalization layer and an AM-Softmax function, which can balance the triple loss and classification loss to update the gradient of the model, so that the model can be trained to converge more quickly, and the trained model can better represent the deep semantic information of the song audio data more effectively. When the output feature vector is used for matching later, it can play a more efficient matching role.

[0153] This embodiment also reflects the scalability and compatibility of the feature extraction model of the present application in application. Specifically, this embodiment allows the feature extraction model to be trained by using training samples corresponding to different downstream tasks in order to serve the needs of different downstream tasks. This allows the feature extraction model to obtain the ability to serve different downstream tasks. Therefore, it is a relatively basic improvement with better economic benefits.

[0154] See also Figure 6 , adapted to the need of matching songs in a music library according to a specified song or a song fragment thereof, in an extended embodiment, after the step S1500, the step of concatenating the global output feature vector and the channel feature vector into a high-dimensional index vector, the following steps are included:

[0155] Step S4100: In response to a query request for audio data, the feature extraction model is called to extract a corresponding high-dimensional index vector as a query vector:

[0156] In this embodiment, the feature extraction model of the present application can be used to serve the needs of song identification by listening, cover recognition, and song identification by humming, and receive query requests submitted by users, determine the song provided or specified by the user as query audio data based on the user's query request, and then identify whether the corresponding query audio data is similar to the songs in the song feature library, identify the same song based on the degree of similarity (song identification by listening, song identification by humming), or determine whether it belongs to a certain original song (cover recognition).

[0157] It can be understood that in the song feature library, the feature extraction model of the present application has been used to extract the corresponding high-dimensional index vector for each song audio data in a song library in advance, and the extraction process can be implemented according to steps S1100 to S1500. The high-dimensional index vector can be a high-dimensional vector obtained by splicing multiple output feature vectors, or a high-dimensional vector obtained by combining multiple output feature vectors stored in a dispersed manner as needed. In this embodiment, it is recommended to use a high-dimensional index vector that has been obtained by simply splicing the global output feature vector and the local output feature vector.

[0158] It is not difficult to understand that for the song audio data that is cut into multiple song segments during the feature extraction stage, a corresponding set of high-dimensional index vectors corresponding to its multiple song segments is generated. In this case, each song segment can be regarded as a song. If the duration of the query audio data is too long, it can also be divided into multiple song segments according to a fixed duration for processing, and then the query situations of each song segment can be comprehensively processed.

[0159] In the song feature library, the mapping relationship data between specific songs and their corresponding high-dimensional index vectors is usually stored. Based on this, the summary information of the specific song audio data corresponding to the high-dimensional index vector can be quickly determined subsequently for outputting the result.

[0160] Whether the query audio data is a complete object or is divided into multiple song segment units, it is all input into the feature extraction model of the present application one by one for feature extraction to obtain its corresponding high-dimensional index vector. This high-dimensional index vector is generated corresponding to the organizational form of the high-dimensional index vectors of the song audio data in the song library that has been extracted by the feature extraction model. That is, if the high-dimensional index vector of each song audio data in the song library is an independent one, then the high-dimensional index vector of the query audio data is also a corresponding independent one; if the former is multiple and scattered obtained from different branch networks and different outputs, then the latter is also a corresponding multiple. That is, the end output structure of the feature extraction model of the present application is the same when extracting features from the song library and when extracting features from the query audio data. For the convenience of illustration in this embodiment, it is defined that the high-dimensional index vector is a single one obtained by simply splicing the outputs of multiple branch networks. Accordingly, the high-dimensional index vector extracted for the query audio data constitutes the query vector corresponding to the query audio data.

[0161] Step S4400: Calculate the similarity between the query vector and multiple high-dimensional index vectors in the preset song feature library to obtain a similarity data sequence. The high-dimensional index vectors stored in the song feature library are obtained by the feature extraction model extracting the corresponding song audio data in the preset song library:

[0162] After obtaining the query vector, a preset similarity calculation formula can be used to calculate the similarity between the query vector and the corresponding high-dimensional index vectors of each song in the song feature library. The similarity calculation formula can adopt algorithms such as cosine similarity algorithm, Euclidean distance algorithm, Pearson coefficient algorithm, Jaccard similarity algorithm, nearest neighbor search algorithm, etc., and can be implemented by any algorithm suitable for calculating the similarity distance between data. Those skilled in the art can implement it flexibly. After the similarity calculation, a similarity data sequence of the high-dimensional index vectors of each song in the song feature library corresponding to the query vector is obtained. In the similarity data sequence, the similarity values corresponding to each song in the song feature library and the query vector are stored.

[0163] It should be noted that if the query vector and the high-dimensional index vectors in the song feature library are both in a scattered form, after separate corresponding calculations, multiple similarity data sequences can be obtained. For the convenience of subsequent calculations, the multiple similarity data sequences can be combined into a single similarity data sequence by taking the mean, weighted mean, or simple summation of the similarity values corresponding to the same song.

[0164] Similarly, if the query audio data is segmented into multiple song segments for feature extraction to obtain multiple query vectors, after each query vector is calculated to obtain the single similarity data sequence, further, the single similarity data sequences corresponding to the multiple query vectors can be summarized to obtain the final similarity data sequence of the query audio data. The summarization method used is the same as the previous one, and can be any form such as taking the mean, simple summation, weighted mean, etc.

[0165] Step S4500: Determine the song audio data corresponding to the maximum similarity whose similarity exceeds the preset threshold in the similarity data sequence as the similar song of the query audio data:

[0166] After determining the final similarity data sequence corresponding to the query audio data, a preset threshold can be further used. The preset threshold can be an empirical threshold. Use this preset threshold to filter the similarity data sequence and filter out all elements whose similarity values exceed the preset threshold. If the number of elements exceeding the preset threshold is 0, it indicates that there is no similar song in the song library that is similar to the query audio data. If multiple similarity values are obtained after screening, only the song corresponding to the maximum similarity value can be selected as the similar song corresponding to the query audio data. Thus, the similar song corresponding to the query audio data is matched.

[0167] In this embodiment, the high-dimensional index vector is extracted from the query audio data by means of the feature extraction model of the present application as the query vector, and the similarity between the query vector and the high-dimensional index vectors of each song in the pre-constructed song feature library is calculated. According to the specific similarity value, the corresponding similar songs are matched for the query audio data, and the matching functions of downstream tasks such as song recognition by listening, humming recognition, and cover recognition can be realized. Accordingly, corresponding services can be provided to a large number of users, and the corresponding similar songs can be found according to the specified song or song segment submitted by the user. Since the present application realizes the indexing of songs through the high-dimensional index vectors of songs obtained by the feature extraction model, and the similarity calculation between high-dimensional index vectors is efficient and fast, therefore, when performing song matching, it is very fast, efficient and accurate.

[0168] Please refer to Figure 7 , on the basis of the previous embodiment, in a specific embodiment, before the step S4400 of calculating the similarity between the query vector and multiple high-dimensional index vectors in the preset song feature library, the following steps are included:

[0169] Step S4200: Calculate the similarity between the query vector and the high-dimensional index vectors of each song audio data in the non-melody feature library, and obtain the corresponding similarity sequence. The high-dimensional index vectors in the non-melody feature library are obtained by the feature extraction model extracting each song audio data of non-melody information:

[0170] In the previous embodiment, when calculating the vector similarity, for the case of matching short-segment songs with a duration less than a predetermined length, the situation that the segmented song segments may not have vocal melodies is not considered, and song matching usually mainly considers the matching of vocal melodies. In view of this situation, the present application can achieve filtering by means of pre-constructing a non-melody feature library. The non-melody feature library stores the high-dimensional index vectors corresponding to the non-melody partial song segments segmented from the song audio data in the music library. The high-dimensional index vectors are also pre-extracted by the feature extraction model of the present application, and thus the non-melody feature library can be used to filter the query audio data.

[0171] Specifically, after obtaining the query vector corresponding to the query audio data in step S4100, the similarity between the query vector and the high-dimensional index vectors of all song audio data in the non-melody feature library can be calculated first in the same way as calculating the similarity before, and the corresponding similarity sequence is obtained.

[0172] Step S4300: Compare whether the similarity of each song audio data in the similarity sequence is lower than a preset threshold. When the similarity of all song audio data is lower than the preset threshold, continue with the subsequent steps, otherwise terminate the subsequent steps:

[0173] To determine whether the query vector is similar to any high-dimensional index vector in the non-melodic feature library, those skilled in the art can set a preset threshold corresponding to the similarity based on prior knowledge or measured empirical data, and then compare each similarity value in the similarity sequence calculated in the previous step with this preset threshold. If the similarity values of all elements in the similarity sequence are lower than this preset threshold, it means that there is no non-melodic song segment in the non-melodic feature library that is similar to the query audio data. Accordingly, the steps S4400 and S4500 can be continued to further determine the similar songs. Otherwise, if there is at least one element whose similarity value is higher than the preset threshold, it means that the query audio data is similar to at least one song segment in the non-melodic feature library, and this song segment is obviously a non-melodic segment. At this time, the execution of subsequent steps can be terminated to save the system resource overhead.

[0174] In this embodiment, by calculating the similarity between the query vector corresponding to the query audio data and the high-dimensional index vector in the non-melodic feature library, and then determining whether the query audio data belongs to the non-melodic song segment in the non-melodic feature library according to the preset threshold, the filtering of the query request is realized, which can save the server resource overhead and improve the efficiency of song query, retrieval, and matching. For the case of short segment recognition, since short segments are more likely to contain content without vocal melody, the effectiveness of the recognition process can be significantly improved, and the processing efficiency of the corresponding downstream tasks can be enhanced.

[0175] Please refer to Figure 8 , to meet the need of similarity judgment for two songs or two song segments to determine whether there is an infringement suspicion. In the extended embodiment, after the step S1400, the step of splicing the global output feature vector and the channel feature vector into a high-dimensional index vector, the following steps are included:

[0176] Step S5100: Obtain the query audio data to be compared:

[0177] In this embodiment, the song audio data extracted in the previous embodiments can be used as the original audio data to be compared, or the original version of the song with copyright. According to the previous embodiments, the high-dimensional index vector corresponding to this original version of the song has been obtained by the feature extraction model of the present application.

[0178] Furthermore, another song to be compared can be continuously obtained, and its corresponding query audio data can be obtained. The song corresponding to the query audio data obtained here can be a song submitted or published by a creative user on an online music platform, or an individual target song obtained by the online music platform independently triggered from the target song library to be determined for infringement.

[0179] Step S5200: calling the feature extraction model to determine the high-dimensional index vector corresponding to the query audio data:

[0180] Since it is necessary to make an infringement determination for the query audio data, similarly, the feature extraction model described in this application is used to extract the corresponding high-dimensional index vector for the query audio data in the same manner as described in the aforementioned embodiments. At this point, the high-dimensional index vector corresponding to the original song and the high-dimensional index vector corresponding to the query audio data are obtained. Combined with the previous description of the correspondence between the organizational forms of the two high-dimensional index vectors used to calculate similarity, it can be seen that the organizational forms of the two high-dimensional index vectors here are also the same.

[0181] Step S5300: Calculate the similarity between the high-dimensional index vector of the query audio data and the high-dimensional index vector of the song audio data:

[0182] Referring to the previous embodiment, the similarity calculation formula corresponding to any similarity algorithm, such as the cosine similarity algorithm, the Euclidean distance algorithm, the Pearson coefficient algorithm, the Jaccard similarity algorithm, the nearest neighbor search algorithm, etc., is used to calculate the similarity between the high-dimensional index vector of the query audio data (also called the query vector) and the high-dimensional index vector of the original song, and the corresponding similarity value can be obtained.

[0183] Step S5400: determine whether the similarity exceeds a preset threshold, and when it exceeds the preset threshold, determine that the query audio data and the song audio data constitute similar songs:

[0184] After obtaining the similarity value between the two songs, a preset threshold value determined by a person skilled in the art based on empirical data or prior knowledge is used to determine whether the two songs are similar songs. Specifically, the similarity value is compared to see whether it exceeds the preset threshold value. If it exceeds the preset threshold value, it means that the query audio data and the audio data of the song as the original song constitute similar songs, and thus it can be determined that the query audio data constitutes an infringing work. On the contrary, if it does not exceed the preset threshold value, it means that the two songs do not constitute similar songs, and no special processing is required, or the relevant users can be simply informed.

[0185] This embodiment exemplarily illustrates the process of using the feature extraction model of the present application to compare two songs to determine whether the two are suspected of infringement. It can be seen that since the feature extraction model of the present application has the ability to more accurately represent the characteristic information of related songs, it can make more accurate judgments when comparing songs for infringement and obtain corresponding infringement judgment results, which helps online service platforms to quickly check various song audio data, stop infringements in a timely manner, and prevent platform infringement risks.

[0186] Please refer to Figure 9 , a song semantic information indexing device provided by the present application is functionally deployed according to the song semantic information indexing method of the present application, including: an encoding processing module 1100, a shared extraction module 1200, a global extraction module 1300, a local extraction module 1400, and an indexing processing module 1500. Among them, the encoding processing module 1100 is used to encode the audio information in the song audio data to obtain corresponding encoded information; the shared extraction module is used to perform multi-level feature extraction on the encoded information in sequence by multiple convolutional blocks in the shared network of the feature extraction model that has been trained to the convergence state, so as to obtain intermediate feature information that extracts the deep semantic information of the song audio data; the global extraction module is used to extract global significant features from the intermediate feature information by using the global branch network of the feature extraction model to obtain a global output feature vector; the local extraction module is used to extract semantic local features from the intermediate feature information by equal segmentation according to channels respectively by using the local branch network of the feature extraction model to obtain a channel output feature vector; the indexing processing module is used to splice the global output feature vector and the channel feature vector into a high-dimensional index vector.

[0187] In a further embodiment, the global extraction module 1300 includes: a global convolution branch module, which is used to perform multi-level feature extraction on the intermediate feature information in sequence by multiple convolutional blocks of the global branch network to obtain deep feature information; a global pooling operation module, which is used to perform a maximum pooling operation on the deep feature information to obtain a global output feature vector that extracts the global significant features of the song audio data.

[0188] In a further embodiment, the local extraction module 1400 includes: a local convolution branch module, which is used to perform multi-level feature extraction on the intermediate feature information in sequence by multiple convolutional blocks of the local branch network to obtain deep feature information; a channel segmentation output module, which is used to equally divide the deep feature information along the channel direction to obtain a plurality of equally divided feature information; a local pooling operation module, which is used to perform an average pooling operation on each of the equally divided feature information respectively to obtain a plurality of equally divided feature vectors, and splice all the equally divided feature vectors into the local output feature vector.

[0189] In an optional embodiment, in the encoding processing module 1100, the audio information is any one of the time-frequency spectrum information, Mel spectrum information, CQT filtering information, pitch contour information, and Chroma feature information of the song audio data.

[0190] In a specific embodiment, in the shared network, at least one of the convolutional blocks applies an attention module to extract key information from the song audio data, and the attention module is a spatial attention module or a channel attention module.

[0191] In a further embodiment, the convolutional block is configured to include the following units: a convolutional transformation unit for performing a convolutional transformation on the information input thereto to obtain transformed feature information; a normalization processing unit for respectively performing instance normalization and batch normalization on the transformed feature information and then combining them into spliced feature information, and activating and outputting the spliced feature information; an intermediate convolutional unit for obtaining residual information after performing multiple convolutional operations and batch normalization on the activated and output spliced feature information; and a residual processing unit for superimposing the residual information on the information input thereto and activating and outputting.

[0192] In an extended embodiment, a training task architecture with the following structure is used to train the feature extraction model: a sample calling module for calling a training sample from a training set and determining the encoded information of the audio information of the training sample, where the training sample is pre-collected song audio data, and the song audio data is a complete song or a segment thereof; a training execution module for inputting the encoded information into the feature extraction model to train it to obtain corresponding output feature vectors; a classification prediction module for respectively performing classification prediction on each of the output feature vectors to map out corresponding classification labels; a gradient update module for calculating the loss value of the feature extraction model using the supervision label corresponding to the training sample and the classification label, and performing gradient update on the feature extraction model according to the loss value; and a loop iteration module for determining whether the loss value reaches a preset threshold. When the preset threshold is not reached, the next training sample in the training set is called to continue iteratively training the feature extraction model until the loss value reaches the preset threshold.

[0193] In an extended embodiment, the song semantic information indexing device further includes: a request response module for responding to a query request for query audio data, and calling the feature extraction model to extract its corresponding high-dimensional index vector as a query vector; a similarity calculation module for calculating the similarity between the query vector and multiple high-dimensional index vectors in a preset song feature library to obtain a similarity data sequence, where the high-dimensional index vectors stored in the song feature library are obtained by the feature extraction model extracting the corresponding song audio data in a preset song library; and a similarity determination module for determining the song audio data corresponding to the maximum similarity in the similarity data sequence whose similarity exceeds the preset threshold as the similar song of the query audio data.

[0194] In a specific embodiment, the song semantic information indexing device further includes: a non-melody similarity calculation module, configured to calculate the similarity between the query vector and the high-dimensional index vectors of each song audio data in the non-melody feature library, to obtain a corresponding similarity sequence, where the high-dimensional index vectors in the non-melody feature library are obtained by the feature extraction model extracting from each song audio data of the non-melody information; a non-melody filtering and processing module, configured to compare whether the similarity of each song audio data in the similarity sequence is lower than a preset threshold, and when the similarity of all song audio data is lower than the preset threshold, continue to run the similarity calculation module, otherwise terminate the operation of other modules.

[0195] In an extended embodiment, the song semantic information indexing device further includes: a query acquisition module, configured to obtain query audio data to be compared; a query indexing module, configured to call the feature extraction model to determine the high-dimensional index vector corresponding to the query audio data; a query calculation module, configured to calculate the similarity between the high-dimensional index vector of the query audio data and the high-dimensional index vector of the song audio data; a similarity comparison module, configured to determine whether the similarity exceeds a preset threshold, and when it exceeds the preset threshold, determine that the query audio data and the song audio data constitute similar songs.

[0196] To solve the above technical problems, an embodiment of the present application further provides a computer device. As Figure 10 shown, it is a schematic internal structure diagram of the computer device. The computer device includes a processor, a computer-readable storage medium, a memory, and a network interface connected through a system bus. Among them, the computer-readable storage medium of the computer device stores an operating system, a database, and computer-readable instructions. The database can store a control information sequence. When the computer-readable instructions are executed by the processor, the processor can implement a song semantic information indexing method. The processor of the computer device is used to provide computing and control capabilities to support the operation of the entire computer device. The memory of the computer device can store computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor can execute the song semantic information indexing method of the present application. The network interface of the computer device is used to connect and communicate with the terminal. Those skilled in the art can understand that Figure 10 the structure shown in

[0197] In this embodiment, the processor is used to execute Figure 9The specific functions of each module and its sub-modules therein, and the memory stores program codes and various types of data required to execute the above modules or sub-modules. The network interface is used for data transmission between user terminals or servers. In this embodiment, the memory stores the program codes and data required to execute all modules / sub-modules in the song semantic information indexing device of the present application, and the server can call the program codes and data of the server to execute the functions of all sub-modules.

[0198] The present application also provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the steps of the song semantic information indexing method according to any embodiment of the present application.

[0199] The present application also provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by one or more processors, the steps of the method according to any embodiment of the present application are implemented.

[0200] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments of the present application can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium can be a computer-readable storage medium such as a magnetic disk, an optical disc, a read-only memory (ROM), or a random access memory (RAM), etc.

[0201] In summary, by improving the feature extraction model of song audio data, the present application enhances the representation learning ability of the deep semantic information of song audio data. Based on the obtained deep semantic information, it can achieve a more accurate and efficient effect when serving song query, retrieval, and matching, and can serve various downstream tasks such as song recognition by listening, humming recognition, cover version recognition, and song infringement comparison, thereby enhancing the comprehensive service ability of the online music platform.

[0202] Those skilled in the art of this technology can understand that the steps, measures, and solutions in the various operations, methods, and processes discussed in the present application can be alternated, changed, combined, or deleted. Further, the other steps, measures, and solutions in the various operations, methods, and processes discussed in the present application can also be alternated, changed, rearranged, decomposed, combined, or deleted. Further, the steps, measures, and solutions in the prior art that are the same as those disclosed in the various operations, methods, and processes in the present application can also be alternated, changed, rearranged, decomposed, combined, or deleted.

[0203] The above are only some embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A method for indexing song semantic information, characterized in that, The method includes the following steps: Encode the audio information in the song audio data to obtain corresponding encoded information; Use multiple convolutional blocks in the shared network of the feature extraction model that has been trained to a convergent state to sequentially perform multi-level feature extraction on the encoded information, so as to obtain intermediate feature information that extracts the deep semantic information of the song audio data. The feature extraction model consists of a shared network and two branch networks. The shared network includes multiple convolutional blocks for gradually extracting the deep semantic information of the encoded information to obtain intermediate feature information. The two branch networks respectively perform extraction of different types of deep semantic information based on the intermediate feature information to obtain corresponding output feature information; Use the global branch network of the feature extraction model to extract global significant features from the intermediate feature information to obtain a global output feature vector. The global branch network, after sequentially performing feature extraction on the intermediate feature information through two convolutional blocks with the same structure as other branch networks, extracts the global significant feature information through a maximum pooling operation on the output of the last convolutional block, and correspondingly outputs a global output feature vector; Use the local branch network of the feature extraction model to separately extract semantic local features from the intermediate feature information by channel equal division to obtain a channel output feature vector. The local branch network, after sequentially performing feature extraction on the intermediate feature information through two convolutional blocks with the same structure as the global branch network, divides the output of the last convolutional block by the number of channels * number of frequency bands * number of frames, divides it equally by channel dimension and outputs, then performs average pooling on each equal division respectively to obtain channel corresponding channel output feature information, and then re-concatenates it into a channel output feature vector; Concatenate the global output feature vector and the channel feature vector into a high-dimensional index vector.

2. The method for indexing song semantic information according to claim 1, characterized in that, The step of using the global branch network of the feature extraction model to extract global significant features from the intermediate feature information to obtain a global output feature vector includes the following steps: Sequentially perform multi-level feature extraction on the intermediate feature information through multiple convolutional blocks of the global branch network to obtain deep feature information; Perform a maximum pooling operation on the deep feature information to obtain a global output feature vector that extracts the global significant features of the song audio data.

3. The method for indexing song semantic information according to claim 1, characterized in that, The step of using the local branch network of the feature extraction model to separately extract semantic local features from the intermediate feature information by channel equal division to obtain a channel output feature vector includes the following steps: Sequentially perform multi-level feature extraction on the intermediate feature information through multiple convolutional blocks of the local branch network to obtain deep feature information; Perform equal division on the deep feature information along the channel direction to obtain multiple equal division feature information; Perform average pooling operations on each equal division feature information respectively to obtain multiple equal division feature vectors, and concatenate all the equal division feature vectors into the local output feature vector.

4. The method for indexing song semantic information according to claim 1, characterized in that, In the step of encoding the audio information in the song audio data, the audio information is any one of the time-frequency spectrum information, Mel spectrum information, CQT filter information, pitch contour information, and Chroma feature information of the song audio data.

5. The method for indexing song semantic information according to claim 1, characterized in that, In the shared network, at least one of the convolutional blocks applies an attention module to extract key information in the song audio data, and the attention module is a spatial attention module or a channel attention module.

6. The method for indexing song semantic information according to claim 1, characterized in that, The convolutional block is used to perform the following steps: Perform a convolutional transformation on the information input thereto to obtain transformed feature information; Perform instance normalization and batch normalization on the transformed feature information respectively and then combine them into concatenated feature information, and activate and output the concatenated feature information; Obtain residual information after performing multiple convolutional operations and batch normalization on the activated and output concatenated feature information; Superimpose the residual information on the information input thereto and activate and output.

7. The method for indexing song semantic information according to any one of claims 1 to 6, characterized in that, The training process of the feature extraction model includes the following steps of iterative training: Call a training sample from the training set, and determine the encoded information of the audio information of the training sample. The training sample is pre-collected song audio data, and the song audio data is a complete song or a segment thereof; Input the encoded information into the feature extraction model to train it to obtain corresponding output feature vectors; Perform classification prediction on each of the output feature vectors respectively to map out corresponding classification labels; Calculate the loss value of the feature extraction model by using the supervision label corresponding to the training sample and the classification label, and perform gradient update on the feature extraction model according to the loss value; Judge whether the loss value reaches a preset threshold. When it does not reach the preset threshold, call the next training sample in the training set to continue to perform iterative training on the feature extraction model until the loss value reaches the preset threshold.

8. The method for indexing song semantic information according to any one of claims 1 to 6, characterized in that, After the step of fusing the global output feature vector and the channel feature vector into a high-dimensional index vector, the following steps are included: In response to a query request for query audio data, call the feature extraction model to extract its corresponding high-dimensional index vector as a query vector; Calculate the similarity between the query vector and multiple high-dimensional index vectors in a preset song feature library to obtain a similarity data sequence. The high-dimensional index vectors stored in the song feature library are obtained by the feature extraction model extracting the corresponding song audio data in a preset song library; Determine the song audio data corresponding to the maximum similarity whose similarity in the similarity data sequence exceeds the preset threshold as the similar song of the query audio data.

9. The method for indexing song semantic information according to claim 8, wherein, Before the step of calculating the similarity between the query vector and multiple high-dimensional index vectors in a preset song feature library, the following steps are included: Calculate the similarity between the query vector and the high-dimensional index vectors of each song audio data in a non-melody feature library to obtain a corresponding similarity sequence. The high-dimensional index vectors in the non-melody feature library are obtained by the feature extraction model extracting each song audio data of non-melody information. Compare whether the similarity of each song audio data in the similarity sequence is lower than a preset threshold. When the similarity of all song audio data is lower than the preset threshold, continue with the subsequent steps; otherwise, terminate the subsequent steps.

10. The method for indexing song semantic information according to any one of claims 1 to 6, wherein, After the step of concatenating the global output feature vector and the channel feature vector into a high-dimensional index vector, the following steps are included: Obtain query audio data to be compared; Call the feature extraction model to determine the high-dimensional index vector corresponding to the query audio data; Calculate the similarity between the high-dimensional index vector of the query audio data and the high-dimensional index vector of the song audio data; Judge whether the similarity exceeds a preset threshold. When it exceeds the preset threshold, determine that the query audio data and the song audio data constitute similar songs.

11. A device for indexing song semantic information, wherein, Include: An encoding processing module for encoding the audio information in the song audio data to obtain corresponding encoded information; A shared extraction module for successively performing multi-level feature extraction on the encoded information by using multiple convolutional blocks in the shared network of the feature extraction model that has been trained to a convergent state, to obtain intermediate feature information that extracts the deep semantic information of the song audio data. The feature extraction model consists of a shared network and two branch networks. The shared network includes multiple convolutional blocks for successively extracting the deep semantic information of the encoded information to obtain intermediate feature information. The two branch networks respectively perform extraction of different types of deep semantic information based on the intermediate feature information to obtain corresponding output feature information; A global extraction module for using the global branch network of the feature extraction model to extract global significant features from the intermediate feature information to obtain a global output feature vector. The global branch network, after successively performing feature extraction on the intermediate feature information through two convolutional blocks with the same structure as other branch networks, extracts the global significant feature information through a maximum pooling operation on the output of the last convolutional block, and correspondingly outputs the global output feature vector; A local extraction module for using the local branch network of the feature extraction model to separately extract semantic local features from the intermediate feature information by channel equal segmentation to obtain a channel output feature vector. The local branch network, after successively performing feature extraction on the intermediate feature information through two convolutional blocks with the same structure as the global branch network, divides the output of the last convolutional block by the number of channels * number of frequency bands * number of frames, and outputs it equally by channel dimension segmentation, and then performs average pooling for each equal segment to obtain the channel output feature information corresponding to the channel, and then re-concatenates it into a channel output feature vector; An index processing module for concatenating the global output feature vector and the channel feature vector into a high-dimensional index vector.

12. A computer device, comprising a central processing unit and a memory, wherein, The central processing unit is used to call and run the computer program stored in the memory to execute the steps of the method according to any one of claims 1 to 10.

13. A computer-readable storage medium, wherein, It stores in the form of computer-readable instructions a computer program implemented according to the method according to any one of claims 1 to 10. When the computer program is called and run by a computer, it executes the steps included in the corresponding method.

14. A computer program product, comprising a computer program / instructions, wherein, When the computer program / instructions are executed by a processor, the steps of the method described in any one of claims 1 to 10 are implemented.

Citation Information

Patent Citations

  • Multi-modal emotion recognition method based on attention enhancing mechanism

    CN112489635A

  • Training method and system for obtaining better speech translation model in generative adversarial

    CN113505611A