Song comparison method and its device, equipment, medium, product
By using pre-trained feature extraction model to extract multi-scale deep semantic information of songs, high-dimensional index vectors are obtained, which solves the problem of low accuracy of cover recognition in the prior art, and achieves more efficient cover recognition and copyright comparison.
Patent Information
- Application Number
- CN202111491601.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-08
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-12-08
AI Technical Summary
The existing technology has low accuracy and low recognition efficiency in cover recognition, especially in the application scenarios of copyright infringement comparison, and cannot effectively identify differences in musical components such as tone, fundamental frequency, rhythm, speed, harmony, lyrics, singing methods, etc.
A feature extraction model pre-trained to the convergent state is used to extract the original song and the multi-scale deep semantic information of the song being compared, obtain a high-dimensional index vector, and determine whether the two constitute a cover relationship through similarity calculation.
The copyright infringement comparison and similar song identification are realized, the accuracy and efficiency of identification are improved, and various types of changes in cover songs can be more accurately identified to avoid misjudgment.
Smart Images

Figure CN114817620B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of music information retrieval, and in particular, to a song comparison method, a corresponding device, a computer device, a computer-readable storage medium, and a computer program product. Background Art
[0002] The song cover recognition technology has a wide range of applications. In exemplary applications, for the need of copyright protection, it is often necessary to compare the original song with the candidate song to identify whether there is sufficient similarity between the two, so as to determine whether the candidate song infringes the copyright of the original song.
[0003] At present, with the popularity of short videos, live broadcasts, and radio stations, the number of cover music is increasing, and the scenarios requiring music recognition are becoming more and more complex. Compared with the original version, the cover version may have differences or even be completely different in music components such as timbre, fundamental frequency, rhythm, speed, harmony, lyrics, singing method, and overall structure. Therefore, the cover recognition technology that can accurately identify the cover relationship between songs is a very challenging research work.
[0004] There are various existing technologies related to cover recognition. Each existing technology has certain deficiencies to some extent. For example: (1) The traditional landmark-based song recognition technology can only recognize songs of the same source version and cannot recognize the cover versions with certain differential information as described above; (2) The traditional melody matching-based humming recognition technology can only recognize clean a cappella / humming and cannot recognize the cover versions with background accompaniment as described above; (3) The traditional cover recognition technology solutions mainly extract audio features such as Pitch Class Profile (PCP), and then use algorithms such as dynamic programming to calculate the similarity distance between songs. Due to the diversity of cover versions, the above solutions can only be applied to cover solutions with minor adaptations, with low accurate recognition rate, slow recognition speed, and limited compatibility.
[0005] In view of the lack of general adaptability, low recognition accuracy, and low recognition efficiency of the existing technology solutions for cover recognition, especially being powerless in the application scenario of copyright infringement comparison, the applicant of the present application attempts to explore a more effective technical solution. Summary of the Invention
[0006] The primary object of the present application is to solve at least one of the above problems and provide a song comparison method, a corresponding device, a computer device, a computer-readable storage medium, and a computer program product.
[0007] To meet the various objects of the present application, the present application adopts the following technical solutions:
[0008] A song comparison method provided to meet one of the purposes of the present application includes the following steps:
[0009] Obtain the audio data of the original song and the song to be compared respectively;
[0010] Use the feature extraction model trained to the convergence state to extract the multi-scale deep semantic information of the audio data of the original song, and correspondingly obtain the original high-dimensional index vector;
[0011] Use the feature extraction model to extract the multi-scale deep semantic information of the audio data of the song to be compared, and correspondingly obtain the compared high-dimensional index vector;
[0012] Calculate the similarity between the original high-dimensional index vector and the compared high-dimensional index vector, and determine whether the corresponding similarity value is greater than the first preset threshold. When it is greater than the first preset threshold, it is determined that the song to be compared and the original song form a cover relationship.
[0013] In an extended embodiment, in the step of calculating the similarity between the original high-dimensional index vector and the compared high-dimensional index vector and determining whether the corresponding similarity value is greater than the first preset threshold, when it is less than the preset threshold, the following steps are executed:
[0014] Determine the song with the relatively shorter audio duration among the original song and the song to be compared as the specified song, and segment the other song with the relatively longer audio duration with the audio duration of the specified song as the measure to obtain multiple song segments of the other song;
[0015] Use the feature extraction model to extract the multi-scale deep semantic information of the multiple song segments of the other song respectively, and correspondingly obtain multiple segment high-dimensional index vectors;
[0016] Take the high-dimensional index vector of the specified song as the specified high-dimensional index vector, calculate the similarity between the specified high-dimensional index vector and each segment high-dimensional index vector respectively, and determine whether the maximum similarity value is higher than the second preset threshold. When it is higher than the second preset threshold, it is determined that the song to be compared forms a cover relationship.
[0017] In a further embodiment, obtaining the audio data of the original song and the song to be compared includes the following steps:
[0018] Obtain the audio data and the lyrics file of the original song;
[0019] Segment the lyrics in the lyrics file of the original song, and extract multiple keywords from them;
[0020] Search online for at least one song according to any combination of the keywords, and use the searched song as the song to be compared;
[0021] Obtain the audio data of the song to be compared.
[0022] In an alternative embodiment, in the step of extracting the multi-scale deep semantic information of the audio data of the original song by using the feature extraction model trained to the convergence state and correspondingly obtaining the original high-dimensional index vector, or in the step of extracting the multi-scale deep semantic information of the audio data of the song to be compared by using the feature extraction model and correspondingly obtaining the compared high-dimensional index vector, the following steps are included:
[0023] Encode the audio data to correspondingly obtain its encoded information;
[0024] Use the feature extraction model to extract a high-dimensional index vector representing the multi-scale deep semantic information of the audio data according to the encoded information.
[0025] In a further embodiment, when the feature extraction model is called, the following steps are executed:
[0026] Use multiple convolutional blocks in the shared network of the feature extraction model trained to the convergence state to sequentially perform multi-level feature extraction on the encoded information to obtain intermediate feature information extracting the deep semantic information of the encoded information;
[0027] Use multiple convolutional blocks in more than two branch networks of the feature extraction model to perform feature extraction on the intermediate feature information at different scales and then convert it into output feature vectors of corresponding scales, and the deep semantic information contained in the output feature vectors of each branch network is different;
[0028] Output the output feature vectors of each branch network as the high-dimensional index vector by the feature extraction model.
[0029] In a further embodiment, after using multiple convolutional blocks in more than two branch networks of the feature extraction model to perform feature extraction on the intermediate feature information at different scales and then convert it into output feature vectors of corresponding scales, it includes any two or more of the following steps:
[0030] Use multiple convolutional blocks in the first branch network to perform feature extraction on the intermediate feature information to obtain global feature information, and pool the global feature information into an output feature vector at the global scale;
[0031] Use multiple convolutional blocks in the second branch network to perform feature extraction on the intermediate feature information and then split it into multiple parts by channels for pooling to correspondingly obtain an output feature vector at the channel scale;
[0032] Use multiple convolutional blocks in the third branch network to perform feature extraction on the intermediate feature information and then split it into multiple parts by frequency bands for pooling to correspondingly obtain an output feature vector at the frequency band scale.
[0033] In a specific embodiment, when the first branch network performs the pooling operation, average pooling and / or maximum pooling operations are adopted to correspondingly obtain one or two output feature vectors of the global scale; when the second branch network performs the pooling operation, average pooling operation is adopted for a single or multiple channels to correspondingly obtain one or more output feature vectors of the channel scale; when the third branch network performs the pooling operation, average pooling operation is adopted for a single or multiple frequency bands to correspondingly obtain one or more output feature vectors of the frequency band scale.
[0034] In an alternative embodiment, the source of the encoded information is any one of the time-frequency spectrum information, Mel spectrum information, CQT filtering information, pitch contour information, and Chroma feature information of the corresponding audio data.
[0035] In a preferred embodiment, in the shared network, at least one of the convolutional blocks applies an attention module to extract key information from the song audio data, and the attention module is a spatial attention module or a channel attention module.
[0036] In a specific embodiment, when the convolutional block is called, the following steps are executed:
[0037] Perform a convolutional transformation on the information input thereto to obtain transformed feature information;
[0038] Perform instance normalization and batch normalization on the transformed feature information respectively and then combine them into spliced feature information, and activate and output the spliced feature information;
[0039] Perform multiple convolutional operations and batch normalization on the activated and output spliced feature information to obtain residual information;
[0040] Superimpose the residual information on the information input thereto and activate and output.
[0041] In an extended embodiment, the training process of the feature extraction model includes the following steps of iterative training:
[0042] Call a training sample from the training set and determine the encoded information of the training sample. The training sample is pre-collected song audio data, and the song audio data is a complete song and its segments;
[0043] Input the encoded information into the feature extraction model to train it to obtain corresponding output feature vectors;
[0044] Perform classification prediction on each of the output feature vectors respectively to map out corresponding classification labels;
[0045] Calculate the loss value of the feature extraction model using the supervision label corresponding to the training sample and the classification label, and perform gradient update on the feature extraction model according to the loss value;
[0046] Determine whether the loss value reaches a preset threshold. When it does not reach the preset threshold, call the next training sample in the training set to continue iterative training on the feature extraction model until the loss value reaches the preset threshold.
[0047] A song comparison device provided to meet one of the purposes of the present application includes: an audio acquisition module, an original version extraction module, a comparison extraction module, and a comprehensive judgment module. Among them, the audio acquisition module is used to acquire the audio data of the original song and the song to be compared respectively; the original version extraction module is used to extract the multi-scale deep semantic information of the audio data of the original song by using a feature extraction model that has been trained to a convergent state, and correspondingly obtain an original high-dimensional index vector; the comparison extraction module is used to extract the multi-scale deep semantic information of the audio data of the song to be compared by using the feature extraction model, and correspondingly obtain a comparison high-dimensional index vector; the comprehensive judgment module is used to calculate the similarity between the original high-dimensional index vector and the comparison high-dimensional index vector, and determine whether the corresponding similarity value is greater than a first preset threshold. When it is greater than the first preset threshold, it is determined that the song to be compared and the original song form a cover relationship.
[0048] A computer device provided to meet one of the purposes of the present application includes a central processing unit and a memory. The central processing unit is used to call and run a computer program stored in the memory to execute the steps of the song comparison method described in the present application.
[0049] A computer-readable storage medium provided to meet another purpose of the present application stores a computer program implemented according to the song comparison method in the form of computer-readable instructions. When the computer program is called and run by a computer, it executes the steps included in the method.
[0050] A computer program product provided to meet another purpose of the present application includes a computer program / instructions. When the computer program / instructions are executed by a processor, they implement the steps of the method described in any embodiment of the present application.
[0051] Compared with the prior art, the advantages of the present application are as follows:
[0052] First, this application uses a feature extraction model pre-trained to a convergent state to obtain high-dimensional index vectors representing the deep semantic information of the original song and the song to be compared, which characterize their invariant styles. By determining the similarity between the two high-dimensional index vectors, it can be determined whether the original song and the song to be compared form a cover relationship, thus enabling functions such as song copyright infringement comparison and similar song identification. Since each high-dimensional index vector is extracted using the same feature extraction model and realizes the deep semantic representation of the corresponding audio data of the song at different semantic scales, the similarity between songs can be accurately determined based on semantics.
[0053] Secondly, since this application realizes multi-scale feature extraction of the deep semantic information of song audio data in the feature extraction model it uses, the obtained high-dimensional index vectors can be made more representative. For example, they can represent the global feature information, significant feature information, channel feature information, frequency band feature information, etc. of the song audio data, thus realizing a more effective indexing of the corresponding song audio data. Based on this, song comparison is performed, which is more compatible with various types of changes in cover songs and can obtain a more accurate and efficient matching effect, avoiding misjudgment in the recognition process.
[0054] In addition, when this application is based on the end-to-end representation learning ability and supplemented with a comparison mechanism, it can achieve obvious scale effects. It can be deployed in the background of an online music service platform to implement a standardized interface, and then serve the needs of various different application scenarios, providing comprehensive and multi-purpose open services, and enhancing the economic advantage of music information retrieval on the platform. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] The above and / or additional aspects and advantages of this application will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, where:
[0056] Figure 1 is a schematic flowchart of a typical embodiment of the song comparison method of this application;
[0057] Figure 2 is a schematic flowchart of an extended embodiment of the song comparison method of this application;
[0058] Figure 3 is a schematic flowchart of the process of obtaining the song to be compared based on the original song in an embodiment of this application;
[0059] Figure 4 is a schematic flowchart of the process of obtaining a high-dimensional index vector based on audio data in an embodiment of this application;
[0060] Figure 5 is a schematic flowchart of the process of running the feature extraction model in an embodiment of this application;
[0061] Figure 6 Schematic diagram of the network architecture of the feature extraction model in an embodiment of the present application;
[0062] Figure 7 Schematic diagram of the network architecture of the feature extraction model in another embodiment of the present application;
[0063] Figure 8 Schematic flow chart showing the working process of the residual convolution block adopted in the feature extraction model of the present application;
[0064] Figure 9 Schematic flow chart of the process of implementing and training the feature extraction model of the present application;
[0065] Figure 10 Principle block diagram of the classification model connected in the training stage of the feature extraction model of the present application;
[0066] Figure 11 Principle block diagram of the song comparison device of the present application;
[0067] Figure 12 Schematic diagram of the structure of a computer device adopted by the present application. Detailed implementation manners
[0068] The embodiments of the present application are described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described by referring to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be construed as a limitation to the present application.
[0069] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present application means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.
[0070] Those skilled in the art can understand that, unless otherwise defined, all terms used herein (including technical terms and scientific terms) have the same meaning as the general understanding of those of ordinary skill in the art to which this application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted in an idealized or overly formal sense unless specifically defined as herein.
[0071] Those skilled in the art can understand that the "client", "terminal", and "terminal device" used herein include both devices with a wireless signal receiver that only has the ability to receive and no ability to transmit, and devices with receiving and transmitting hardware that can perform two-way communication on a two-way communication link. Such devices can include: cellular or other communication devices such as personal computers, tablets, etc., which have a single-line display or a multi-line display or a cellular or other communication device without a multi-line display; PCS (Personal Communications Service), which can combine voice, data processing, fax, and / or data communication capabilities; PDA (Personal Digital Assistant), which can include a radio frequency receiver, pager, Internet / intranet access, web browser, notepad, calendar, and / or GPS (Global Positioning System) receiver; conventional laptop and / or palm-top computers or other devices, which are conventional laptop and / or palm-top computers or other devices with and / or including a radio frequency receiver. The "client", "terminal", and "terminal device" used herein can be portable, transportable, installed in a vehicle (air, sea, and / or land), or suitable for and / or configured to operate locally and / or in a distributed manner at any other location on the earth and / or in space. The "client", "terminal", and "terminal device" used herein can also be a communication terminal, an Internet access terminal, a music / video playing terminal, such as a PDA, MID (Mobile Internet Device), and / or a mobile phone with music / video playing functions, or can also be devices such as a smart TV, a set-top box, etc.
[0072] The hardware referred to by names such as "server", "client", and "service node" in this application is essentially an electronic device with the equivalent capabilities of a personal computer. It is a hardware device with the necessary components revealed by the von Neumann principle, including a central processing unit (including an arithmetic unit and a controller), a memory, an input device, and an output device. The computer program is stored in its memory, and the central processing unit loads the program stored in the external memory into the internal memory for execution, executes the instructions in the program, and interacts with the input / output devices to complete specific functions.
[0073] It should be noted that the concept of "server" in this application can similarly be extended to apply to server clusters. According to the network deployment principles understood by those skilled in the art, the various servers should be logically divided. Physically, these servers can either be independent of each other but can be invoked through interfaces, or integrated into a single physical computer or a set of computer clusters. Those skilled in the art should understand this flexibility and should not use it to restrict the implementation manner of the network deployment method of this application.
[0074] One or several technical features of this application, unless expressly specified, can either be deployed on the server and accessed by the client remotely invoking the online service interface provided by the server, or directly deployed and run on the client for access.
[0075] The neural network models cited or possibly cited in this application, unless expressly specified, can either be deployed on a remote server and remotely invoked on the client, or deployed on a client with sufficient device capabilities for direct invocation. In some embodiments, when it runs on the client, its corresponding intelligence can be obtained through transfer learning to reduce the requirements for the client's hardware operating resources and avoid excessive occupation of the client's hardware operating resources.
[0076] All kinds of data involved in this application, unless expressly specified, can either be remotely stored on the server or stored on the local terminal device, as long as it is suitable for being invoked by the technical solution of this application.
[0077] Those skilled in the art should be aware that although the various methods of this application are described based on the same concept and thus show commonality with each other, unless otherwise specified, these methods can all be executed independently. Similarly, for the various embodiments disclosed in this application, they are all proposed based on the same inventive concept. Therefore, concepts with the same expression, as well as concepts that are only appropriately transformed for convenience although the concept expressions are different, should be equivalently understood.
[0078] For each embodiment to be disclosed in this application, unless explicitly stated as mutually exclusive, the relevant technical features involved in each embodiment can be cross - combined to flexibly construct new embodiments, as long as such combination does not deviate from the creative spirit of this application and can meet the needs in the prior art or solve certain deficiencies in the prior art. Those skilled in the art should be aware of this flexibility.
[0079] A song comparison method of this application can be programmed as a computer program product and deployed to run in a computer device. It can be implemented as a front - end product of a client device or as an online service product. Thus, the client can access the interface opened after the computer program product runs in the form of a web program or an application program, and achieve human - computer interaction with the process of the computer program product through a graphical user interface.
[0080] Please refer to Figure 1 , in a typical embodiment of the song comparison method of this application, the following steps are included:
[0081] Step S1100: Obtain the audio data of the original song and the song to be compared respectively:
[0082] In an exemplary application scenario, the technical solution of this application is responsible for comparing the similarity between the original song and other songs, and the other songs are the songs to be compared. Therefore, it is necessary to determine the original song and the songs to be compared in advance.
[0083] In one embodiment, both the original song and the song to be compared can be specified or provided by the user; in another embodiment, the song to be compared can be obtained through associated retrieval according to information associated with the original song, such as its lyric file.
[0084] After determining the specified information of the original song and the song to be compared, the corresponding audio data can be obtained. The audio data can be either local data or remote data that can be pulled to the local.
[0085] When multiple songs to be compared are provided, each step of this application can be executed for each song to be compared one by one to achieve the purpose of comparing the original song with each song to be compared separately.
[0086] The songs referred to in this application, including the original song and the songs to be compared, are not limited in file format, including but not limited to any format such as MP3, WMA, M4A, WAV, etc. Of course, it can also be audio data obtained by separating the audio from various video files. Those skilled in the art should be able to understand this flexibility.
[0087] For the original song and the song to be compared, in order to compare them with each other, it is necessary to extract their multi-scale deep semantic information by means of step S1200 and step S1300 respectively, as described below.
[0088] Step S1200: Use the feature extraction model that has been trained to the convergence state to extract the multi-scale deep semantic information of the audio data of the original song, and correspondingly obtain the original high-dimensional index vector:
[0089] The feature extraction model for extracting the deep semantic information of songs implemented in this application based on the convolutional neural network model is pre-trained to the convergence state. After training, it is enabled to learn the ability to extract the deep semantic information of multiple scales of the audio data of songs according to the encoded information, and to realize the representation learning of the style-invariant features of the corresponding audio data of songs, so that it can be used for requirements such as query, retrieval, and matching between songs.
[0090] The feature extraction model of this application is implemented to be suitable for extracting the deep semantic information of multiple scales of the same audio data, and representing these deep semantic information as single or multiple high-dimensional index vectors, so as to realize the feature representation of the audio data from multiple different aspects and / or different angles. The high-dimensional index vector is essentially a high-dimensional vector, which plays an indexing and representative role for the encoded information of the corresponding audio data at the semantic level. These different scales include the global scale based on the encoded information, or the frequency band scale, channel scale, etc. based on the encoded information for feature extraction. For a song, selecting the deep semantic information of two or more arbitrary scales corresponding to its encoded information and representing it as a high-dimensional index vector can realize the feature representation of the multi-scale deep semantic information of the corresponding song.
[0091] After the feature extraction model implemented according to the above principle is trained to convergence, the service interface can be opened for the technical solution of this embodiment to call. Input the encoded information of the original song to it, and the feature extraction model performs feature extraction on the basis of this encoded information to obtain the corresponding high-dimensional index vector. As for the principle and process of converting the audio data into the corresponding encoded information, it can be flexibly set by those skilled in the art, and any corresponding encoded information such as the time-frequency spectrum information, Mel spectrum information, CQT filtering information, pitch contour information, Chroma feature information, etc. of the audio data can be realized. The recommended encoding principle and process will be further revealed in other embodiments of this application later, and will not be elaborated here for the time being.
[0092] It should be understood that since the feature extraction model can extract the deep semantic information of a song from multiple scales, when converting the deep semantic information of these different scales into the high-dimensional index vectors described above, there can be different organizational forms. For example, the high-dimensional index vector can be represented as a single high-dimensional vector, and generally this single high-dimensional vector as a whole represents the deep semantic information of a song; or, the high-dimensional index vector can be represented as multiple discrete high-dimensional vectors according to the scale correspondence relationship, and each high-dimensional vector corresponds to one scale. In any case, those skilled in the art can flexibly organize these high-dimensional vectors according to the need for actual scale semantic information, with the criterion of facilitating the invocation of the representation data of the overall deep semantic information of the song.
[0093] For this step, after the feature extraction model extracts the features from the encoded information of the original song, the corresponding high-dimensional index vector of the original song can finally be obtained. As the original high-dimensional index vector, it can be used for subsequent similarity matching.
[0094] Step S1300: Use the feature extraction model to extract the multi-scale deep semantic information of the audio data of the song to be compared, and correspondingly obtain the high-dimensional index vector to be compared:
[0095] Similar to the previous step of extracting the multi-scale deep semantic information of the audio data of the original song, this step is parallel to the previous step. Still call the feature extraction model that has been trained to the convergence state to extract the multi-scale deep semantic information from the encoded information of the audio data of the song to be compared, and obtain the corresponding high-dimensional index vector as the high-dimensional index vector to be compared. As for the encoding principle and process related to the encoded information obtained here, it will be further revealed in the embodiments below by the same token.
[0096] It should be emphasized that step S1200 and step S1300 can be implemented in parallel to improve the overall comparison speed.
[0097] Step S1400: Calculate the similarity between the original high-dimensional index vector and the high-dimensional index vector to be compared, and determine whether the corresponding similarity value is greater than the first preset threshold. When it is greater than the first preset threshold, it is determined that the song to be compared and the original song form a cover relationship:
[0098] Based on the original high-dimensional index vector corresponding to the original song and the compared high-dimensional index vector corresponding to the song to be compared, a preset similarity calculation formula can be applied to calculate the similarity, so as to calculate the similarity value between the original song and the song to be compared. The similarity calculation formula can adopt algorithms such as cosine similarity algorithm, Euclidean distance algorithm, Pearson coefficient algorithm, Jaccard similarity algorithm, nearest neighbor search algorithm, etc., and can be implemented by any algorithm suitable for calculating the similarity distance between data, and those skilled in the art can implement it flexibly. After the similarity calculation, the similarity value between the original high-dimensional index vector of the original song and the compared high-dimensional index vector of the song to be compared is obtained.
[0099] After determining the similarity value, a preset threshold, hereinafter referred to as the first preset threshold, can be further used. The first preset threshold can be an empirical threshold or an experimental threshold. Compare the similarity value with the first preset threshold. If the similarity value is greater than (including equal to) the first preset threshold, it indicates that the original song and the song to be compared are sufficiently similar songs. Therefore, it can be determined that the two constitute a cover relationship. In the scenario of song copyright infringement comparison, it can be determined accordingly that the song to be compared constitutes a suspected infringement of the original song. On the contrary, if the similarity value is less than the first preset threshold, it indicates that the original song and the song to be compared are not sufficiently similar, and it can be simply determined that the two do not constitute a cover relationship. In the scenario of song copyright infringement comparison, it can be determined accordingly that the song to be compared does not constitute a suspected infringement of the original song. Of course, the relatively complex situation caused by malicious cover processing increasing the difficulty of cover identification is not considered here. For this, more refined judgments will be realized through other embodiments later, which will not be elaborated here for the time being.
[0100] So far, based on the given original song and the song to be compared, the present application has realized the identification of whether the two constitute a cover relationship and can output the comparison result.
[0101] In other embodiments that will be successively disclosed in the present application later, various situations based on the changes of this typical embodiment will be further disclosed, which will not be elaborated here for the time being. Only based on the introduction of this typical embodiment, it can be understood that the implementation of the present application has rich advantages, including but not limited to the following aspects:
[0102] First, the present application uses a feature extraction model pre-trained to a converged state to obtain high-dimensional index vectors representing the deep semantic information of the original song and the song to be compared, which characterize their invariant styles. By determining the similarity between the two high-dimensional index vectors, it can be determined whether the original song and the song to be compared form a cover relationship, thus enabling functions such as song copyright infringement comparison and similar song identification. Since each high-dimensional index vector is extracted using the same feature extraction model and realizes the deep semantic representation of the corresponding audio data of the song at different semantic scales, the similarity between songs can be accurately determined based on semantics.
[0103] Secondly, since the present application realizes multi-scale feature extraction of the deep semantic information of song audio data in the feature extraction model it uses, the obtained high-dimensional index vectors can be made more representative. For example, they can represent the global feature information, significant feature information, channel feature information, frequency band feature information, etc. of the song audio data, thus realizing a more effective indexing of the corresponding song audio data. Based on this, song comparison is performed, which is more compatible with various types of changes in cover songs and can obtain a more accurate and efficient matching effect, avoiding misjudgment in the recognition process.
[0104] In addition, when the present application is based on the end-to-end representation learning ability and supplemented by a comparison mechanism, it can achieve obvious scale effects. It can be deployed in the background of an online music service platform to implement a standardized interface, and then serve the needs of various different application scenarios, providing an integrated and multi-purpose open service, and enhancing the economic advantage of music information retrieval on the platform.
[0105] Considering some situations where cover song recognition is difficult. For example, the song to be compared is composed of multiple local song segments that are far apart in time from the original song. In this case, if only a single high-dimensional index vector is used to represent the deep semantic information of the song to be compared and a single high-dimensional index vector is used to represent the deep semantic information of the original song, it often leads to a large difference between the two, and accurate recognition cannot be simply achieved by directly calculating the similarity between the two high-dimensional index vectors. Conversely, the same is true when several short-duration song segments of the original song are inserted into the longer song to be compared.
[0106] Considering this situation, it is necessary to further deepen the technical solution of the present application. Therefore, please refer to Figure 2 In the extended embodiment, in the step S1400 of calculating the similarity between the original high-dimensional index vector and the high-dimensional index vector to be compared and determining whether the corresponding similarity value is greater than the first preset threshold, when it is less than the preset threshold, the following steps are executed:
[0107] Step S1500: Determine the song with the relatively shorter audio duration among the original song and the song to be compared as the designated song, and segment the other song with the relatively longer audio duration based on the audio duration of the designated song to obtain multiple song segments of the other song:
[0108] When the audio durations of the original song and the song to be compared are inconsistent, one of them with the relatively shorter audio duration can be selected as the designated song (for example, the original song). Then, based on the audio duration of the designated song, segment the other song (for example, the song to be compared), and divide the other song into multiple song segments according to this audio duration. It can be understood that each song segment will obtain a corresponding audio data. Similarly, corresponding encoding can be performed according to the encoding principle and process of the present application to obtain corresponding encoding information for subsequent processing.
[0109] Step S1600: Use the feature extraction model to extract the multi-scale deep semantic information of multiple song segments of the other song respectively, and correspondingly obtain multiple segment high-dimensional index vectors:
[0110] The high-dimensional index vector of the designated song has been extracted in the previous steps. Therefore, in this step, only the feature extraction of the encoding information of the audio data corresponding to multiple song segments of the other song needs to be performed. For this, the feature extraction model of the present application can still be used to extract the multi-scale deep semantic information of each of the song segments, and correspondingly obtain the high-dimensional index vectors of each song segment as segment high-dimensional index vectors.
[0111] Step S1700: Use the high-dimensional index vector of the designated song as the designated high-dimensional index vector, calculate the similarity between the designated high-dimensional index vector and each segment high-dimensional index vector respectively, and determine whether the maximum similarity value is higher than the second preset threshold. When it is higher than the second preset threshold, determine that the song to be compared constitutes a cover relationship:
[0112] The high-dimensional index vector of the designated song, here called the designated high-dimensional index vector, has been determined in advance. In order to compare it with the segment high-dimensional index vectors of each song segment, the similarity between the designated high-dimensional index vector and each of the segment high-dimensional index vectors can be calculated respectively to obtain the corresponding similarity values, thereby obtaining a similarity sequence.
[0113] Similarly, when calculating the similarity, a preset similarity calculation formula can be applied to calculate the similarity, so as to calculate the similarity values between the specified song and each song segment. The similarity calculation formula can adopt algorithms such as cosine similarity algorithm, Euclidean distance algorithm, Pearson coefficient algorithm, Jaccard similarity algorithm, nearest neighbor search algorithm, etc., and can be implemented by any algorithm suitable for calculating the similarity distance between data, and those skilled in the art can implement it flexibly. After the similarity calculation, the similarity values between the specified high-dimensional index vector of the specified song and the segment high-dimensional index vectors of each song segment are obtained.
[0114] After determining the similarity value, a preset threshold, hereinafter referred to as the second preset threshold, can be further used. The second preset threshold can be an empirical threshold or an experimental threshold. Compare the maximum similarity value in the similarity sequence with the second preset threshold. If the maximum similarity value is higher than (including equal to) the second preset threshold, it indicates that the specified song and at least one song segment constitute a sufficiently similar song. Therefore, it can be determined that the original song and the compared song constitute a cover relationship. In the scenario of song copyright infringement comparison, it can be determined accordingly that the compared song constitutes a suspected infringement of the original song. On the contrary, if the maximum similarity value is lower than the second preset threshold, it indicates that the specified song and all song segments do not constitute sufficient similarity, and it can be determined that the original song and the compared song do not constitute a cover relationship. In the scenario of song copyright infringement comparison, it can be determined accordingly that the compared song does not constitute a suspected infringement of the original song.
[0115] The second preset threshold and the first preset threshold of this application are generally set separately. In a simplified embodiment, the two can also be equal-valued data to simplify the setting.
[0116] In this embodiment, in response to the situation where the original high-dimensional index vector of the original song and the compared high-dimensional index vector of the compared song are determined not to constitute a cover relationship, the one with the shorter audio duration is further determined as the specified song, and the other song is segmented into multiple song segments, and then the similarity is calculated one by one between the high-dimensional index vector of the specified song and the high-dimensional index vectors of each song segment. According to the similarity value, it is judged whether each song segment constitutes a cover relationship with the specified song, realizing a more refined cover relationship recognition. For the situation where the cover content is relatively discrete, it can better achieve compatible recognition, thereby improving the recognition accuracy. Especially in the scenario of song infringement comparison, it can more quickly and efficiently determine suspected infringement products.
[0117] To facilitate the architecture of online services and realize the automatic online monitoring of infringement behaviors based on the original song, please refer to Figure 3 In a further embodiment, the step S1100, obtaining the audio data of the original song and the compared song, includes the following steps:
[0118] Step S1110: Obtain the audio data and lyrics file of the original song:
[0119] The original song user can entrust an online service platform to search for suspected infringing products of their original song, and can provide the original song to the online service platform through pre-configuration, real-time submission, etc. In any case, the corresponding computer device of the present application can obtain the audio data and its lyrics file corresponding to the original song from the obtained storage address in a controlled or timed manner.
[0120] Step S1120: Segment the lyrics in the lyrics file of the original song and extract multiple keywords from them:
[0121] The content stored in the lyrics file can be collectively referred to as the lyrics in a broad sense. The lyrics in a broad sense mainly include the song name, composer information, lyricist information, multiple lines of lyrics corresponding to the melody, and the time stamps corresponding to each line of lyrics, etc. It can be seen that the text information contained in the lyrics file is rich. Accordingly, the present application can segment the lyrics in the lyrics file with the keyword extraction model implemented by natural language technology, and then extract multiple keywords from them. The technologies for implementing the keyword extraction model include multiple technical approaches such as statistical features, word graphs, and topic models. Exemplary algorithms include TF-IDF, TextRank, etc., and a pre-trained neural network model such as Bert, Electra, etc. can also be used to implement it. In this regard, those skilled in the art can call flexibly. Generally speaking, those skilled in the art can extract multiple keywords from the lyrics file of the original song with the help of a variety of existing technologies.
[0122] Step S1130: Online search for at least one song according to any combination of the keywords, and use the searched song as the song to be compared:
[0123] When multiple keywords are obtained through keyword extraction, multiple search expressions can be obtained by arbitrarily combining the multiple keywords. For example, when {keyword_1, keyword_2, keyword_3} exists, six search expressions can be combined according to the principle of permutation and combination. On this basis, call the interface provided by the online search engine, input any one or any number of the search expressions, and conduct an online search to obtain the corresponding search results. Based on the search results, those skilled in the art can know the possible songs in the search results after page analysis with various mature technical means, and these songs can be regarded as the songs to be compared in the present application for subsequent processing.
[0124] In an improved and simplified embodiment, an online service platform can legally access a music library held by itself or by a third party with an open interface. With the corresponding online search engine of the music library, database operations can be performed on the music library, and the compared songs can be retrieved more quickly and conveniently.
[0125] Step S1140: Obtain the audio data of the compared song:
[0126] After determining each of the compared songs, the audio data of each compared song can be further downloaded or copied for subsequent processing in this application.
[0127] In this embodiment, the lyrics file of the original song can be used as the keyword source. By means of online keyword search, various online songs can be monitored, a comparison range can be determined, and then the cover relationship of the songs within this range can be further identified with the help of this application, which is particularly helpful for realizing intelligent monitoring of song copyright infringement and achieving the effect of purifying the online music ecological environment.
[0128] Please refer to Figure 4 , in an optional embodiment, in the step S1200 of extracting the multi-scale deep semantic information of the audio data of the original song by using a feature extraction model trained to a convergent state and correspondingly obtaining the original high-dimensional index vector, and / or, in the step S1300 of extracting the multi-scale deep semantic information of the audio data of the compared song by using the feature extraction model and correspondingly obtaining the compared high-dimensional index vector, the following steps are included:
[0129] Step S0001: Encode the audio data to obtain its encoding information accordingly:
[0130] Before using the feature extraction model to extract the deep semantic information of various audio data in this application, the audio data needs to be encoded accordingly to obtain its corresponding encoding information, and the deep semantic information is extracted based on the encoding information.
[0131] As mentioned above, the audio data can be audio data in any format such as MP3, WMA, M4A, WAV, etc., or audio data obtained by separating the audio from various video files. The audio data usually consists of multiple voice data packets in the time domain. On this basis, by performing corresponding transformation processing on the voice data packets according to the specific type of encoding information, the corresponding encoding information can be obtained.
[0132] The encoded information mainly serves as information related to describing the style-invariant features in audio data, and there can be various types, including but not limited to spectrogram information, Mel spectrogram information, CQT filtering information, pitch contour information, Chroma feature information, etc., which are extracted from the voice data packets of the audio data. Such information can be encoded using corresponding algorithms to obtain the encoded information of the corresponding type. In this application, any one of the above types of encoded information can be used to implement feature extraction in this application. In practice, it is recommended to encode with the CQT filtering information that has been measured to be optimal to obtain the encoded information.
[0133] Those skilled in the art understand that the above various encoded information can be encoded using corresponding algorithms. During the encoding process, it is necessary to first perform conventional processing such as pre-emphasis, framing, and windowing on the audio data, and then perform time-domain or frequency-domain analysis, that is, to implement speech signal analysis. The purpose of pre-emphasis is to enhance the high-frequency part of the speech signal and make the spectrum smooth; generally, pre-emphasis is achieved through a first-order high-pass filter. Before analyzing the speech signal, it is also necessary to frame it. Usually, the length of each frame of the speech signal is set to 20 ms. Considering the frame shift factor, there can be an overlap of 10 ms between adjacent frames. To implement framing, windowing operations can be performed on the speech signal. Different window selections will affect the results of speech signal analysis. More commonly, the window function corresponding to the Hamming window (Hamm) is used to perform the windowing operation.
[0134] On the basis of completing the preprocessing required for the speech signal analysis of the song audio data, time-domain and frequency-domain analysis can be further performed on it to achieve encoding and obtain the corresponding encoded information:
[0135] For the spectrogram information, by performing pre-emphasis, framing, windowing, and short-time Fourier transform (STFT) on the speech data of each voice data packet in the time domain to transform it into the frequency domain, the data corresponding to the spectrogram is obtained, thereby constituting the spectrogram information.
[0136] The Mel spectrogram information can be obtained by filtering the spectrogram information using a Mel-scale filter bank. Similarly, by taking the logarithm of the Mel spectrogram information and performing a DCT transform, the corresponding Mel cepstrum information can also be obtained. It can be understood that the Mel spectrogram information and its Mel cepstrum information can better describe the style-invariant features in the song, such as pitch, intonation, timbre, etc.
[0137] For the CQT filtering information, since in music, all tones are composed of the 12-tone equal temperament of several octaves, that is, the 12-tone equal temperament, corresponding to the twelve semitones on one octave of a piano. The frequency ratio between adjacent semitones is 2 1 / 12Obviously, for two octaves of the same pitch class, the higher octave has a frequency that is twice that of the lower octave. Therefore, in music, sounds are exponentially distributed, but the audio spectra obtained by Fourier transform are linearly distributed, and the frequency points of the two cannot correspond one by one, which will cause errors in the estimated values of certain scale frequencies. Therefore, the CQT time-frequency transformation algorithm can be used to replace the Fourier transform method for speech analysis. CQT, Constant Q Transform, that is, the constant Q transform, refers to a filter bank in which the center frequencies are distributed according to an exponential law, the filter bandwidths are different, but the ratio of the center frequency to the bandwidth is a constant Q. Different from the Fourier transform, the horizontal axis frequency of its spectrum is not linear, but based on log2, and the filter window length can be changed according to the different spectral line frequencies to obtain better performance. Since the distribution of CQT is the same as that of the scale frequencies, by calculating the CQT spectrum of the music signal, the amplitude values of the music signal at the frequencies of each note can be directly obtained, which is more perfect for the signal processing of music. Therefore, this embodiment recommends using this information to perform corresponding encoding to obtain corresponding encoding information as the input of the neural network model of this application.
[0138] The pitch class profile information mentioned above, including PCP (Pitch Class Profile), HPCP (Harmonic Pitch Class Profile) are both acceptable, aiming to extract the corresponding fundamental tone sequence from the song audio data, and after regularization, merging, and segmentation, it is transformed into a melody profile sequence, and then transformed into the corresponding feature representation using the standard pitch difference generated by the standard pitch. The encoding information constructed based on the pitch class profile information has good robustness to environmental noise.
[0139] The Chroma feature information mentioned above is a general term for the chroma vector (Chroma Vector) and the chromagram (Chromagram). The chroma vector is a vector containing 12 elements, and these elements respectively represent the energy of the 12 pitch classes within a period of time (such as 1 frame), and the energy of the same pitch class in different octaves is accumulated. The chromagram is a sequence of chroma vectors. Specifically, after performing a short-time Fourier transform on the speech data packet of the song audio data to transform it from the time domain to the frequency domain, some noise reduction processing is performed, and then tuning is carried out; the absolute time is converted into frames according to the length of the selected window, and the energy of each pitch within each frame is recorded to form a pitch spectrogram; based on the pitch spectrogram, the energy (in terms of loudness) of the notes of the same time, the same pitch class, and different octaves is superimposed on the element of this pitch class in the chroma vector to form a chromagram. The data corresponding to this chromagram is the Chroma feature information mentioned above.
[0140] Any of the above specific coding information can be used as the input of the feature extraction model of the present application. For the convenience of processing by the feature extraction model, the coding information can be organized according to a certain preset format. For example, the coding information corresponding to each voice packet is organized into a row vector, and for the entire audio data, the row vectors of its respective voice data packets are organized row by row in time sequence to obtain a two-dimensional matrix as its complete coding information. And so on, which can be preset to adapt to the feature extraction model and can be flexibly implemented by those skilled in the art.
[0141] It should be noted that the coding principle described here is applicable not only to the original song, the compared song, and the song segment of the compared song, but also to the processing of training samples by the feature extraction model during the training stage, which should be understood by those skilled in the art.
[0142] Step S0002: Use the feature extraction model to extract a high-dimensional index vector representing the multi-scale deep semantic information of the audio data according to the coding information.
[0143] As mentioned above, the feature extraction model of the present application is pre-trained until convergence before being put into use. When it is necessary to obtain the high-dimensional index vector of the audio data, the feature extraction model of the present application can be called accordingly to process the coding information of the audio data, which will not be elaborated here.
[0144] This embodiment further introduces the principle and process of a preferred variety of coding information in the present application, which is more convenient for the implementation of the present application. Among them, it is recommended to use CQT feature information as the coding information, and through actual measurement, the beneficial effects of the present application can be further demonstrated.
[0145] Please refer to Figure 5 , in the in-depth embodiment, when the feature extraction model is called, the following steps are executed:
[0146] Step S2100: Use multiple convolutional blocks in the shared network of the feature extraction model that has been trained to a convergent state to sequentially perform multi-level feature extraction on the coding information, and obtain intermediate feature information that extracts the deep semantic information of the coding information.
[0147] The feature extraction model is constructed based on the multi-branch idea in the present application and can be flexibly deformed to meet the requirements of different embodiments of the present application. In a typical embodiment of the feature extraction model, such as Figure 6As shown in the principle block diagram, the feature extraction model is composed of a shared network and multiple branch networks. Among them, the shared network includes multiple convolutional blocks for gradually extracting deep semantic information of encoded information to obtain intermediate feature information; the multiple branch networks respectively perform extraction of different types of deep semantic information based on the intermediate feature information to obtain corresponding output feature information. Among the multiple branch networks, each contains a part of the same structure, which contains multiple convolutional blocks for gradually extracting deep semantic information. After the output of the last convolutional block, different processing can be performed according to the different functions of each branch network.
[0148] The convolutional block can be implemented by a convolutional layer based on CNN and RNN, and preferably a convolutional block based on the residual convolution principle. In order to achieve the role of context grooming to extract key information in the song audio data, an attention mechanism can be applied in any one of the convolutional blocks, and a corresponding attention module is added, specifically a spatial attention module (Spatial Attention Module, SAM) or a channel attention module (Channel Attention Module, CAM). In an enhanced embodiment, instance normalization operation (IN) and batch normalization (BN) operations are applied in the convolutional block to divide the information input into it into two parts. One part performs instance normalization operation to learn style-invariant features, and the other part performs batch normalization operation to achieve normalization. Therefore, the so-called IBN architecture is applied. By applying this architecture, music attribute invariant features with highly diverse styles of song audio data can be learned, such as notes, rhythm, timbre, etc., while retaining the version information.
[0149] Accordingly, it is not difficult to understand that the feature extraction model adapts to different application scenarios, enables different branch networks, and uses a preselected training set to train it to a convergent state first, and then the corresponding feature extraction ability can be obtained, so as to be suitable for performing tasks corresponding to the application scenarios, and extract the output feature information corresponding to the song audio data from the encoded information of the song audio data input into it. Regarding the training process of the feature extraction model, it will be given in the exemplary embodiments of this application and will not be elaborated here for the time being.
[0150] In this step, in the architecture as Figure 6 shown, after the encoded information is gradually subjected to feature extraction by multiple convolutional blocks of the shared network, especially after key information extraction by the last convolutional block among them, intermediate feature information that extracts the key information of the encoded information can be obtained. This intermediate feature information is divided into multiple paths and output to the multiple branch networks, so as to perform extraction of deep semantic information from different angles in each branch network.
[0151] Step S2200: After performing feature extraction on the intermediate feature information at different scales using multiple convolutional blocks in two or more branch networks of the feature extraction model, convert it into output feature vectors at corresponding scales. The deep semantic information contained in the output feature vectors of each branch network is different:
[0152] As described above, Figure 6 In the architecture shown, each branch network can be flexibly selected and combined. Therefore, according to the specific architecture obtained by the combination, it is possible to determine how many branch networks there are specifically. The intermediate feature information output by the shared network is respectively input into each of the branch networks for further feature extraction processing.
[0153] According to Figure 6 the architecture shown, each branch network has the same structural part, which includes two convolutional blocks. After the two convolutional blocks sequentially perform feature extraction on the feature information output therein, the output after extraction can be processed differently according to the specific structure of different branch networks.
[0154] Specifically, different branch networks can perform different processes in their non-identical structural parts according to the different deep semantic information they extract. For example: maximum pooling or average pooling output can be performed on one of the branch networks; a Dropout layer can be connected to one of the branch networks to randomly discard redundant features therein and then perform maximum pooling output; in another branch network, the intermediate feature information output by the last convolutional block can be equally channel-segmented and then average-pooled respectively for output; in another branch network, the intermediate feature information output by the last convolutional block can be equally frequency-band segmented and then average-pooled respectively for output, and so on. By performing various different processes on the feature information output by the last convolutional block, output feature information containing different deep semantic information can be obtained. These output feature information respectively describe the deep semantic information of the song audio data from different scales, including the global information and various local information of the song audio data, such as the global information abstracting the significant features of the coding information of the song audio data, the local information abstracting the channel or frequency-band features of the coding information of the song audio data, and so on. Accordingly, multiple output feature information with different representations can be obtained, and these output feature information can be independently called or arbitrarily combined as needed.
[0155] In this application, the output feature information output by each branch network is normalized into the representation of an output feature vector. Therefore, multiple branch networks can correspondingly obtain multiple output feature vectors. Each output feature vector represents the deep semantic information of the song audio data in different aspects or at different scales, and the deep semantic information contained in each output feature vector is different from each other.
[0156] In use, usually two or more branch networks are adopted to obtain two or more output feature vectors, so as to use two or more deep semantic information to perform feature representation on the song audio data. For example, the output feature vector for representing the global information of the song audio data can be combined with the output feature vector for representing the channel information of the song audio data for use, or the output feature vector for representing the global information of the song audio data can be combined with the output feature vector for representing the band information of the song audio data for use, or the output feature vector for representing the channel information of the song audio data can be combined with the output feature vector for representing the band information of the song audio data for use, or all output feature vectors can be combined for use. And so on, which can be called by those skilled in the art as needed.
[0157] Step S2300: The feature extraction model outputs the output feature vectors of each branch network as the high-dimensional index vectors described above:
[0158] The output feature vectors obtained by each branch network can be finally converted into high-dimensional index vectors for storage or directly used. The high-dimensional index vector is a high-dimensional vector used to index the corresponding song audio data. Since each branch network has normalized its output feature information into an output feature vector, in this case, the high-dimensional index vector can be flexibly processed according to the specific use of the feature extraction model. For example, for the application requirements of only storing for backup and calling separately, each output feature vector can be used as multiple corresponding high-dimensional index vectors. Another example is that for specific tasks such as exemplary song copyright infringement comparison, all output feature vectors output by all the already constructed branch networks can be orderly spliced according to the needs of the specific task to obtain a single high-dimensional index vector, and this high-dimensional index vector can be stored or immediately used for matching. Thus, the representation learning of the song audio data is realized through the high-dimensional index vector.
[0159] In addition to the various application methods disclosed in this application, there may be various different uses for the mining and utilization based on the high-dimensional index vectors obtained from this application, which can be flexibly used by those skilled in the art according to the principles disclosed herein, and none of them will affect the manifestation of the creativity of this application.
[0160] Through the above introduction of the execution process and network architecture of the feature extraction model, it can be understood that this embodiment includes very rich beneficial effects, including but not limited to the following aspects:
[0161] First, the feature extraction model uses the audio information of the song audio data to encode the corresponding coding information to obtain the style invariant features of the song audio data, and then extracts intermediate feature information from the coding information through a shared network. Based on the intermediate feature information, the deep semantic information of the song audio data is extracted from different angles through multiple branch networks to obtain the corresponding output feature information. Finally, these output feature information are used as the high-dimensional index vector corresponding to the song audio data to complete the end-to-end representation learning of the song audio data.
[0162] Secondly, since the feature extraction model adopts a combination of a shared network and multiple branch networks to realize multi-angle feature extraction of deep semantic information of song audio data, the obtained high-dimensional index vector can be more representative, such as representing the global feature information, significant feature information, channel feature information, frequency band feature information, etc. of the song audio data, thereby achieving more effective indexing of the corresponding song audio data, and performing downstream processing such as retrieval, query, and matching of song audio data on this basis, more accurate and efficient matching effects can be obtained, which can be used in a variety of application scenarios such as cover recognition, song identification, humming recognition, and song infringement determination.
[0163] In addition, the output feature vectors obtained by the multiple branch networks of the feature extraction model can be combined into a single high-dimensional index vector or used independently as different high-dimensional index vectors. They can be flexibly determined according to the required deep semantic information. They have a wide range of uses and are flexible in usage. When processing representation learning of massive song audio data, they can achieve relatively obvious scale results. They can be deployed in the background of the online music service platform to realize a standardized interface, and then serve the needs of a variety of different application scenarios, provide comprehensive and multi-purpose open services, and enhance the platform's economic advantages in music information retrieval.
[0164] In a further embodiment, the step S2200, using multiple convolution blocks in more than two branch networks in the feature extraction model to extract features of different scales on the intermediate feature information and converting them into output feature vectors of corresponding scales, includes any two or more of the following steps:
[0165] Step S2210: Use multiple convolution blocks in the first branch network to extract features from the intermediate feature information to obtain global feature information, and pool the global feature information into an output feature vector of global scale:
[0166] Figure 6In the first branch network given as an example, after the intermediate feature information is subjected to feature extraction step by step through two convolutional blocks with the same structure as other branch networks, the output of the last convolutional block is divided into two paths. One path directly performs average pooling operation to obtain its overall feature information, and the other path randomly discards some time-frequency region information through a Dropout layer, and then extracts the significant feature information in the global through a max pooling operation. Thus, two global output feature vectors are correspondingly output. According to this architecture, in the model training stage, on the one hand, the generalization ability of the model to audio with local time-frequency domain changes such as segment deletion and segment insertion in song audio data is improved, and on the other hand, it also plays a role in preventing the model from overfitting to a certain extent. In addition, one of the two global output feature vectors captures the overall feature and the other captures the significant feature, improving the recognition ability of the model.
[0167] Step S2220: After using multiple convolutional blocks in the second branch network to perform feature extraction on the intermediate feature information, it is segmented into multiple parts according to channels for pooling, and the output feature vectors at the channel scale are correspondingly obtained:
[0168] Since the feature information output by each convolutional block is usually represented as "number of channels * number of frequency bands * number of frames", it can be segmented according to the number of channels. Figure 6 In the second branch network given as an example, after the intermediate feature information is subjected to feature extraction step by step through two convolutional blocks with the same structure as other branch networks, the output of the last convolutional block is segmented according to channels, divided into multiple paths such as two paths for output, and then respectively passed through a 1*1 convolutional layer and subjected to average pooling to obtain the channel output feature information corresponding to the two parts of the channels. In this process, the two channel branches focus on capturing the local features of the audio. For audio with large adaptation differences and a large amount of information submerged by strong noise or other interfering sounds, feature representations can be established from a few local significant common features.
[0169] Step S2230: After using multiple convolutional blocks in the third branch network to perform feature extraction on the intermediate feature information, it is segmented into multiple parts according to frequency bands for pooling, and the output feature vectors at the frequency band scale are correspondingly obtained:
[0170] Figure 6In the third branch network given by way of example, after the intermediate feature information is subjected to feature extraction step by step through two convolutional blocks having the same structure as other branch networks, the output of the last convolutional block is subjected to average pooling and then segmented by frequency band, divided into multiple paths, for example, two paths of output. After average pooling, band output feature information corresponding to two parts of the frequency band is obtained. In this process, each band branch focuses on extracting the feature information of the corresponding frequency band, and has a significant effect on resisting the frequency band selective attenuation in a harsh pickup environment, balancing the contributions of high and low frequency information in feature composition, and resisting the addition or deletion of content in a fixed range frequency band (such as adding or reducing a kind of drum sound) or strong interference in a fixed frequency band range.
[0171] It can be understood that multiple output feature vectors obtained in the same branch network can be further processed into the same output feature vector through concatenation or average pooling. In this regard, those skilled in the art can implement it flexibly.
[0172] In this embodiment, rich branch networks are used to extract multi-scale feature information from song audio data, so that the obtained output feature vectors can obtain rich deep semantic information representations, which represent both the global information and significant information of the song audio data, and also represent the relevant local information of the song audio data by channels and frequency bands. Considering that the intermediate feature information has captured the key information of the song audio data under the action of the shared network, therefore, this embodiment realizes the indexing value of the song audio data from multiple aspects. When the subsequent obtained high-dimensional index vector is used for querying, retrieving, and matching, the accuracy in all aspects can be improved.
[0173] Since this embodiment can capture the deep semantic information of song audio data in multiple aspects, it is particularly suitable for the feature extraction of song audio data with a relatively large amount of data, and is particularly suitable for the application scenario of long song processing. For such application scenarios, it can achieve a more accurate matching effect.
[0174] Please refer to Figure 7 , on the basis of the previous embodiment, the network structure of the feature extraction model of the present application is improved. It can be seen that Figure 7 The difference between the network architecture in Figure 6 and the network architecture in Figure 7 is that in , the output of the last convolutional block of the first branch network directly undergoes maximum pooling to obtain the global output feature vector, capturing the significant feature information of the encoded information of the song audio data; in the second branch network, the output of the last convolutional block is equally segmented into four parts of channel-corresponding feature information. After the channel-corresponding feature information of each part is respectively subjected to average pooling processing, they are re-concatenated into the corresponding output feature vector. It is not difficult to understand that through the segmentation and construction of local branches, the obtained output feature vectors can learn better local feature information.
[0175] This embodiment exemplarily provides a Figure 6 The modification based on the network architecture shown is relatively lightweight. Based on this example, it is not difficult to understand that the focus of the creative spirit of this application lies in the flexible combination of multiple branch networks. Based on the principles disclosed in this application, those skilled in the art can select a feature extraction model constructed by different branch network combinations to adapt to different specific uses according to the characteristics of multi-scale deep semantic information possessed by the output feature vectors obtained by each branch network, and transform various other embodiments of this application to meet needs such as humming recognition, song recognition, cover recognition, and song infringement comparison.
[0176] See also Figure 8 In a specific embodiment, when the convolution block is called, the following steps are performed:
[0177] Step S3100: Perform convolution transformation on the input information to obtain transformation feature information:
[0178] For any convolution block in the feature extraction model of the present application, each convolution block first performs a convolution operation on the information input therein, whether it is the encoded information or the intermediate feature information output by the previous convolution block, to obtain the corresponding transformed feature information.
[0179] Step S3200: performing instance normalization and batch normalization processing on the transformed feature information respectively, and then combining them into spliced feature information, and activating and outputting the spliced feature information:
[0180] After the first convolution, an instance batch normalization layer (IN) is applied to process the transformed feature information. The transformed feature information is divided into two paths, and a batch normalization block (BN) is used to batch normalize half of the channels, while the instance normalization layer is applied to the other channels to perform instance normalization, which enables the corresponding convolution block to capture the style-invariant features of the song audio data, thereby achieving better utilization of the song representation with diverse styles in a single data. After different normalization processes, the two parts of the channels can be spliced into the same spliced feature information for activation output.
[0181] Step S3300: The spliced feature information of the activated output is subjected to multiple convolution operations and batch normalization to obtain residual information:
[0182] The concatenated feature information of the activated output further undergoes convolution operations through multiple convolutional layers to further extract features. After each such convolutional layer, a batch normalization layer is connected for normalization processing and then output. Among them, the last convolutional layer is implemented with a 1*1 convolutional kernel to avoid the attenuation of the representation learning ability of the entire feature extraction model due to multiple instance normalization processes through multiple convolutional blocks. Accordingly, the finally output feature information is the residual information in the residual convolution process.
[0183] Step S3400: Superimpose the residual information onto the information input therein to activate the output.
[0184] Finally, according to the principle of residual convolution, referring to the transformed feature information obtained from the first convolution, superimpose it with the aforementioned residual information and then activate the output, and the intermediate feature information output after the residual convolution operation of the current convolutional block can be obtained.
[0185] In this embodiment, a convolutional block required for constructing the feature extraction model of the present application is applied based on the combination of residual convolution and instance batch normalization. The residual convolution network therein is improved based on the basic model of the Resnet series, and at the same time, the IBN architecture is superimposed. The thus constructed feature extraction model is easier to train and can achieve a more accurate feature extraction effect, and is particularly suitable for the feature extraction of song audio data.
[0186] Please refer to Figure 9 , in an extended embodiment, the training process of the feature extraction model includes the following iterative training steps:
[0187] Step S4100: Call a training sample from the training set and determine the encoding information of the training sample. The training sample is pre-collected song audio data, and the song audio data is a complete song and its fragments.
[0188] Those skilled in the art can understand that different training sets for training the feature extraction model can be constructed to adapt to different downstream tasks. Each training set contains a sufficient number of training samples, and each training sample is prepared with a corresponding supervision label.
[0189] The training samples described above can be pre-collected by those skilled in the art. Each training sample is a song audio data. Depending on the different downstream tasks, these song audio data can be a complete song, a song MIDI melody segment, a song with accompaniment, the a cappella part of a song, a song segment without melody, a song segment with melody, and so on. Different singing versions of the same song can be combined into the same category, that is, corresponding to the same supervision label, to enhance the generalization ability of the model classification. When the duration of the song audio data in the training sample is too long, it can also be segmented into multiple song segments according to a certain preset duration as multiple training samples and associated with the same supervision label for training. When segmenting the song, the time stamps of the lyrics of the song can be referred to for implementation, so that the song segments are segmented based on one or more complete lyrics.
[0190] Preferably, to adapt to the characteristics that the present application may extract the features of the whole song and also extract the features of song segments during the song comparison stage, during the training stage, for each song, training samples can be constructed both for the audio data of its whole song and for the audio data of multiple song segments of this song respectively, and these two types of training samples are associated with the supervision label of the same song.
[0191] For the training samples in the training set, for the convenience of model training, the encoding information corresponding to its song audio data can be prepared in advance, or the corresponding encoding information can also be obtained by real-time encoding when each song audio data is called for training the feature extraction model. As for the specific encoding principle, it can be processed by referring to the corresponding process disclosed in the foregoing of this application.
[0192] Step S4200: Input the encoding information into the feature extraction model for training to obtain the corresponding output feature vectors respectively:
[0193] During the training process of a training sample, the encoding information corresponding to this training sample is output to the feature extraction model for feature extraction. The feature extraction principle can be referred to the description of the principle of the feature extraction model in the foregoing embodiments, and will not be elaborated here. In this process, the feature extraction model realizes the representation learning of the training sample and obtains the corresponding output feature vectors respectively.
[0194] Step S4300: Perform classification prediction on each of the output feature vectors respectively to map out the corresponding classification labels:
[0195] In this application, the training task of the feature extraction model is understood as a classification task. Therefore, by connecting the output feature vectors of each path of the feature extraction model to the corresponding pre-prepared classification models, examining the classification results of each classification model, and supervising them with the corresponding supervision labels, the training of the model can be implemented. Based on this principle, in the training stage, when training the feature extraction model implemented in any embodiment of this application, a classification model is connected to each output feature vector output end of each branch network of it.
[0196] The classification model adopted is as Figure 10 shown in the structure. A batch normalization layer is used to perform batch normalization operations on the output feature vectors, and then through a fully connected layer, the output feature vectors are mapped to the classification space. The classification probabilities corresponding to each classification label are calculated through the classification function, so as to determine that the one with the largest classification probability is the classification label corresponding to the training sample.
[0197] The classifier in the classification model can be constructed by using a multi-classifier implemented by the Softmax function, or can be constructed by using a multi-classifier implemented by the AM-Softmax function that can enhance the intra-class compactness and expand the inter-class sparsity. Obviously, the latter has better classification advantages.
[0198] Step S4400: Calculate the loss value of the feature extraction model by using the supervision label corresponding to the training sample and the classification label, and perform gradient update on the feature extraction model according to the loss value:
[0199] In the classification model, the batch normalization layer is adopted to achieve the balance of the triplet loss and the cross-entropy classification loss. Subsequently, the triplet loss can be calculated for the batch normalization layer, and the cross-entropy classification loss can be calculated for the fully connected layer. By combining these two losses, the optimization of the output feature vectors can be realized.
[0200] Accordingly, after the corresponding classification label of the training sample is predicted, the loss value between the supervision label and the classification label can be calculated according to its corresponding supervision label, and then the feature extraction model is subjected to gradient update according to the loss value, and the weight parameters of each link of the entire model are corrected to promote the convergence of the model.
[0201] Since there are multiple branch networks, each branch network may have multiple outputs of output feature vectors, and correspondingly there are multiple classification models. Therefore, when calculating the loss value, a weighted method can be adopted, that is, the triplet loss and the classification loss in each classification model are first weighted and summed to obtain the loss value corresponding to each output feature vector, and then the loss values corresponding to each output feature vector are weighted and summed again to obtain the final loss value, and the entire feature extraction model is subjected to gradient update with this loss value.
[0202] Step S4500: Determine whether the loss value reaches a preset threshold. When it does not reach the preset threshold, call the next training sample in the training set to continue the iterative training of the feature extraction model until the loss value reaches the preset threshold:
[0203] For the loss value calculated for each training sample, determine whether it infinitely approaches the value of 0, or determine whether it reaches the preset threshold. When these judgment conditions are met, it can be determined that the feature extraction model has been trained to the convergence state. Accordingly, the training of the model can be terminated, and the feature extraction model can be put into the production stage, for example, used to extract features from the songs in the song library or serve other downstream tasks, etc. If the convergence state is not reached, the next training sample in the training set can be continuously called to continue the iterative training of the feature extraction model until the feature extraction model is trained to the convergence state.
[0204] This embodiment discloses the training principle and process of the feature extraction model of the present application. It can be seen from this embodiment that by training the feature extraction model with the prepared training set, the feature extraction model can learn the ability to extract the corresponding output feature vector from the encoded information of the song audio data, realizing effective representation learning of the deep semantic information of the song audio data. Moreover, the output feature vectors of multiple scales of the same song audio data can be jointly trained, with higher training efficiency and richer model functions. When it is put into the production stage, the deep semantic information corresponding to multiple scales of the same song audio data can be quickly obtained.
[0205] Since the classification model of this embodiment adopts a multi-classifier implemented with batch normalization and the AM-Softmax function, it can balance the triplet loss and the classification loss for gradient update of the model, enabling the model to be trained to convergence more quickly, and the trained model can better perform more effective representation learning on the deep semantic information of the song audio data. When the subsequent output feature vectors are combined and used as needed, they can more effectively characterize the feature information of the song audio data and play a more efficient matching role.
[0206] This embodiment also reflects the scalability and compatibility of the feature extraction model of the present application in terms of application. Specifically, this embodiment allows, for the need of serving different downstream tasks, by training the feature extraction model with the training samples corresponding to different downstream tasks, the feature extraction model can obtain the ability to serve different downstream tasks. Therefore, it belongs to a relatively basic improvement and has better economic utility.
[0207] Please refer to Figure 11, a song comparison device provided by the present application is functionally deployed to adapt to the song comparison method of the present application, including: an audio acquisition module 1100, an original version extraction module 1200, a song to be compared extraction module 1300, and a comprehensive judgment module 1400. Among them, the audio acquisition module 1100 is used to acquire the audio data of the original song and the song to be compared respectively; the original version extraction module 1200 is used to extract the multi-scale deep semantic information of the audio data of the original song by using a feature extraction model that has been trained to a convergent state, and correspondingly obtain an original high-dimensional index vector; the song to be compared extraction module 1300 is used to extract the multi-scale deep semantic information of the audio data of the song to be compared by using the feature extraction model, and correspondingly obtain a song to be compared high-dimensional index vector; the comprehensive judgment module 1400 is used to calculate the similarity between the original high-dimensional index vector and the song to be compared high-dimensional index vector, and judge whether the corresponding similarity value is greater than a first preset threshold. When it is greater than the first preset threshold, it is determined that the song to be compared and the original song form a cover relationship.
[0208] In an extended embodiment, the song comparison device includes a structure that runs when the comprehensive judgment module 1400 determines that the similarity value is less than the first preset threshold. This structure includes: a song segmentation module, which is used to determine the song with a relatively shorter audio duration among the original song and the song to be compared as the specified song, and segment the other song with a relatively longer audio duration based on the audio duration of the specified song to obtain multiple song segments of the other song; a segmented extraction module, which is used to extract the multi-scale deep semantic information of the multiple song segments of the other song by using the feature extraction model, and correspondingly obtain multiple segment high-dimensional index vectors; a segmented judgment module, which is used to use the high-dimensional index vector of the specified song as the specified high-dimensional index vector, calculate the similarity between the specified high-dimensional index vector and each segment high-dimensional index vector respectively, and judge whether the maximum similarity value is higher than a second preset threshold. When it is higher than the second preset threshold, it is determined that the song to be compared forms a cover relationship.
[0209] In a further embodiment, the audio acquisition module 1100 includes: an original version acquisition sub-module, which is used to acquire the audio data and lyrics file of the original song; a word segmentation extraction sub-module, which is used to segment the lyrics in the lyrics file of the original song and extract multiple keywords therefrom; an online search sub-module, which is used to online search for at least one song according to any combination of the keywords, and use the searched song as the song to be compared; a song to be compared acquisition sub-module, which is used to acquire the audio data of the song to be compared.
[0210] In an alternative embodiment, the original extraction module 1200 and / or the comparison extraction module 1300 includes: an audio encoding sub-module for encoding the audio data to obtain its encoding information; and a model calling sub-module for using the feature extraction model to extract a high-dimensional index vector representing the multi-scale deep semantic information of the audio data according to the encoding information.
[0211] In an enhanced embodiment, the feature extraction model is implemented as the following structure, which includes: a shared extraction module configured to sequentially perform multi-level feature extraction on the encoding information using multiple convolutional blocks in the shared network of the feature extraction model that has been trained to a convergent state, to obtain intermediate feature information that has extracted the deep semantic information of the encoding information; a branch extraction module configured to perform feature extraction of different scales on the intermediate feature information using multiple convolutional blocks in two or more branch networks of the feature extraction model, and then convert it into output feature vectors of corresponding scales, where the deep semantic information contained in the output feature vectors of each branch network is different; and a processing output module configured to output the output feature vectors of each branch network as the high-dimensional index vector.
[0212] In a further embodiment, the branch extraction module includes any two or more of the following modules: a first extraction sub-module configured to perform feature extraction on the intermediate feature information using multiple convolutional blocks in the first branch network to obtain global feature information, and pool the global feature information into an output feature vector of the global scale; a second extraction sub-module configured to perform feature extraction on the intermediate feature information using multiple convolutional blocks in the second branch network, then split it into multiple parts by channels for pooling, and correspondingly obtain output feature vectors of the channel scale; and a third extraction sub-module configured to perform feature extraction on the intermediate feature information using multiple convolutional blocks in the third branch network, then split it into multiple parts by frequency bands for pooling, and correspondingly obtain output feature vectors of the frequency band scale.
[0213] In a specific embodiment, when the first branch network performs the pooling operation, it uses average pooling and / or maximum pooling operations to correspondingly obtain one or two output feature vectors of the global scale; when the second branch network performs the pooling operation, it uses average pooling operations for single or multiple channels to correspondingly obtain one or more output feature vectors of the channel scale; when the third branch network performs the pooling operation, it uses average pooling operations for single or multiple frequency bands to correspondingly obtain one or more output feature vectors of the frequency band scale.
[0214] In an alternative embodiment, the source of the encoding information is any one of the time-frequency spectrum information, Mel spectrum information, CQT filtering information, pitch contour information, and Chroma feature information of the corresponding audio data.
[0215] In a preferred embodiment, in the shared network, at least one of the convolutional blocks applies an attention module to extract key information from the song audio data, and the attention module is a spatial attention module or a channel attention module.
[0216] In an optimized embodiment, the convolutional block is implemented as the following structure, which includes: an initial convolution unit for performing convolution transformation on the input information to obtain transformed feature information; a normalization processing unit for respectively performing instance normalization and batch normalization on the transformed feature information and then combining them into spliced feature information, and activating and outputting the spliced feature information; a residual calculation unit for obtaining residual information after performing multiple convolution operations and batch normalization on the activated and output spliced feature information; and an activation output unit for superimposing the residual information on the input information and activating and outputting it.
[0217] In an extended embodiment, the feature extraction model is placed in a training task implemented by the following structure for iterative training, which includes: a sample calling module for calling a training sample from the training set and determining the encoding information of the training sample, where the training sample is pre-collected song audio data, and the song audio data is a complete song and its segments; a representation learning module for inputting the encoding information into the feature extraction model for training to obtain corresponding output feature vectors; a classification prediction module for respectively performing classification prediction on each of the output feature vectors to map out corresponding classification labels; a loss calculation module for calculating the loss value of the feature extraction model by using the supervision label corresponding to the training sample and the classification label, and performing gradient update on the feature extraction model according to the loss value; and an iterative decision module for determining whether the loss value reaches a preset threshold. When the preset threshold is not reached, the next training sample in the training set is called to continue iterative training on the feature extraction model until the loss value reaches the preset threshold.
[0218] To solve the above technical problems, an embodiment of the present application also provides a computer device. As Figure 12As shown, it is a schematic diagram of the internal structure of a computer device. The computer device includes a processor, a computer-readable storage medium, a memory, and a network interface connected via a system bus. Among them, the computer-readable storage medium of the computer device stores an operating system, a database, and computer-readable instructions. The database can store a control information sequence. When the computer-readable instructions are executed by the processor, the processor can implement a song comparison method. The processor of the computer device is used to provide computing and control capabilities to support the operation of the entire computer device. The memory of the computer device can store computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor can execute the song comparison method of the present application. The network interface of the computer device is used to connect and communicate with a terminal. Those skilled in the art can understand that Figure 12 the structure shown is only a block diagram of some structures related to the solution of the present application and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0219] In this embodiment, the processor is used to execute Figure 11 the specific functions of each module and its sub-modules in. The memory stores the program codes and various types of data required to execute the above modules or sub-modules. The network interface is used for data transmission between the user terminal and the server. The memory in this embodiment stores the program codes and data required to execute all modules / sub-modules in the song comparison device of the present application. The server can call the program codes and data of the server to execute the functions of all sub-modules.
[0220] The present application also provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, one or more processors are caused to execute the steps of the song comparison method according to any embodiment of the present application.
[0221] The present application also provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by one or more processors, the steps of the method according to any embodiment of the present application are implemented.
[0222] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments of the present application can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium can be a computer-readable storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0223] In summary, the present application uses a feature extraction model to achieve representation learning of multi-scale deep semantic information of the original song and the compared song, obtaining corresponding high-dimensional index vectors. Based on the respective high-dimensional index vectors, similarity comparison is performed to determine whether the two constitute a cover relationship, achieving a more accurate and efficient comparison effect. It can serve multiple downstream tasks such as song recognition by listening, song recognition by humming, cover recognition, and song infringement comparison, improving the comprehensive service ability of online music platforms.
[0224] Those skilled in the art of the present technology can understand that the various operations, methods, steps, measures, and solutions in the processes discussed in the present application can be alternated, changed, combined, or deleted. Further, other steps, measures, and solutions in the various operations, methods, and processes discussed in the present application can also be alternated, changed, rearranged, decomposed, combined, or deleted. Further, the steps, measures, and solutions in the prior art that are the same as those disclosed in the present application can also be alternated, changed, rearranged, decomposed, combined, or deleted.
[0225] The above are only some embodiments of the present application. It should be noted that for those of ordinary skill in the art of the present technology, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A song comparison method, characterized in that, it includes the following steps: Obtain the audio data of the original song and the song to be compared respectively; Use a feature extraction model that has been trained to a convergent state to extract the multi-scale deep semantic information of the audio data of the original song, and correspondingly obtain the original high-dimensional index vector; Use the feature extraction model to extract the multi-scale deep semantic information of the audio data of the song to be compared, and correspondingly obtain the compared high-dimensional index vector; Calculate the similarity between the original high-dimensional index vector and the compared high-dimensional index vector, and judge whether the corresponding similarity value is greater than a first preset threshold. When it is greater than the first preset threshold, it is determined that the song to be compared and the original song form a cover relationship; When the feature extraction model is called, the following steps are executed: Use multiple convolutional blocks in the shared network of the feature extraction model that has been trained to a convergent state to sequentially perform multi-level feature extraction on the encoded information of the audio data, and obtain intermediate feature information that has extracted the deep semantic information of the encoded information; Use multiple convolutional blocks in more than two branch networks in the feature extraction model to perform feature extraction on the intermediate feature information at different scales, and then convert it into output feature vectors of corresponding scales. The deep semantic information contained in the output feature vectors of each branch network is different; The feature extraction model outputs the output feature vectors of each branch network as the high-dimensional index vector.
2. The song comparison method according to claim 1, characterized in that, In the step of calculating the similarity between the original high-dimensional index vector and the compared high-dimensional index vector and judging whether the corresponding similarity value is greater than the first preset threshold, when it is less than the preset threshold, the following steps are executed: Determine the song with the relatively shorter audio duration among the original song and the song to be compared as the specified song, and segment the other song with the relatively longer audio duration with the audio duration of the specified song as the metric to obtain multiple song segments of the other song; Use the feature extraction model to extract the multi-scale deep semantic information of the multiple song segments of the other song respectively, and correspondingly obtain multiple segment high-dimensional index vectors; Use the high-dimensional index vector of the specified song as the specified high-dimensional index vector, calculate the similarity between the specified high-dimensional index vector and each segment high-dimensional index vector respectively, and judge whether the maximum similarity value is higher than a second preset threshold. When it is higher than the second preset threshold, it is determined that the song to be compared forms a cover relationship.
3. The song comparison method according to claim 1, characterized in that, Obtaining the audio data of the original song and the song to be compared includes the following steps: Obtain the audio data and lyrics file of the original song; Segment the lyrics in the lyrics file of the original song, and extract multiple keywords from them; Online search for at least one song according to any combination of the keywords, and use the searched song as the song to be compared; Obtain the audio data of the song to be compared.
4. The song comparison method according to claim 1, characterized in that, In the step of extracting the multi-scale deep semantic information of the audio data of the original song by using the feature extraction model trained to the convergence state and correspondingly obtaining the original high-dimensional index vector, or, in the step of extracting the multi-scale deep semantic information of the audio data of the song to be compared by using the feature extraction model and correspondingly obtaining the compared high-dimensional index vector, the following steps are included: Encode the audio data to obtain its encoding information accordingly; Use the feature extraction model to extract a high-dimensional index vector representing the multi-scale deep semantic information of the audio data according to the encoding information.
5. According to the song comparison method described in any one of claims 1 to 4, characterized in that, After using multiple convolutional blocks in two or more branch networks in the feature extraction model to perform feature extraction on the intermediate feature information at different scales, it is converted into output feature vectors of corresponding scales, including any two or more of the following steps: Use multiple convolutional blocks in the first branch network to perform feature extraction on the intermediate feature information to obtain global feature information, and pool the global feature information into an output feature vector at the global scale; Use multiple convolutional blocks in the second branch network to perform feature extraction on the intermediate feature information, then split it into multiple parts according to channels for pooling, and correspondingly obtain output feature vectors at the channel scale; Use multiple convolutional blocks in the third branch network to perform feature extraction on the intermediate feature information, then split it into multiple parts according to frequency bands for pooling, and correspondingly obtain output feature vectors at the frequency band scale.
6. According to the song comparison method described in claim 5, characterized in that: When the first branch network performs the pooling operation, average pooling and / or maximum pooling operations are used to correspondingly obtain one or two output feature vectors at the global scale; When the second branch network performs the pooling operation, average pooling operation is used for single or multiple channels to correspondingly obtain one or more output feature vectors at the channel scale; When the third branch network performs the pooling operation, average pooling operation is used for single or multiple frequency bands to correspondingly obtain one or more output feature vectors at the frequency band scale.
7. According to the song comparison method described in claim 4, characterized in that, The source of the encoding information is any one of the time-frequency spectrum information, Mel spectrum information, CQT filtering information, pitch contour information, and Chroma feature information of the corresponding audio data.
8. According to the song comparison method described in any one of claims 1 to 4, characterized in that, In the shared network, at least one of the convolutional blocks applies an attention module to extract key information in the song audio data, and the attention module is a spatial attention module or a channel attention module.
9. According to the song comparison method described in any one of claims 1 to 4, characterized in that, When the convolutional block is called, the following steps are performed: Perform a convolutional transformation on the information input thereto to obtain transformed feature information; Perform instance normalization and batch normalization processing on the transformed feature information respectively, then combine them into spliced feature information, and activate and output the spliced feature information. The concatenated feature information of the activation output is subjected to multiple convolutional operations and batch normalization processing to obtain residual information; The residual information is superimposed on the information input thereto to activate the output.
10. The song comparison method according to any one of claims 1 to 4, characterized in that, the training process of the feature extraction model includes the following steps of iterative training: Call a training sample from the training set, and determine the encoded information of the training sample. The training sample is pre-acquired song audio data, and the song audio data is a complete song and its segments; Input the encoded information into the feature extraction model for training to obtain corresponding output feature vectors; Perform classification prediction on each of the output feature vectors respectively to map corresponding classification labels; Calculate the loss value of the feature extraction model by using the supervision label corresponding to the training sample and the classification label, and perform gradient update on the feature extraction model according to the loss value; Judge whether the loss value reaches a preset threshold. When the preset threshold is not reached, call the next training sample in the training set to continue iterative training of the feature extraction model until the loss value reaches the preset threshold.
11. A song comparison device, characterized in that, it implements the method according to any one of claims 1 to 10, including: An audio acquisition module for respectively acquiring the audio data of the original song and the song to be compared; An original extraction module for extracting the multi-scale deep semantic information of the audio data of the original song by using a feature extraction model that has been trained to a convergent state, and correspondingly obtaining an original high-dimensional index vector; A to-be-compared extraction module for extracting the multi-scale deep semantic information of the audio data of the song to be compared by using the feature extraction model, and correspondingly obtaining a to-be-compared high-dimensional index vector; A comprehensive judgment module for calculating the similarity between the original high-dimensional index vector and the to-be-compared high-dimensional index vector, and judging whether the corresponding similarity value is greater than a first preset threshold. When it is greater than the first preset threshold, it is determined that the song to be compared and the original song form a cover relationship.
12. A computer device, including a central processing unit and a memory, characterized in that, the central processing unit is used to call and run a computer program stored in the memory to execute the steps of the method according to any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that, it stores a computer program implemented according to the method according to any one of claims 1 to 10 in the form of computer-readable instructions. When the computer program is called and run by a computer, it executes the steps included in the corresponding method.
14. A computer program product, including a computer program / instructions, characterized in that, when the computer program / instructions are executed by a processor, the steps of the method described in any one of claims 1 to 10 are implemented.
Citation Information
Patent Citations
Content-based voice frequency semantic feature similarity comparative method
CN102841932A
Voice frequency searching method based on M-tree
CN103324691A