Song relationship identification method, computer device and computer program product

By automatically identifying the degree and duration of repetition in song audio, the problem of low efficiency in song relationship identification is solved, achieving fast and accurate song relationship identification.

CN116166838BActive Publication Date: 2026-06-05TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
Filing Date
2022-12-29
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

In existing technologies, the efficiency of song relationship identification is low, and it is difficult to cope with the massive existing data and daily incremental data in the audio library. Relying on manual identification is time-consuming and labor-intensive.

Method used

By acquiring the repetition degree and duration of the audio of the song to be identified, a computer program is used to automatically identify the relationship between the audio of the songs, including determining the audio content and duration of the first song audio that is covered by the second song audio, and the audio content and duration of the second song audio that is covered by the first song audio, and then determining the relationship between the songs based on these indicators.

Benefits of technology

It enables rapid identification of relationships between song audio, improves the efficiency of song relationship identification, reduces manual intervention, and enhances recognition efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116166838B_ABST
    Figure CN116166838B_ABST
Patent Text Reader

Abstract

The application relates to a song relationship identification method, computer equipment and a computer program product, which can improve the identification efficiency of song relationships. The method comprises the following steps: obtaining first song audio and second song audio of a song relationship to be identified; determining a first repetition degree of the first song audio with respect to the second song audio according to first audio content covered by the second song audio in the first song audio, and obtaining a first duration of the first audio content; determining a second repetition degree of the second song audio with respect to the second song audio according to second audio content covered by the first song audio in the second song audio, and obtaining a second duration of the second audio content; and determining a song relationship between the first song audio and the second song audio according to the first repetition degree, the first duration, the second repetition degree and the second duration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing technology, and in particular to a song relationship identification method, computer device, and computer program product. Background Technology

[0002] With the development of computer technology, users can use audio and video applications to create derivative works from audio and share them. Accurate and efficient detection of relationships between songs can improve audio service quality and user experience.

[0003] In related technologies, the main reliance is on manual identification of the relationships between different songs, such as review by specialized staff or information provided by internet users. However, this current method is time-consuming and labor-intensive, suffers from low identification efficiency, and struggles to handle the massive amounts of existing data and daily incremental data in audio libraries. Summary of the Invention

[0004] Therefore, it is necessary to provide a song relationship identification method, apparatus, computer device, computer-readable storage medium, and computer program product to address the aforementioned technical problems.

[0005] Firstly, this application provides a method for identifying song relationships. The method includes:

[0006] Obtain the audio of the first and second songs whose song relationship needs to be identified;

[0007] Based on the first audio content in the first song audio that is covered by the second song audio, determine the first degree of repetition of the first song audio with respect to the second song audio, and obtain the first duration of the first audio content;

[0008] Based on the second audio content in the second song audio that is covered by the first song audio, determine the second degree of repetition of the second song audio relative to the second song audio, and obtain the second duration of the second audio content;

[0009] The song relationship between the first song audio and the second song audio is determined based on the first degree of repetition, the first duration, the second degree of repetition, and the second duration.

[0010] Secondly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:

[0011] Obtain the audio of the first and second songs whose song relationship needs to be identified;

[0012] Based on the first audio content in the first song audio that is covered by the second song audio, determine the first degree of repetition of the first song audio with respect to the second song audio, and obtain the first duration of the first audio content;

[0013] Based on the second audio content in the second song audio that is covered by the first song audio, determine the second degree of repetition of the second song audio relative to the second song audio, and obtain the second duration of the second audio content;

[0014] The song relationship between the first song audio and the second song audio is determined based on the first degree of repetition, the first duration, the second degree of repetition, and the second duration.

[0015] Thirdly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:

[0016] Obtain the audio of the first and second songs whose song relationship needs to be identified;

[0017] Based on the first audio content in the first song audio that is covered by the second song audio, determine the first degree of repetition of the first song audio with respect to the second song audio, and obtain the first duration of the first audio content;

[0018] Based on the second audio content in the second song audio that is covered by the first song audio, determine the second degree of repetition of the second song audio relative to the second song audio, and obtain the second duration of the second audio content;

[0019] The song relationship between the first song audio and the second song audio is determined based on the first degree of repetition, the first duration, the second degree of repetition, and the second duration.

[0020] The aforementioned song relationship identification method, computer device, and computer program product can acquire first and second song audio files to be identified, determine a first degree of repetition of the first song audio file relative to the second song audio file based on first audio content covered by the second song audio file in the first song audio file, and obtain a first duration of the first audio content; and determine a second degree of repetition of the second song audio file relative to the second song audio file based on second audio content covered by the first song audio file in the second song audio file, and obtain a second duration of the second audio content. Furthermore, the song relationship between the first and second song audio files can be determined based on the first degree of repetition, the first duration, the second degree of repetition, and the second duration. In this embodiment, by identifying the audio content covered by the two song audio files and the duration of the covered audio content, the song relationship between the first and second song audio files can be quickly identified, effectively improving the efficiency of song relationship identification. Attached Figure Description

[0021] Figure 1 This is a flowchart illustrating a song relationship identification method in one embodiment;

[0022] Figure 2 This is a flowchart illustrating one step in determining first audio content in one embodiment;

[0023] Figure 3 This is a flowchart illustrating one step in determining a first similarity in one embodiment;

[0024] Figure 4a This is a schematic diagram illustrating the first similarity between a first audio segment and the audio of a second song in one embodiment;

[0025] Figure 4b This is a schematic diagram illustrating the second similarity between a second audio segment and the audio of a first song in one embodiment;

[0026] Figure 5a This is a schematic diagram illustrating the first similarity between another first audio segment and the audio of a second song in one embodiment;

[0027] Figure 5b This is a schematic diagram illustrating the second similarity between another second audio segment and the audio of the first song in one embodiment;

[0028] Figure 6a This is a schematic diagram illustrating the first similarity between another first audio segment and the audio of a second song in one embodiment;

[0029] Figure 6b This is a schematic diagram illustrating the second similarity between another second audio segment and the audio of the first song in one embodiment;

[0030] Figure 7a This is a schematic diagram illustrating the first similarity between another first audio segment and the audio of a second song in one embodiment;

[0031] Figure 7b This is a schematic diagram illustrating the second similarity between another second audio segment and the audio of the first song in one embodiment;

[0032] Figure 8 This is a flowchart illustrating another song relationship identification method in one embodiment;

[0033] Figure 9 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0034] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0035] In one embodiment, such as Figure 1 As shown, a method for identifying song relationships is provided. This embodiment illustrates the application of this method to a server. It is understood that this method can also be applied to a terminal, or to a system including both a terminal and a server, and is implemented through interaction between the terminal and the server. The terminal can be, but is not limited to, various personal computers, laptops, smartphones, or tablets; the server can be a standalone server or a server cluster consisting of multiple servers. The server can have a corresponding data storage system for storing the data that the server needs to process, such as the audio of the songs whose relationships need to be identified. This data storage system can be integrated on the server or located in the cloud or on other network servers.

[0036] In this embodiment, the method may include the following steps:

[0037] S101, Obtain the audio of the first song and the audio of the second song whose song relationship is to be identified.

[0038] In practical applications, two song audio files can be acquired. To facilitate differentiation between the two audio files later, they can be referred to as the first song audio and the second song audio, respectively. Therefore, the two acquired song audio files can be considered as the first and second song audio files to be used to identify the relationship between the songs. The first and / or second song audio files can be audio from an audio library or audio extracted from a video, such as background music. They can be complete song audio or audio from a segment of a song. Correspondingly, the lengths of the first and second song audio files can be equal or unequal.

[0039] In this step, the method of triggering the acquisition of the first and second song audios for identifying the song relationship is not limited. For example, upon receiving a song relationship identification request, the two song audios related to the request can be used as the first and second song audios. For instance, when receiving a user's request to find other versions of a target song audio (e.g., a medley or excerpt), it can be determined that a song relationship identification request has been received. The target song audio can then be used as the first song audio, the second song audio can be retrieved from the audio library, and the song relationship between the two can be identified. Based on the song relationship identification result, other versions of the target song audio can be obtained. Alternatively, the user can send a song relationship identification request for two song audios, and upon receiving the request, the two song audios carried in the request can be used as the first and second song audios, respectively.

[0040] For example, when a song audio publishing request is received, the song audio to be published can be used as the first song audio, and the second song audio can be obtained from the audio library. The song relationship between the first song audio and the second song audio can be identified, and the song relationship identification result can be used to determine whether the song audio to be published uses audio content from other songs.

[0041] In one example, to improve the efficiency of identifying song relationships, the first and second song audios to be identified can be related audios. For example, two audios that meet at least one of the following conditions can be identified as related audios: the audio content is similar (the similarity of at least part of the audio content of the first and second song audios is higher than a threshold), or the audio names are the same or the similarity is higher than a threshold.

[0042] S102, based on the first audio content in the first song audio that is covered by the second song audio, determine the first degree of repetition of the first song audio with respect to the second song audio, and obtain the first duration of the first audio content.

[0043] The degree of repetition, also known as overall similarity, characterizes the extent to which the audio content of one song repeats at least a portion of the audio content of another song. This repetition degree is not limited to representing the complete reproduction of one song from another; it can also characterize the reproduction of a portion of another song's content. For example, if song A is generated by repeatedly repetitive segments of song B, although song A differs significantly from the entire song B, since all audio content of song A is based on the audio content of song B, it can be determined that song A has a high degree of repetition with respect to song B. In this embodiment, for ease of distinction, the degree of repetition of the first song audio with respect to the second song audio can be referred to as the first degree of repetition.

[0044] In practice, after obtaining the audio of the first song, the audio content within the first song that is covered by the audio of the second song can be identified as the first audio content. The first audio content can refer to one or more audio segments within the entire first song audio that are covered by the audio of the second song. These covered first audio segments can be different audio segments or contain the same audio segments. For example, the first audio content can be divided into five audio segments: A, B, C, D, and E. If audio segments A and C are both covered by the audio of the second song, then audio segments A and C can both be identified as audio segments covered by the audio of the second song.

[0045] Furthermore, after determining the first audio content in the first song audio that is covered by the second song audio, the first degree of repetition of the first song audio with respect to the second song audio can be determined based on the first audio content, and the first duration of the first audio content can be obtained.

[0046] S103, based on the second audio content in the second song audio that is covered by the first song audio, determine the second degree of repetition of the second song audio relative to the second song audio, and obtain the second duration of the second audio content.

[0047] The degree to which the second song audio repeats the first song audio is referred to as the second degree of repetition.

[0048] In this step, after obtaining the second song audio, the audio content covered by the first song audio in the second song audio can be identified as the second audio content. The second audio content can refer to one or more audio segments covered by the first song audio in the entire second song audio. The covered second audio content can be different audio segments or can contain the same audio segment.

[0049] Furthermore, after determining the second audio content in the second song audio that is covered by the first song audio, the second degree of repetition of the second song audio relative to the first song audio can be determined based on the second audio content, and the second duration of the second audio content can be obtained.

[0050] S104, determine the song relationship between the first song audio and the second song audio based on the first degree of repetition, the first duration, the second degree of repetition, and the second duration.

[0051] Specifically, after obtaining the first degree of repetition, the first duration, the second degree of repetition, and the second duration, the degree of repetition between the first song audio and the second song audio can be determined based on the first degree of repetition and the second degree of repetition.

[0052] Furthermore, since both the first and second audio content appear simultaneously in the first and second song audios, but the manner in which they appear may differ—for example, the repeated appearance of a certain audio segment or the reproduction of multiple different audio segments—the differences will be reflected in the first and second durations. Therefore, based on the first and second durations, it can be determined whether the first audio content covered in the first song audio and the second audio content covered in the second song include multiple repeated audio segments.

[0053] Therefore, the relationship between the first and second song audios can be determined based on the first degree of repetition, the first duration, the second degree of repetition, and the second duration.

[0054] In the above-described song relationship identification method, a first song audio and a second song audio of the song relationship to be identified can be obtained. Based on the first audio content in the first song audio that is covered by the second song audio, a first degree of repetition of the first song audio with respect to the second song audio is determined, and a first duration of the first audio content is obtained. Furthermore, based on the second audio content in the second song audio that is covered by the first song audio, a second degree of repetition of the second song audio with respect to the second song audio is determined, and a second duration of the second audio content is obtained. Then, the song relationship between the first song audio and the second song audio can be determined based on the first degree of repetition, the first duration, the second degree of repetition, and the second duration. In this embodiment, by identifying the audio content covered by each other and the duration of the covered audio content, the song relationship between the first song audio and the second song audio can be quickly identified, effectively improving the efficiency of song relationship identification.

[0055] In one embodiment, such as Figure 2 As shown, the following steps may also be included before S102:

[0056] S201, Obtain multiple first audio segments of the first song audio.

[0057] In the specific implementation, after obtaining the audio of the first song, the audio of the first song can be segmented, and multiple audio segments of the audio of the first song can be obtained based on the segmentation results. In order to facilitate the distinction between the audio segments of the first song and the audio segments of the second song, the audio segments of the first song can be referred to as the first audio segment.

[0058] S202, for each first audio segment, determine the first similarity between the first audio segment and the audio of the second song.

[0059] After obtaining multiple first audio segments, the similarity between each first audio segment and the second song audio can be determined. To distinguish it from the similarity between the second audio segment and the first song audio in the following text, the similarity between the first audio segment and the second song audio can be called the first similarity.

[0060] Specifically, the first similarity can be used to characterize the degree of similarity between the audio content in the first audio segment and the audio content of the second song. The first similarity can be determined based on the same audio content in the first audio segment and the second song. For example, if the audio content in the first audio segment can be obtained in the second song, then it can be determined that the first audio segment and the second song have a high degree of similarity.

[0061] S203, based on the first audio segment with a first similarity greater than the similarity threshold, determine the first audio content in the first song audio that is covered by the second song audio.

[0062] After obtaining the first similarity between each first audio segment and the second song audio, multiple first audio segments can be filtered based on a preset similarity threshold (such as 0.3) to determine the first audio segments with a first similarity greater than the similarity threshold. Then, based on the first audio segments with a first similarity greater than the similarity threshold, the first audio content in the first song audio that is covered by the second song audio can be determined.

[0063] In this embodiment, the first song audio can be divided into multiple first audio segments, and the similarity between each first audio segment and the second song audio can be determined. This allows for the rapid location of the first audio content covered by the second song audio from different positions in the first song audio, providing a basis for subsequently determining the degree of repetition of the first song audio with respect to the second song audio.

[0064] In one embodiment, determining the first similarity between the first audio segment and the second song audio in step S202 may include the following steps:

[0065] Obtain the segment fingerprint features of the first audio segment and the song fingerprint features of the second song audio; based on the segment fingerprint features of the first audio segment and the song fingerprint features of the second song audio, determine the first similarity between the first audio segment and the second song audio.

[0066] Here, fingerprint features, also known as audio fingerprints, refer to information extracted from audio using a preset algorithm, where distinctive digital features are presented as identifiers. For example, feature frequency points can be extracted from the audio spectrum, and an audio fingerprint can be generated based on multiple feature frequency points. In this embodiment, the fingerprint features of an audio segment can be called segment fingerprint features, and the fingerprint features of a song audio can be called song fingerprint features.

[0067] In specific implementation, after obtaining multiple first audio segments of the first song audio, a segment fingerprint feature can be obtained for each first audio segment. This segment fingerprint feature can be obtained by extracting features from all audio content in the first audio segment. Furthermore, a song fingerprint feature can also be obtained for the second song audio; that is, features can be extracted from all audio content in the second song audio to generate a fingerprint feature for the entire second song audio, which serves as the song fingerprint feature.

[0068] Since fingerprint features can characterize distinctive audio content, that is, information that distinguishes one audio content from other audio content, in this embodiment, the first similarity between the first audio segment and the second song audio can be determined by comparing the segment fingerprint features of the first audio segment and the song fingerprint features of the second song audio.

[0069] In this embodiment, by comparing the segment fingerprint features of the first audio segment with the song fingerprint features of the second song audio, the audio content in the first audio segment can be quickly compared with the audio content of the second song audio to determine the first similarity.

[0070] In one embodiment, such as Figure 3 As shown, determining the first similarity between the first audio segment and the second song audio based on the segment fingerprint features of the first audio segment and the song fingerprint features of the second song audio can include the following steps:

[0071] S301, obtain the common fingerprint feature between the segment fingerprint feature of the first audio segment and the song fingerprint feature of the second song audio.

[0072] In this embodiment, the common part between the segment fingerprint features of the first audio segment and the song fingerprint features of the second song audio can be obtained, and the fingerprint features of the common part can be used as the common fingerprint features.

[0073] For example, the longest common subsequence of the segment fingerprint features of the first audio segment and the song fingerprint features of the second song audio segment can be obtained, and this longest common subsequence can be used as the common fingerprint feature. The longest common subsequence consists of multiple characters, which are the common parts obtained from the segment fingerprint features and the song fingerprint features. These multiple characters can be continuous or non-contiguous.

[0074] For example, if English characters are used to refer to fingerprint features, in one example, the fragment fingerprint feature of the first audio segment can be BCEG, while the song fingerprint feature of the second song audio can be AEBCJHGHA. It can be determined that although the fragment fingerprint feature BCEG does not appear continuously in the song fingerprint feature, without changing the order of fingerprint features, the fragment fingerprint feature of the first audio segment and the song fingerprint feature of the second song audio both include the fingerprint feature BCG. Therefore, the fingerprint feature of the common part can be regarded as the common fingerprint feature.

[0075] S302, determine the shortest fingerprint feature among the segment fingerprint features of the first audio segment and the song fingerprint features of the second song audio.

[0076] In a specific implementation, the feature length of the segment fingerprint feature of the first audio segment and the feature length of the song fingerprint feature of the second song audio can be obtained. By comparing the feature lengths of the two, the fingerprint feature with the shorter feature length can be determined as the shortest fingerprint feature.

[0077] S303, determine the first similarity between the first audio segment and the second song audio based on the ratio of the feature length of the common fingerprint feature to the feature length of the shortest fingerprint feature.

[0078] After determining the shortest fingerprint feature, the feature length of the common fingerprint feature can be determined, and the ratio of the feature length of the common fingerprint feature to the feature length of the shortest fingerprint feature can be obtained. Then, the first similarity between the first audio segment and the second song audio can be determined based on this ratio.

[0079] In this embodiment, the first similarity between the first audio segment and the second song audio can be determined based on the common fingerprint features between the segment fingerprint features of the first audio segment and the song fingerprint features of the second song audio. This allows for the rapid and accurate determination of the first similarity based on any audio content in the first audio segment that is similar to the second song audio.

[0080] In one embodiment, S102, based on the first audio content in the first song audio that is covered by the second song audio, determines the first degree of repetition of the first song audio to the second song audio, which may include the following steps:

[0081] Determine the number of first audio segments whose first similarity is greater than a similarity threshold; based on the number of first audio segments and the total number of segments of multiple first audio segments of the first song audio, obtain the first degree of repetition of the first song audio for the second song audio.

[0082] In a specific implementation, after obtaining the first similarity of multiple first audio segments, the number of first audio segments with a first similarity greater than the similarity threshold can be counted to determine the number of first audio segments with a first similarity greater than the similarity threshold.

[0083] After obtaining the number of the first audio segments, the number of the first audio segments can be compared with the total number of segments of multiple first audio segments to determine the proportion of similar first audio segments in all first audio segments. For example, the ratio of the number of first audio segments to the total number of segments can be obtained, thereby determining the first degree of repetition of the first song audio to the second song audio.

[0084] In this embodiment, by comparing the number of first audio segments with a first similarity greater than a similarity threshold with the total number of segments of the first audio segment, the proportion of the first audio segment that has repeated the audio content of the second song in the entire first song audio can be determined, and the degree of repetition of the first song audio to the second song audio can be accurately identified.

[0085] In one embodiment, prior to S103, the method may further include the following steps:

[0086] S401, Obtain multiple second audio segments of the second song audio.

[0087] In the specific implementation, after obtaining the audio of the second song, the audio of the second song can be segmented, and multiple audio segments of the second song audio can be obtained based on the segmentation results. In order to facilitate the distinction between the audio segments of the second song audio and the audio segments of the first song audio, the audio segments of the second song audio can be referred to as the second audio segments.

[0088] S402, for each second audio segment, determine the second similarity between the second audio segment and the audio of the first song.

[0089] After obtaining multiple second audio segments, the similarity between each second audio segment and the first song audio can be determined. To distinguish it from the similarity between the first audio segment and the second song audio in the following text, the similarity between the second audio segment and the first song audio can be called the second similarity.

[0090] Specifically, the second similarity can be used to characterize the degree of similarity between the audio content in the second audio segment and the audio content in the first song audio. The second similarity can be determined based on the same audio content in the second audio segment and the first song audio. For example, if the audio content in the second audio segment can be obtained from similar or identical content in the first song audio, then it can be determined that the second audio segment and the first song audio have a high degree of similarity.

[0091] S403, based on the second audio segment with a second similarity greater than the similarity threshold, determine the second audio content in the second song audio that is covered by the first song audio.

[0092] After obtaining the second similarity between each second audio segment and the first song audio, multiple second audio segments can be filtered based on a preset similarity threshold to determine the second audio segments with a second similarity greater than the similarity threshold. Then, based on the second audio segments with a second similarity greater than the similarity threshold, the second audio content in the second song audio that is covered by the first song audio can be determined.

[0093] In this embodiment, the second song audio can be divided into multiple second audio segments, and the similarity between each second audio segment and the first song audio can be determined. This allows for the rapid location of the second audio content covered by the first song audio from different positions in the second song audio, providing a basis for subsequently determining the degree of repetition of the second song audio with respect to the first song audio.

[0094] In one embodiment, determining the second similarity between the second audio segment and the first song audio in step S402 may include the following steps:

[0095] Obtain the segment fingerprint features of the second audio segment and the song fingerprint features of the first song audio; based on the segment fingerprint features of the second audio segment and the song fingerprint features of the first song audio, determine the second similarity between the second audio segment and the first song audio.

[0096] In practical applications, after obtaining multiple second audio segments of the second song audio, a segment fingerprint feature can be obtained for each second audio segment. This segment fingerprint feature can be obtained by extracting features from all audio content in the second audio segment. Furthermore, the song fingerprint feature of the first song audio can also be obtained; that is, features can be extracted from all audio content in the first song audio to generate a fingerprint feature for the entire first song audio, which serves as the song fingerprint feature.

[0097] Since fingerprint features can characterize distinctive audio content, that is, information that distinguishes one audio content from other audio content, in this embodiment, the second similarity between the second audio segment and the first song audio can be determined by comparing the segment fingerprint features of the second audio segment and the song fingerprint features of the first song audio.

[0098] In this embodiment, by comparing the segment fingerprint features of the second audio segment with the song fingerprint features of the first song audio, the audio content in the second audio segment can be quickly compared with the audio content of the first song audio to determine the second similarity.

[0099] In one embodiment, determining the second similarity between the second audio segment and the first song audio based on the segment fingerprint features of the second audio segment and the song fingerprint features of the first song audio may include the following steps:

[0100] S501, obtain the common fingerprint feature between the segment fingerprint feature of the second audio segment and the song fingerprint feature of the first song audio.

[0101] In this embodiment, the common part between the segment fingerprint features of the second audio segment and the song fingerprint features of the first song audio can be obtained, and the fingerprint features of the common part can be used as the common fingerprint features.

[0102] For example, the longest common subsequence of the segment fingerprint features of the second audio segment and the song fingerprint features of the first song audio segment can be obtained, and this longest common subsequence can be used as the common fingerprint feature. The longest common subsequence consists of multiple characters, which are the common parts obtained from the segment fingerprint features and the song fingerprint features. These multiple characters can be continuous or non-contiguous.

[0103] For example, if English characters are used to refer to fingerprint features, in one example, the fragment fingerprint feature of the second audio segment can be ILKN, while the song fingerprint feature of the first song audio can be WEINBKHA. It can be determined that although the fragment fingerprint feature ILKN does not appear continuously in the song fingerprint feature, without changing the order of fingerprint features, both the fragment fingerprint feature of the second audio segment and the song fingerprint feature of the first song audio include the fingerprint feature IK. Therefore, the fingerprint feature of the common part can be regarded as the common fingerprint feature.

[0104] S502, determine the shortest fingerprint feature among the segment fingerprint features of the second audio segment and the song fingerprint features of the first song audio.

[0105] In a specific implementation, the feature length of the segment fingerprint feature of the second audio segment and the feature length of the song fingerprint feature of the first song audio can be obtained. By comparing the feature lengths of the two, the fingerprint feature with the shorter feature length can be determined as the shortest fingerprint feature.

[0106] S503, determine the second similarity between the second audio segment and the first song audio based on the ratio of the feature length of the common fingerprint feature to the feature length of the shortest fingerprint feature.

[0107] After determining the shortest fingerprint feature, the feature length of the common fingerprint feature can be determined, and the ratio of the feature length of the common fingerprint feature to the feature length of the shortest fingerprint feature can be obtained. Then, the second similarity between the second audio segment and the first song audio can be determined based on this ratio.

[0108] In this embodiment, the second similarity between the second audio segment and the first song audio can be determined based on the common fingerprint features between the segment fingerprint features of the second audio segment and the song fingerprint features of the first song audio. This allows for the rapid and accurate determination of the second similarity based on any audio content in the second audio segment that is similar to the first song audio.

[0109] In one embodiment, S103, determining the second degree of repetition of the second song audio relative to the first song audio based on the second audio content in the second song audio that is covered by the first song audio, may include the following steps:

[0110] Determine the number of second audio segments whose second similarity is greater than a similarity threshold; based on the number of second audio segments and the total number of segments of multiple second audio segments of the second song audio, obtain the second degree of repetition of the second song audio relative to the first song audio.

[0111] In a specific implementation, after obtaining the second similarity of multiple second audio segments, the number of second audio segments with a second similarity greater than the similarity threshold can be counted to determine the total number of second audio segments with a second similarity greater than the similarity threshold.

[0112] After obtaining the number of second audio segments, the number of second audio segments can be compared with the total number of segments of multiple second audio segments to determine the proportion of similar second audio segments in all second audio segments. For example, the ratio of the number of second audio segments to the total number of segments can be obtained, thereby determining the second degree of repetition of the second song audio to the first song audio.

[0113] In this embodiment, by comparing the number of second audio segments with a second similarity greater than a similarity threshold with the total number of second audio segments, the proportion of second audio segments that have repeated the audio content of the first song in the entire second song audio can be determined, and the degree of repetition of the second song audio to the first song audio can be accurately identified.

[0114] In one embodiment, S104, determining the song relationship between the first song audio and the second song audio based on the first repetition level, the first duration, the second repetition level, and the second duration, may include the following steps:

[0115] S601, if both the first degree of repetition and the second degree of repetition are greater than the repetition degree threshold, and the duration difference between the first duration and the second duration is less than the duration threshold, then it is determined that the audio of the first song and the audio of the second song are the same.

[0116] In practice, if both the first and second repetition levels are greater than a repetition threshold, it can be determined that most of the content in the first song audio originates from the second song audio, and vice versa. Furthermore, the first duration of the first audio content and the second duration of the second audio content can be compared. Specifically, since the repeated audio content appears in both the first and second song audios, and the first and second durations are determined based on the repeated audio content in both songs, if the difference between the first and second durations is less than a preset duration threshold, it can be determined that the repeated audio content in the first song audio is not significantly different from the repeated audio content in the second song audio. Therefore, if most of the content in the first and second song audios originates from each other, and the difference in the repeated content is not significant, it can be determined that the first and second song audios are identical.

[0117] Specifically, for the first duration, each first audio segment of the first audio content can be obtained, and the duration of the segments of the first audio segments covered by the second audio content can be summed to obtain the first duration; for the second duration, each second audio segment of the second audio content can be obtained, and the duration of the segments of the second audio segments covered by the first audio content can be summed to determine the second duration.

[0118] For example, given a set of audio files containing the first and second songs whose relationships need to be identified. Figure 4a The diagram shows the first similarity scores of each audio segment in the first song's audio (the horizontal axis represents the segment number, and the vertical axis represents the similarity score). Figure 4b The second similarity of each second audio segment in the second song audio is shown; by Figure 4a and Figure 4bIt can be seen that the similarity between the audio segments of the first and second songs is high, meaning both the first and second repetition levels exceed the repetition threshold. The first and second durations are equal, meaning the similarity between the 20 audio segments of the first and second songs is higher than the similarity threshold, and given the same segment length, the first and second durations are equal. Therefore, Figure 4a The corresponding first song audio and Figure 4b The corresponding second song audio is the same audio.

[0119] S602, if the first degree of repetition is greater than the repetition threshold, the second degree of repetition is less than the repetition threshold, and the duration difference between the first duration and the second duration is less than the duration threshold, then the first song audio is determined to be a segment version of the second song audio.

[0120] Among them, the fragment version can refer to an audio version formed by extracting one or more audio segments from the second song audio (the proportion of multiple audio segments in the second song audio is less than the repetition threshold). For example, the second song audio can be composed of second audio segments G, H, I, J, K, etc. By extracting the second audio segments H and I and generating the song audio, the song audio is a fragment version of the second song audio.

[0121] If the first degree of repetition is greater than the repetition threshold, while the second degree of repetition is less than the repetition threshold, then it can be determined that most of the content in the first song audio comes from the second song audio, but only a small amount of audio content in the second song audio comes from the first song audio. In other words, the audio content repeated in the first song audio is only a small part of the content in the second song audio.

[0122] Furthermore, since the extracted audio content appears simultaneously in both the first and second song audios, if the difference between the first and second durations is less than a preset duration threshold, it can be determined that the audio content in the first song audio is not significantly different from a portion of the audio content in the second song audio. Therefore, it can be determined that the first song audio is a fragment version of the second song audio.

[0123] For example, for another set of audio files containing the first and second songs whose relationship needs to be identified, Figure 5a The first similarity scores of each first audio segment in the first song's audio are shown. Figure 5b The second similarity of each second audio segment in the second song audio is shown.

[0124] Depend on Figure 5a and Figure 5bIt can be seen that the first song audio has a high degree of first-order repetition (the similarity of audio segments 0-3 is all above the similarity threshold), while the second song audio has a low degree of second-order repetition (only audio segments 14-18 have a similarity above the similarity threshold). Furthermore, the first duration corresponding to the four audio segments in the first song audio that exceed the similarity threshold is not significantly different from the second duration corresponding to the five audio segments in the second song audio. Therefore, Figure 5a The corresponding first song audio is Figure 5b The corresponding audio snippet of the second song.

[0125] S603, if the first degree of repetition is greater than the repetition degree threshold, the second degree of repetition is less than the repetition degree threshold, and the duration difference between the first duration and the second duration is greater than the duration threshold, then the first song audio is determined to be a repetitive segment version of the second song audio.

[0126] Among them, the repeated segment version can refer to an audio version formed by repeating a certain audio segment in the second song audio multiple times. For example, the second song audio can be composed of second audio segments G, H, I, J, K, etc. By repeatedly repeating the second audio segment G and splicing it into GGG, the song audio composed of multiple audio segments GGG is the repeated segment version of the second song audio.

[0127] Specifically, if the first degree of repetition is greater than the repetition threshold, while the second degree of repetition is less than the repetition threshold, it can be determined that the repeated audio content in the first song audio is only a small part of the content in the second song audio.

[0128] Furthermore, if the difference between the first and second durations exceeds a preset duration threshold, it can be determined that the repeated audio content in the first song audio differs significantly from that in the second song audio in terms of the number of repetitions. For example, the first song audio may contain three identical audio segments 'a', while the second song audio may only contain one audio segment 'a'. Therefore, it can be determined that the first song audio is a repeated segment version of the second song audio.

[0129] For example, for another set of audio files containing the first and second songs whose relationship needs to be identified, Figure 6a The first similarity scores of each first audio segment in the first song's audio are shown. Figure 6b The second similarity of each second audio segment in the second song audio is shown.

[0130] Depend on Figure 6a and Figure 6bIt can be seen that the first song audio has a high degree of repetition (the similarity of all 11 audio segments is higher than the similarity threshold), while the second song audio has a low degree of repetition (only the 4th and 5th audio segments have a similarity higher than the similarity threshold). Furthermore, the first duration corresponding to the 11 audio segments in the first song audio that have a similarity higher than the threshold differs significantly from the second duration corresponding to the 2 audio segments in the second song audio. Therefore, Figure 6a The corresponding first song audio is Figure 6b The corresponding version of the second song audio clip.

[0131] S604, if the first degree of repetition is less than the repetition threshold, the second degree of repetition is greater than the repetition threshold, and the duration difference between the second duration and the first duration is less than the duration threshold, then the first song audio is determined to be a medley version of the second song audio.

[0132] Among them, the medley version can also be called the continuous mix version, which can be made up of audio clips from different songs.

[0133] Conversely, if the first degree of repetition is less than the repetition threshold, while the second degree of repetition is greater than the repetition threshold, it can be determined that only a small portion of the content in the first song audio comes from the second song audio, but a large portion of the audio content in the second song audio comes from the first song audio. In other words, the audio content repeated in the second song audio is only a small portion of the content in the first song audio.

[0134] Furthermore, since the repeated audio content appears simultaneously in both the first and second song audios, if the difference between the first and second durations is less than a preset duration threshold, it can be determined that the repeated audio content in the first song audio is not significantly different from the repeated audio content in the second song audio. Therefore, it can be determined that the first song audio is a medley version of the second song audio.

[0135] For example, for another set of audio files containing the first and second songs whose relationship needs to be identified, Figure 7a The first similarity scores of each first audio segment in the first song's audio are shown. Figure 7b The second similarity of each second audio segment in the second song audio is shown.

[0136] Depend on Figure 7a and Figure 7bIt can be seen that the first repetition level of the first song audio is extremely low (only the 17th audio segment reaches the similarity threshold), while the second repetition level of the second song audio is relatively high (both audio segments have similarities higher than the similarity threshold). Simultaneously, the first duration corresponding to the one audio segment in the first song audio that exceeds the similarity threshold is close to the second duration corresponding to the two audio segments in the second song audio, and the difference in their durations is less than the duration threshold. Therefore, Figure 7a The corresponding first song audio is Figure 7b The corresponding second song audio medley version.

[0137] In this embodiment, by comparing the first degree of repetition and the second degree of repetition, and by comparing the first duration and the second duration, multiple song relationships can be quickly and accurately identified, effectively improving the efficiency of song relationship identification.

[0138] In one embodiment, S101, which involves obtaining the audio of the first song and the audio of the second song to be identified regarding the song relationship, may include the following steps:

[0139] Obtain the slice fingerprint features of each of the multiple audio slices of the first song audio; for each audio slice, determine the candidate song audio that matches the slice fingerprint features of the audio slice from the audio library; based on the candidate song audio associated with each audio slice, obtain the second song audio to be identified in relation to the first song audio.

[0140] In practical applications, after obtaining the first song audio, it can be divided into multiple audio segments, and the fingerprint features of each audio segment can be obtained, resulting in segment fingerprint features for each audio segment. For each audio segment, multiple audio files in the audio library can be searched and matched based on the segment fingerprint features to determine candidate song audio files that match the segment fingerprint features. For example, song audio files with the same or similar fingerprint features can be used as candidate song audio files.

[0141] Furthermore, based on the candidate song audios obtained by matching each audio slice, the second song audio to be identified in relation to the first song audio can be obtained. For example, after obtaining multiple candidate song audios, deduplication can be performed, and the remaining candidate song audios after deduplication can be used as the second song audio.

[0142] In this embodiment, matching candidate song audios can be obtained based on the slice fingerprint features of each audio slice of the first song audio, and the second song audio can be obtained based on the candidate song audios. This can quickly find the second song audio that may have a preset song relationship with the first song audio from massive audio data, avoiding the need to identify song relationships for song audios that have no relation to each other, and improving the identification efficiency.

[0143] In one embodiment, after S104, the method may further include:

[0144] Generate audio content theft tags for the first song audio and / or the second song audio based on the song relationships.

[0145] As an example, an audio content misappropriation tag can indicate how one song's audio content is misappropriated from another song's audio content.

[0146] Specifically, after obtaining the relationship between the first and second song audio files, it is possible to determine how the first song audio files repeat the audio content of the second song audio files. If the first and second song audio files have different publishers, it can be determined that the first song audio files have plagiarized the audio content of the second song audio files.

[0147] In this step, after obtaining the audio of the first song and the audio of the second song, it is possible to further determine whether the publishers of the first song audio and the second song audio are the same. If so, it can be determined that the first song audio was generated by a legitimate publisher who further processed the second song audio. In this case, an audio content theft tag does not need to be generated, but the two can be associated for easier subsequent processing. If not, it can be determined that audio theft has occurred, and an audio content theft tag can be generated based on the song relationship.

[0148] For example, if the identified song relationship is that the first song audio is the same as the second song audio, an audio content theft tag can be generated indicating that the first song audio completely steals the second song audio, or the second song audio completely steals the second song audio. If the first song audio is a duplicate, fragment, or medley version of the second song audio, an audio content theft tag can be generated indicating that the first song audio partially steals the second song audio. Of course, in other examples, the song relationship can also be directly used as the audio content theft tag, thus more clearly indicating the theft method between the song audios.

[0149] In this embodiment, by generating audio content theft tags for the first or second song audio, it is easier to further process the audio songs with theft based on the tags, thereby improving the quality of audio services.

[0150] To enable those skilled in the art to better understand the above steps, the following example illustrates the embodiments of this application, but it should be understood that the embodiments of this application are not limited thereto.

[0151] like Figure 8As shown, after obtaining the audio of the first song, the audio fingerprint of the first song audio can be obtained, including the fingerprint sequence of the first audio segment of the first song audio (i.e., the segment fingerprint feature mentioned above), the fingerprint sequence of the audio slice of the first song audio (i.e., the slice fingerprint feature mentioned above), and the fingerprint sequence of the entire first song audio (i.e., the song fingerprint feature mentioned above). The fingerprint sequence of the first audio segment and the fingerprint sequence of the audio slice of the first song audio can be the same or different, which can be determined by the division method of the audio segment and the audio slice.

[0152] After obtaining the fingerprint sequence of the audio slice of the first song audio, the fingerprint sequence of each audio slice can be matched with the fingerprint sequences of multiple song audios in the audio library to obtain at least one second song audio. Then, the fingerprint sequence of the audio segment of the second song audio and the fingerprint sequence of the entire song audio can be compared with the fingerprint sequence of the entire first song audio and the fingerprint sequence of the first audio segment of the first song audio to determine the song relationship between the first song audio and the second song audio and generate corresponding tags.

[0153] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0154] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 9 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores song audio. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer program implements a song relationship identification method.

[0155] Those skilled in the art will understand that Figure 9 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0156] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:

[0157] Obtain the audio of the first and second songs whose song relationship needs to be identified;

[0158] Based on the first audio content in the first song audio that is covered by the second song audio, determine the first degree of repetition of the first song audio with respect to the second song audio, and obtain the first duration of the first audio content;

[0159] Based on the second audio content in the second song audio that is covered by the first song audio, determine the second degree of repetition of the second song audio relative to the second song audio, and obtain the second duration of the second audio content;

[0160] The song relationship between the first song audio and the second song audio is determined based on the first degree of repetition, the first duration, the second degree of repetition, and the second duration.

[0161] In one embodiment, the processor also performs the steps described in the other embodiments when executing the computer program.

[0162] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, performs the following steps:

[0163] Obtain the audio of the first and second songs whose song relationship needs to be identified;

[0164] Based on the first audio content in the first song audio that is covered by the second song audio, determine the first degree of repetition of the first song audio with respect to the second song audio, and obtain the first duration of the first audio content;

[0165] Based on the second audio content in the second song audio that is covered by the first song audio, determine the second degree of repetition of the second song audio relative to the second song audio, and obtain the second duration of the second audio content;

[0166] The song relationship between the first song audio and the second song audio is determined based on the first degree of repetition, the first duration, the second degree of repetition, and the second duration.

[0167] In one embodiment, the computer program, when executed by a processor, also implements the steps described in the other embodiments above.

[0168] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0169] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0170] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0171] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for identifying song relationships, characterized in that, The method includes: Obtain the audio of the first and second songs whose song relationship needs to be identified; Based on each first audio segment with a first similarity greater than a similarity threshold, the first audio content in the first song audio that is covered by the second song audio is determined; the first similarity is the similarity between the first audio segment in the first song audio and the second song audio. Based on the first audio content, determine the first degree of repetition between the first song audio and the second song audio, and obtain the first duration of the first audio content; Based on each second audio segment with a second similarity greater than a similarity threshold, the second audio content in the second song audio that is covered by the first song audio is determined; the second similarity is the similarity between the second audio segment in the second song audio and the first song audio. Based on the second audio content, determine the second degree of repetition of the second song audio to the first song audio, and obtain the second duration of the second audio content; If both the first degree of repetition and the second degree of repetition are greater than the repetition threshold, and the duration difference between the first duration and the second duration is less than the duration threshold, then it is determined that the audio of the first song and the audio of the second song are the same. If the first degree of repetition is greater than the repetition threshold, the second degree of repetition is less than the repetition threshold, and the duration difference between the first duration and the second duration is less than the duration threshold, then the first song audio is determined to be a segment version of the second song audio. If the first degree of repetition is greater than the repetition threshold, the second degree of repetition is less than the repetition threshold, and the duration difference between the first duration and the second duration is greater than the duration threshold, then the first song audio is determined to be a repetitive segment version of the second song audio. If the first degree of repetition is less than the repetition threshold, the second degree of repetition is greater than the repetition threshold, and the duration difference between the second duration and the first duration is less than the duration threshold, then the first song audio is determined to be a medley version of the second song audio.

2. The method according to claim 1, characterized in that, Before determining the first audio content in the first song audio that is covered by the second song audio based on each first audio segment with a first similarity greater than a similarity threshold, the method further includes: Obtain multiple first audio segments of the first song audio; For each first audio segment, determine the first similarity between the first audio segment and the audio of the second song.

3. The method according to claim 2, characterized in that, Determining the first similarity between the first audio segment and the second song audio includes: Obtain the segment fingerprint feature of the first audio segment, and obtain the song fingerprint feature of the second song audio; Based on the segment fingerprint features of the first audio segment and the song fingerprint features of the second song audio, a first similarity between the first audio segment and the second song audio is determined.

4. The method according to claim 3, characterized in that, The step of determining the first similarity between the first audio segment and the second song audio based on the segment fingerprint features of the first audio segment and the song fingerprint features of the second song audio includes: Obtain the common fingerprint feature between the segment fingerprint feature of the first audio segment and the song fingerprint feature of the second song audio; Determine the shortest fingerprint feature between the segment fingerprint feature of the first audio segment and the song fingerprint feature of the second song audio. The first similarity between the first audio segment and the second song audio is determined based on the ratio of the feature length of the common fingerprint feature to the feature length of the shortest fingerprint feature.

5. The method according to claim 1, characterized in that, The step of determining the degree of repetition between the first song audio and the second song audio based on the first audio content includes: Determine the number of first audio segments whose first similarity is greater than a similarity threshold; Based on the number of the first audio segments and the total number of segments of the first audio segments of the first song audio, the first repetition degree of the first song audio relative to the second song audio is obtained.

6. The method according to claim 1, characterized in that, The process of obtaining the first and second song audios to identify the relationship between the songs includes: Obtain the slice fingerprint features of each of the multiple audio slices of the first song's audio. For each audio slice, candidate song audios that match the slice fingerprint features of the audio slice are identified from the audio library; Based on the candidate song audios associated with each audio slice, obtain the second song audio that is to be identified as having a song relationship with the first song audio.

7. The method according to any one of claims 1-6, characterized in that, After determining the song relationship between the first song audio and the second song audio based on the first repetition degree, the first duration, the second repetition degree, and the second duration, the method further includes: Based on the song relationships, generate audio content theft tags for the audio of the first song and / or the audio of the second song.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.