Pirated audio detection method, device and computer program product

By recalling similar audios in the audio library that match the audio characteristics to be detected and using the audio classification model to be identified, the problem of low detection efficiency of pirated audio in the prior art is solved, and efficient and accurate detection of pirated audio in the audio library is achieved.

CN114925231BActive Publication Date: 2025-05-16TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210567984.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-24
Publication Date
2025-05-16
Estimated Expiration
2042-05-24

AI Technical Summary

Technical Problem

In the prior art, pirated audio detection efficiency is low and cannot effectively process massive audio data, resulting in the inability to cope with the detection requirements of incremental data in the audio library.

Method used

By recalling similar audio matching the audio characteristics to be detected from the audio library, forming the audio group to be detected, and entering the audio diagram into the trained audio classification model, obtaining the audio classification results, determining the benchmark audio, and finally identifying the pirated audio in the audio group to be detected.

Benefits of technology

It realizes efficient and accurate detection of pirated audio in the audio library, and the automated process can process massive data, improving detection efficiency and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114925231B_ABST
    Figure CN114925231B_ABST
Patent Text Reader

Abstract

The present application relates to the field of audio technology, and provides a method for detecting pirated audio, a computer device, and a computer program product. The present application can realize efficient and accurate detection of pirated audio in an audio library. The method comprises: determining similar audio that matches the audio features of the audio to be detected from the audio library, obtaining an audio group to be detected including the audio to be detected and its similar audio, and then inputting the audio image of each audio in the audio group to be detected into a trained audio classification model, obtaining the audio classification result of each audio in the audio group to be detected output by the model, determining the benchmark audio in the audio group to be detected according to the classification result, and identifying the pirated audio in the audio group to be detected based on the benchmark audio.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of audio technology, and in particular to a pirated audio detection method, a computer device, and a computer program product. Background Art

[0002] With the development of the Internet and audio technology, various audio applications provide users with a variety of audio services such as audio playback, audio search, and song identification. Accurate and efficient detection of pirated audio can improve the quality of audio services and user experience.

[0003] Current technology mainly uses manual review to identify pirated audio in the audio library. However, this method is time-consuming and labor-intensive, has the technical problem of low detection efficiency, and cannot cope with the massive amount of stock data and daily incremental data in the audio library. Summary of the invention

[0004] Based on this, it is necessary to provide a pirated audio detection method, computer equipment and computer program product to address the above technical issues.

[0005] In a first aspect, the present application provides a method for detecting pirated audio. The method comprises:

[0006] Determine similar audio that matches the audio feature of the audio to be tested from the audio library, and obtain an audio group to be tested that includes the audio to be tested and the similar audio;

[0007] Inputting the audio image of each audio in the audio group to be inspected into the trained audio classification model, and obtaining the audio classification result of each audio in the audio group to be inspected output by the audio classification model; the audio classification result is used to indicate whether the audio is genuine audio;

[0008] Determining benchmark audio in the audio group to be tested according to the audio classification result;

[0009] Based on the benchmark audio, pirated audio in the audio group to be checked is identified.

[0010] In a second aspect, the present application further provides a computer device. The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the following steps when executing the computer program:

[0011] Determine similar audio that matches the audio features of the audio to be tested from the audio library, and obtain an audio group to be tested that includes the audio to be tested and the similar audio; input the audio image of each audio in the audio group to be tested into a trained audio classification model, and obtain the audio classification result of each audio in the audio group to be tested output by the audio classification model; the audio classification result is used to indicate whether the audio is genuine audio; determine the benchmark audio in the audio group to be tested based on the audio classification result; and identify pirated audio in the audio group to be tested based on the benchmark audio.

[0012] In a third aspect, the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0013] Determine similar audio that matches the audio features of the audio to be tested from the audio library, and obtain an audio group to be tested that includes the audio to be tested and the similar audio; input the audio image of each audio in the audio group to be tested into a trained audio classification model, and obtain the audio classification result of each audio in the audio group to be tested output by the audio classification model; the audio classification result is used to indicate whether the audio is genuine audio; determine the benchmark audio in the audio group to be tested based on the audio classification result; and identify pirated audio in the audio group to be tested based on the benchmark audio.

[0014] The above-mentioned pirated audio detection method, computer device and computer program product determine similar audio that matches the audio features of the audio to be tested from the audio library, obtain an audio group to be tested including the audio to be tested and its similar audio, then input the audio image of each audio in the audio group to be tested into the trained audio classification model to obtain the audio classification result of each audio in the audio group to be tested output by the model, determine the benchmark audio in the audio group to be tested based on the classification result, and identify the pirated audio in the audio group to be tested based on the benchmark audio. This scheme uses an audio recognition system to recall similar audio that matches the audio features of the audio to be tested from the audio library and form an audio group to be tested with it, then input the image of each audio in the audio group to be tested into the audio classification model to determine the benchmark audio in the audio group to be tested, and finally determine whether other audio in the audio group to be tested is pirated audio based on the benchmark audio. The entire detection process is automated, which can not only detect whether the audio to be tested is pirated audio, but also detect whether the recalled similar audio is pirated audio, thereby achieving efficient and accurate detection of pirated audio in the audio library. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 A schematic diagram of a process of detecting pirated audio in one embodiment;

[0016] Figure 2A schematic flow chart of the steps of recalling similar audio in one embodiment;

[0017] Figure 3 A schematic flow chart of the steps of determining audio of the same type in one embodiment;

[0018] Figure 4 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0019] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0020] The pirated audio detection method provided in the embodiment of the present application can be applied to a computer device such as a server for execution. The server can be implemented as an independent server or a server cluster composed of multiple servers.

[0021] In one embodiment, Figure 1 As shown, a pirated audio detection method is provided, comprising the following steps:

[0022] Step S101, determining similar audio that matches the audio feature of the audio to be tested from the audio library, and obtaining an audio group to be tested including the audio to be tested and the similar audio.

[0023] This step is mainly to recall the audio that matches the audio features of the audio to be tested from the audio library after obtaining the audio to be tested. The audio that matches the audio features of the audio to be tested is called the same type of audio. The number of the same type of audio is generally multiple, and then the audio to be tested and the multiple same type of audio form an audio group to be tested. In practical applications, an audio recognition system can be used to determine the same type of audio from the audio library. The audio recognition system refers to a system with audio recognition capabilities, which is mainly used to identify audio similar to the audio to be tested from the audio library. The audio recognition system specifically identifies similar audio in the audio library as the same type of audio based on the audio features of the audio to be tested. Among them, for audio features, audio features such as Landmark audio fingerprints can be used.

[0024] In a specific application, taking songs as audio as an example, the audio recognition system can adopt a landmark-based song recognition system. After obtaining a song to be tested, the landmark-based song recognition system can determine songs that match the landmark audio fingerprint of the song from the song library as similar songs, and these similar songs and the song to be tested are combined into a song group to be tested.

[0025] Step S102: input the audio image of each audio in the audio group to be tested into the trained audio classification model to obtain the audio classification result of each audio in the audio group to be tested output by the audio classification model.

[0026] As an example, if the audio in the audio group to be tested is a song, the audio picture can be the album picture of each song. For example, the album picture of the song can include at least one of the following information: singer portrait, singer name, album name, album release time, and the names of each song included in the album.

[0027] In this step, the audio images of each audio in the audio group to be tested can be obtained, and the audio images of each audio in the audio group to be tested can be input into the trained audio classification model, and the audio classification model outputs the audio classification result of the corresponding audio according to the audio images, thereby obtaining the audio classification result of each audio in the audio group to be tested, and the audio classification result can indicate whether the corresponding audio is genuine audio.

[0028] Specifically, for the audio classification model, a residual network model used for target classification in the field of computer vision can be used, which can be a binary classification model based on audio images. In the specific implementation, a Resnet50 deep neural network model or other models can be used. Although the binary classification model based on audio images can relatively accurately distinguish whether the audio is genuine audio or pirated audio, in order to further improve the recognition accuracy of pirated audio and avoid identifying pirated audio based on the single dimension of audio images, this embodiment can first determine the genuine audio from the audio group to be tested based on the audio classification results.

[0029] Among them, pirated audio refers to the audio content that is consistent with the original audio / genuine audio, but the other information of the audio is inconsistent. Taking songs as an example, pirated songs are consistent with the original song audio, but the singer name and song name are inconsistent with the original song. Such songs are called pirated songs.

[0030] Step S103: determining the benchmark audio in the audio group to be tested according to the audio classification result.

[0031] In this step, after obtaining the audio classification results of each audio in the audio group to be tested output by the audio classification model, it can be determined whether the audio is a benchmark audio based on whether the audio classification result indicates that the corresponding audio is a genuine audio. Specifically, if the audio classification result of an audio in the audio group to be tested is a genuine audio, then the audio can be determined as the benchmark audio in the audio group to be tested. If there are multiple audios in the audio group to be tested whose audio classification results are genuine audio, then the benchmark audio can be further determined from the multiple genuine audios.

[0032] Step S105: identifying pirated audio in the audio group to be checked based on the benchmark audio.

[0033] This step mainly utilizes the benchmark audio in the audio group to be tested determined in step S104 to identify whether other audio in the audio group to be tested is pirated audio. Specifically, the benchmark audio can be compared with other audio in the audio group to be tested based on other audio information, and whether it is pirated audio can be determined based on the comparison result. Exemplarily, other audio information can be information such as the audio publisher and creator, and the audio publisher and creator of the benchmark audio can be compared with the audio publisher and creator of other audio in the audio group to be tested, and other audio whose audio publisher and creator are inconsistent with the benchmark audio can be identified as pirated audio.

[0034] The above-mentioned pirated audio detection method recalls similar audios that match the audio features of the audio to be tested from the audio library and forms an audio group to be tested with them, then inputs the accompanying images of each audio in the audio group to be tested into the audio classification model to determine the benchmark audio in the audio group to be tested, and finally determines whether other audios in the audio group to be tested are pirated audios based on the benchmark audio. The entire detection process is automated, which can not only detect whether the audio to be tested is pirated audio, but also detect whether the recalled similar audio is pirated audio, thereby achieving efficient and accurate detection of pirated audio in the audio library.

[0035] Regarding the step S101 of determining the same type of audio that matches the audio feature of the audio to be detected from the audio library, in one embodiment, Figure 2 As shown, the following steps may be included:

[0036] Step S201, segment the audio to be inspected to obtain multiple audio segments.

[0037] The audio to be tested has a certain duration, and the audio to be tested can be segmented at a certain duration interval to obtain multiple audio segments. For example, for a song to be tested with a duration of 3 minutes, the song to be tested can be segmented at a duration interval of 6 seconds to obtain a total of 30 audio segments.

[0038] Step S202: for each audio segment, determine a candidate audio group of the same type that matches the audio feature of the audio segment from the audio library to obtain a plurality of candidate audio groups of the same type.

[0039] In this step, for each audio segment obtained by segmentation in step S201, audio matching its audio features can be recalled from the audio library as candidate similar audio, and the number of candidate similar audio is generally multiple, and these candidate similar audios are used to form candidate similar audio groups for corresponding audio segments, and each audio segment corresponds to a candidate similar audio group, and multiple candidate similar audio groups are obtained. In one example, multiple audio segments can be input into an audio recognition system respectively, and the audio recognition system matches the candidate similar audio of the current audio segment from the audio library, and the current audio segment is an audio segment among the multiple audio segments.

[0040] Step S203, obtaining similar audio that matches the audio feature of the audio to be detected based on the candidate similar audios contained in each candidate similar audio group.

[0041] Specifically, each audio clip corresponds to a candidate audio group of the same type. In this step, the candidate audios of the same type contained in each candidate audio group of the same type can be used as the audio of the same type that matches the audio features of the audio to be tested. In a specific implementation, if the audio is represented by an audio identification number (audio ID), the audio IDs of each candidate audio group of the same type are intersected, and the audio corresponding to the audio ID in the intersection is used as the audio of the same type that matches the audio features of the audio to be tested.

[0042] This embodiment divides the audio to be tested into segments to recall candidate similar audio groups that match each audio segment, and then obtains similar audio of the audio to be tested based on the candidate similar audio contained in each candidate similar audio group, thereby improving the accuracy of recalling similar audio of the audio to be tested in the audio library.

[0043] Further, in one embodiment, if Figure 3 As shown, the above step S203 may specifically include:

[0044] Step S301, obtaining the feature matching time of each candidate audio of the same type and multiple audio segments in each candidate audio group, and the matching degree of each candidate audio of the same type.

[0045] As described in the above embodiment, audios matching the audio features can be recalled from the audio library as candidate audios of the same type, and these candidate audios of the same type will constitute candidate audio groups of the corresponding audio segments. Wherein, while recalling the candidate audios of the same type, the matching degree of the candidate audios of the same type and the corresponding audio segments can also be output, for example, the matching degree of each candidate audio of the same type can be output by the audio recognition system. In addition, when determining the candidate audios of the same type, the candidate audios of the same type are recalled in the audio library based on the audio features of the audio segments, and the matching of the audio features can correspond to a certain time point of the audio, which is called the feature matching time. For example, when a certain pitch feature of the audio segment is identified to match the pitch feature of an audio in the audio library, the audio can be used as a candidate audio of the same type, and the time point corresponding to the pitch feature in the audio segment can also be obtained, and the time point corresponding to the pitch feature in the candidate audio of the same type can be output, such as the pitch feature appears at a moment a in the audio segment and appears at a moment b in the candidate audio of the same type, and the a moment and the b moment are called the feature matching time. Thus, this step can also obtain the feature matching time of each candidate audio of the same type and multiple audio segments in each candidate audio group.

[0046] Step S302: for each candidate similar audio group, determine the candidate similar audios in the candidate similar audio group whose matching degree satisfies the matching degree threshold condition and / or whose feature matching time difference satisfies the time difference threshold condition, and obtain a screened candidate similar audio group.

[0047] This step is to further screen each candidate audio of the same type in each recalled candidate audio group, wherein the screening basis may include at least one of the matching degree and the feature matching time difference.

[0048] If the screening process is based on the matching degree, this process can be called matching degree screening. When screening based on the matching degree, all candidate similar audios in each group of candidate similar audio groups are filtered / screened through a matching degree threshold. For any candidate similar audio, its matching degree is compared with the matching degree threshold. If its matching degree is greater than the matching degree threshold, it can be determined that the candidate similar audio meets the matching degree threshold condition, and the candidate similar audio is retained in the candidate similar audio group to which it belongs; if its matching degree is less than or equal to the matching degree threshold, the candidate similar audio can be removed from the candidate similar audio group to which it belongs. After processing this process, each candidate similar audio group screened by matching degree can be obtained.

[0049] If the screening is based on the feature matching time difference, this process can also be called matching time screening. Specifically, for each group of candidate similar audios in the candidate similar audio group that has been screened by the matching degree, they are filtered / screened by a time difference threshold. For a candidate similar audio in any group, the difference between the feature matching time of the candidate similar audio and the feature matching time of the corresponding audio segment is calculated to obtain the feature matching time difference of the candidate similar audio, and then the feature matching time difference is compared with the aforementioned time difference threshold. If its feature matching time difference is less than the time difference threshold, it can be determined that the candidate similar audio meets the time difference threshold condition, and the candidate similar audio is retained in the candidate similar audio group that has been screened by the matching degree; if its feature matching time difference is greater than or equal to the time difference threshold, the candidate similar audio can be removed from the candidate similar audio group that has been screened by the matching degree. After this process, each candidate similar audio group that has been screened by the matching time can be obtained.

[0050] Therefore, by screening and filtering the candidate similar audios in each candidate similar audio group by at least one of the above methods, a screened candidate similar audio group can be obtained.

[0051] Step S303, obtaining similar audio that matches the audio feature of the audio to be detected based on the candidate similar audios contained in each screened candidate similar audio group.

[0052] In this step, the candidate audios of the same type contained in each of the screened candidate audio groups of the same type may be used as the audios of the same type that match the audio features of the audio to be detected.

[0053] This embodiment further combines at least one of the matching degree and the feature matching time on the basis of recalling candidate similar audios by audio segment matching, and obtains similar audios after screening the candidate similar audios, thereby ensuring the similarity between the audio to be tested and the recalled similar audios, and further improving the accuracy of recalling similar audios from the audio library.

[0054] The process of recalling similar audio as described in steps S201 to S203 and steps S301 to S303 is described below by taking the song to be tested as the audio to be tested, the Landmark-based song recognition system as the audio recognition system, and the Landmark audio fingerprint as the audio feature as an example:

[0055] First, the song to be tested is segmented to obtain multiple song fragments to be tested, and then each song fragment to be tested is matched through a landmark-based song recognition system, and the landmark-based song recognition system returns the matching song identifier, feature matching time and matching degree.

[0056] Among them, for the matching process of the song clip to be tested by the Landmark-based song recognition system, specifically, the audio feature adopted by the Landmark-based song recognition system is the Landmark audio fingerprint. For any audio, the extraction process of the Landmark audio fingerprint is mainly to transform the time domain audio into the frequency domain through Fourier transform to obtain the time-frequency spectrum. The horizontal axis of the time-frequency spectrum represents the time index and the vertical axis represents the frequency index. Then, the local peak point of the time-frequency point signal in the time-frequency spectrum is extracted, and the remaining peak points around the local peak point are found to form a peak point combination (t1, f1, t2, f2), that is, both the time-frequency point (t1, f1) and the time-frequency point (t2, f2) have peak points. The information is further simplified into simplified information such as (f1, f2, t2-t1), and then the simplified information (f1, f2, t2-t1) can be hashed to obtain a hash value. In this way, the corresponding hash value can be calculated for each song in the song library, and then the song library index can be constructed according to the hash value of each song. For the song fragment to be tested, the song recognition system can calculate its hash value and use it as the index hash value, and use the index hash value to match the song library index, and return the matching song identifier, feature matching time and matching degree.

[0057] At this point, each song segment to be tested will correspond to a group of candidate similar song groups recalled. In order to ensure the consistency between the songs to be tested and the recalled similar songs, the candidate similar songs recalled for each song segment to be tested can be first screened by a matching threshold, and the candidate similar songs with a matching degree greater than the matching threshold are retained to obtain each candidate similar song group screened by the matching degree. Then, for each candidate similar song in each candidate similar song group screened by the matching degree, the difference between the feature matching time of the candidate similar song and the feature matching time of the corresponding song segment to be tested is calculated to obtain the feature matching time difference, and the candidate similar songs with a feature matching time difference less than the time difference threshold are retained to obtain each candidate similar song group screened by the matching time. Finally, the candidate similar songs contained in each candidate similar song group screened by the matching time are taken as the same songs as the song to be tested.

[0058] In one embodiment, determining the benchmark audio in the audio group to be tested according to the audio classification result in step S103 specifically includes:

[0059] The audio heat of each audio in the audio group to be tested is obtained; and the benchmark audio in the audio group to be tested is determined according to the audio heat and the audio classification result.

[0060] In this embodiment, the audio heat of each audio in the audio group to be tested is obtained, and the benchmark audio in the audio group to be tested is determined by combining the audio heat and the audio classification results provided by the audio classification model. In a specific application, the audio heat can be determined based on information such as the number of plays, the number of clicks, and the release date of the audio. This embodiment determines the benchmark audio in the audio group to be tested by combining the audio heat and the audio classification results provided by the model, which can avoid the problem of affecting the accuracy of pirated audio detection by relying solely on the single dimension of the audio classification results provided by the model, and relatively speaking, the original audio is more sophisticated and has a higher heat, so that the benchmark audio in the audio group to be tested can be more accurately determined, further improving the robustness of pirated audio detection.

[0061] Furthermore, in one embodiment, the above embodiment determines the benchmark audio in the audio group to be inspected according to the audio heat and the audio classification result, specifically including:

[0062] The audio in the audio group to be tested whose audio heat meets the preset heat condition and whose audio classification result is genuine audio is determined as the benchmark audio.

[0063] This embodiment mainly selects audios that are determined to be genuine audios by the audio classification model and have high audio popularity as benchmark audios from the audio group to be tested. Specifically, audios with audio popularity greater than or equal to a preset popularity threshold can be determined as audios that meet the preset popularity conditions, and audios with the highest audio popularity can also be determined as audios that meet the preset popularity conditions. Then, audios that meet the preset popularity conditions can be selected from at least one audio that is classified as genuine audio as a benchmark audio.

[0064] In one example, when obtaining benchmark audio from multiple genuine audios, only one genuine audio can be selected as the benchmark audio. Specifically, the audio content of each audio in the audio group to be tested is consistent, and the difference can only be the difference in the audio description information of each audio. It can be understood that each audio is actually generated based on the same audio content. Therefore, the most credible audio can be obtained from multiple genuine audios. This audio can be regarded as the audio content source of other similar audios in the audio group to be tested and used as the benchmark audio.

[0065] In one embodiment, the step S105 of identifying pirated audio in the audio group to be checked based on the benchmark audio may include:

[0066] Obtain the benchmark audio and audio description information of the audio to be compared in the audio group to be tested; determine whether the audio to be compared is pirated audio based on the consistency between the audio description information of the audio to be compared and the audio description information of the benchmark audio.

[0067] In this embodiment, after identifying the benchmark audio in the audio group to be tested, other audio in the audio group to be tested except the benchmark audio can be used as the audio to be compared, and the audio description information of each audio to be compared is compared with the audio description information of the benchmark audio to determine the consistency between the two, and determine whether the audio to be compared is pirated audio based on the consistency. Specifically, the audio description information can be text information used to describe the audio, and can include the duration of the audio, the creator of the audio, the name of the audio, etc. By comparing whether the audio description information is consistent, it is determined whether the audio to be compared in the audio group to be tested is pirated audio. If the audio description information of the two is consistent, it can be determined that the audio to be compared is the original audio. If the audio description information of the two is inconsistent, it can be determined that the audio to be compared is pirated audio. In this embodiment, combined with the characteristics of pirated audio that tampered with the audio description information of the original audio, the benchmark audio and the audio to be compared are compared for consistency of audio description information, so as to more accurately and efficiently identify the pirated audio in the audio group to be tested.

[0068] Furthermore, in some embodiments, the audio description information may include multiple audio description sub-information corresponding to multiple audio description items, that is, the audio description information may include multiple audio description sub-information. Taking a song as an audio, for example, the audio description item may include the singer name and song title of the song. The audio description sub-information corresponding to the singer name refers to the specific name of the singer, and the audio description sub-information corresponding to the song title refers to the specific name of the song. Based on this, the above embodiment determines whether the audio to be compared is pirated audio based on the consistency between the audio description information of the audio to be compared and the audio description information of the benchmark audio, including:

[0069] If the audio description sub-information corresponding to the audio to be compared and the benchmark audio in at least one audio description item is inconsistent, it is determined that the audio to be compared is pirated audio.

[0070] In this embodiment, specifically, for each audio description item, the audio description sub-information of the audio to be compared and the benchmark audio can be compared, and each audio description item can correspond to a comparison result, which indicates whether the audio description sub-information of the two is consistent on the audio description item. Among them, if the audio description sub-information corresponding to the audio to be compared and the benchmark audio on at least one audio description item is inconsistent, such as the name of the singer is inconsistent, or the name of the song is inconsistent, then it can be determined that the audio to be compared is pirated audio. If the audio description sub-information corresponding to the audio to be compared and the benchmark audio on each audio description item is consistent, then it can be determined that the audio to be compared is original audio. The solution of this embodiment allows users to set multiple audio description items in a targeted manner to more accurately and efficiently detect pirated audio in the audio group to be tested.

[0071] In one embodiment, the above method may further include the following steps of training an audio classification model, specifically including:

[0072] Obtain audio images of genuine audio and pirated audio; use the audio images of genuine audio as positive samples, and use the audio images of pirated audio as negative samples; based on a preset loss function, train the audio classification model to be trained using the positive samples and the negative samples, and obtain the audio classification model when the loss value of the preset loss function meets the preset loss threshold condition.

[0073] Specifically, taking into account the characteristics of pirated audio, in addition to tampering with the audio name, publisher name and other information of the genuine audio, pictures different from the genuine audio will be used as audio illustrations. This embodiment obtains an audio classification model based on audio illustration training of genuine audio and pirated audio, which is used to identify genuine and pirated audio based on audio illustrations.

[0074] In this embodiment, audio images of genuine audio and pirated audio are first obtained, the audio images of genuine audio are taken as positive samples, and the audio images of pirated audio are taken as negative samples. The residual network model used for target classification, such as Resnet50, is used as the audio classification model to be trained. The preset loss function may adopt a cross-entropy loss function. Based on the cross-entropy loss function, the audio classification model to be trained is trained using positive samples and negative samples. When the loss value of the preset loss function meets the preset loss threshold condition, such as the loss value is less than or equal to the preset loss threshold, a trained audio classification model based on the audio image is obtained. The model may be a binary classification model based on the audio image, which is used to identify whether the audio is genuine based on the input audio image.

[0075] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0076] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4As shown. The computer device includes a processor, a memory and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store audio and other data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a pirated audio detection method is implemented.

[0077] Those skilled in the art will understand that Figure 4 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0078] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.

[0079] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0080] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.

[0081] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0082] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0083] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A method for detecting pirated audio, characterized in that: The method comprises: Determine similar audio that matches the audio feature of the audio to be tested from the audio library, and obtain an audio group to be tested that includes the audio to be tested and the similar audio; Inputting the audio image of each audio in the audio group to be inspected into the trained audio classification model, and obtaining the audio classification result of each audio in the audio group to be inspected output by the audio classification model; the audio classification result is used to indicate whether the audio is genuine audio; Determining benchmark audio in the audio group to be tested according to the audio classification result; Based on the benchmark audio, pirated audio in the audio group to be checked is identified.

2. The method according to claim 1, characterized in that The step of determining the benchmark audio in the to-be-tested audio group according to the audio classification result includes: Obtaining the audio heat of each audio in the audio group to be checked; According to the audio heat and the audio classification result, the benchmark audio in the audio group to be tested is determined.

3. The method according to claim 2, characterized in that The step of determining the benchmark audio in the to-be-tested audio group according to the audio heat and the audio classification result includes: The audio in the audio group to be inspected whose audio popularity meets the preset popularity condition and whose audio classification result is genuine audio is determined as the benchmark audio.

4. The method according to any one of claims 1 to 3, characterized in that: The step of identifying pirated audio in the to-be-checked audio group based on the benchmark audio includes: Acquire audio description information of the benchmark audio and the audio to be compared in the audio group to be tested, where the audio to be compared is other audio except the benchmark audio; Whether the audio to be compared is pirated audio is determined based on the consistency between the audio description information of the audio to be compared and the audio description information of the benchmark audio.

5. The method according to claim 4, characterized in that The audio description information includes a plurality of audio description sub-information corresponding to a plurality of audio description items; and determining whether the audio to be compared is pirated audio according to the consistency between the audio description information of the audio to be compared and the audio description information of the benchmark audio includes: If the audio to be compared is inconsistent with the audio description sub-information corresponding to the benchmark audio in at least one audio description item, it is determined that the audio to be compared is pirated audio.

6. The method according to claim 1, characterized in that The step of determining similar audio that matches the audio feature of the audio to be detected from the audio library includes: Segmenting the audio to be inspected to obtain multiple audio segments; For each audio segment, determining a candidate audio group of the same type that matches the audio feature of the audio segment from the audio library to obtain a plurality of candidate audio groups of the same type; According to the candidate audios of the same type contained in each candidate audio group of the same type, the audio of the same type that matches the audio feature of the audio to be detected is obtained.

7. The method according to claim 6, characterized in that The obtaining, based on the candidate audios of the same type commonly included in each candidate audio group of the same type, the audio of the same type that matches the audio feature of the audio to be detected comprises: Obtaining feature matching time of each candidate audio of the same type in each candidate audio group and the multiple audio segments, and a matching degree of each candidate audio of the same type; For each candidate similar audio group, determine the candidate similar audio in the candidate similar audio group whose matching degree satisfies the matching degree threshold condition and / or whose feature matching time difference satisfies the time difference threshold condition, to obtain a screened candidate similar audio group; wherein the feature matching time difference is the difference between the feature matching time of the candidate similar audio and the feature matching time of the corresponding audio segment; Based on the candidate audios of the same type contained in each of the screened candidate audio groups of the same type, the audios of the same type that match the audio features of the audio to be detected are obtained.

8. The method according to claim 1, characterized in that The method further comprises: Get audio illustrations of genuine and pirated audio; The audio image of the genuine audio is used as a positive sample, and the audio image of the pirated audio is used as a negative sample; Based on a preset loss function, the audio classification model to be trained is trained using the positive samples and the negative samples, and when the loss value of the preset loss function meets a preset loss threshold condition, an audio classification model is obtained.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Music copyright identification method based on characteristic

    CN107967922A

  • Audio classification method and device and storage medium

    CN112380382A