Processing Method, Device, Equipment and Storage Medium Based on Audio Fingerprint
By making matching judgments and updating the number of matches before the audio fingerprint is entered into the library, the problem of insufficient storage space and low recognition rate of the audio fingerprint database is solved, and efficient clustering and accurate recognition of audio fingerprints are achieved.
Patent Information
- Application Number
- CN202210010673.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-06
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-01-06
AI Technical Summary
Existing audio fingerprint databases are prone to fingerprint coverage and incomplete matching problems when storing a large amount of new fingerprint information, resulting in a lower recognition rate.
Before entering the audio fingerprint information, determine whether there is a matching fingerprint information in the preset fingerprint library. If it exists, it is prohibited to enter the library, and the number of matches is updated in the resource information library to achieve clustering of audio fingerprints and saving storage space.
It effectively saves the storage space of the fingerprint library, improves the accuracy and efficiency of audio recognition, and provides auxiliary information for music copyright identification and audio recommendation tasks, reducing the cost of third-party service query.
Smart Images

Figure CN114420166B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of audio processing, and in particular, to a processing method, apparatus, device, and storage medium based on audio fingerprints. Background Art
[0002] Audio fingerprint algorithms can extract specific acoustic features in audio as its "fingerprint" and search and match it with the "fingerprints" in the database, so as to achieve the purpose of identifying a large number of music. At present, it has been widely used in related tasks such as music retrieval, monitoring of copyrighted content, and deduplication of content libraries.
[0003] Currently, when fingerprint information that is different from the stored fingerprint information appears, the new fingerprint information will be stored in the fingerprint database. However, the size of the database for storing fingerprints is limited. In many application scenarios, the number of audio increases rapidly, and the number of fingerprints stored in the database also increases accordingly. When the database cannot store more new fingerprint information, a large number of fingerprint information will be overwritten and replaced, which easily leads to incomplete fingerprints of some audio, increasing the difficulty of matching and reducing the recognition rate. Summary of the Invention
[0004] Embodiments of the present invention provide a processing method, apparatus, device, and storage medium based on audio fingerprints, which can optimize the existing processing scheme based on audio fingerprints.
[0005] In a first aspect, embodiments of the present invention provide a processing method based on audio fingerprints, the method comprising:
[0006] Determining first fingerprint information of a first audio resource to be stored in the database;
[0007] Matching the first fingerprint information with each fingerprint information in a preset fingerprint database;
[0008] If there is second fingerprint information that matches the first fingerprint information successfully, then prohibiting the execution of the operation of storing the first fingerprint information in the database, and updating the matching times of the in-database resource identifier associated with the second fingerprint information in a preset in-database resource information database.
[0009] In a second aspect, embodiments of the present invention provide a processing apparatus based on audio fingerprints, the apparatus comprising:
[0010] A fingerprint information determination module, configured to determine first fingerprint information of a first audio resource to be stored in the database;
[0011] A fingerprint information matching module, configured to match the first fingerprint information with each fingerprint information in a preset fingerprint database;
[0012] The warehousing control module is used to, if there is second fingerprint information that matches the first fingerprint information successfully, prohibit the execution of the warehousing operation of the first fingerprint information, and update the matching times of the in-warehouse resource identifier associated with the second fingerprint information in a preset warehousing resource information database.
[0013] In a third aspect, an embodiment of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the processing method based on audio fingerprints provided by the embodiment of the present invention is implemented.
[0014] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the processing method based on audio fingerprints provided by the embodiment of the present invention is implemented.
[0015] In the processing solution based on audio fingerprints provided by the embodiment of the present invention, the first fingerprint information of the first audio resource to be warehoused is determined, and the first fingerprint information is matched with each fingerprint information in a preset fingerprint database. If there is second fingerprint information that matches the first fingerprint information successfully, the execution of the warehousing operation of the first fingerprint information is prohibited, and the matching times of the in-warehouse resource identifier associated with the second fingerprint information are updated in a preset warehousing resource information database. By adopting the above technical solution, when audio fingerprint information needs to be warehoused, it is first determined whether there is matching fingerprint information in the fingerprint database. If there is, the warehousing is prohibited, and the matching times of the associated in-warehouse resource identifier are updated in the resource information database. The audio fingerprints are clustered through the in-warehouse resource identifier in the resource information database, thereby effectively saving the storage space of the fingerprint database. Moreover, for specific business scenarios such as music copyright identification tasks and audio recommendation tasks, effective auxiliary information can be provided, reducing the cost of querying copyright information by third-party services and improving the accuracy of providing audio recommendations, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a schematic flowchart of a processing method based on audio fingerprints provided by an embodiment of the present invention;
[0017] Figure 2 It is a schematic flowchart of another processing method based on audio fingerprints provided by an embodiment of the present invention;
[0018] Figure 3 It is a structural block diagram of a processing device based on audio fingerprints provided by an embodiment of the present invention;
[0019] Figure 4 It is a structural block diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. In addition, it should be noted that for the sake of description, only parts related to the present invention rather than all structures are shown in the drawings. In addition, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0021] Figure 1 FIG. is a schematic flow chart of a processing method based on audio fingerprint provided by an embodiment of the present invention. This method can be executed by a processing device based on audio fingerprint, where the device can be implemented by software and / or hardware and is generally integrated in a computer device. As Figure 1 shown, the method includes:
[0022] Step 101, determine the first fingerprint information of the first audio resource to be stored in the library.
[0023] In the embodiment of the present invention, the specific type of the audio resource is not limited. It may include the audio data in the audio file, and may also include the audio data in the video file. For example, the background music data in the video file corresponding to the short video, etc.
[0024] Exemplarily, when a preset audio storage instruction is received, the audio resource indicated by the preset audio storage instruction can be determined as the first audio resource. The preset audio storage instruction can be understood as an instruction for requesting to perform an operation of adding the fingerprint information (hereinafter referred to as fingerprint information) of the indicated audio resource to a preset fingerprint library, that is, an operation of storing the first fingerprint information of the first audio resource in the library is required. For example, when a user uploads a file carrying audio through a client, a preset audio storage instruction indicating the file can be generated. Among them, the preset fingerprint library can be understood as a database for storing audio fingerprint information, and the specific storage method is not limited.
[0025] In the embodiment of the present invention, the specific source of the first audio resource is not limited. It may be an audio resource in a preset audio resource library, an audio resource locally stored in a computer device, or a real-time recorded audio resource, etc. The preset audio resource library can be understood as a library for storing audio resources in a server. In the preset audio resource library, an audio identifier (such as music_id) can be set for the audio resource. For an audio resource that does not exist in the preset audio resource library, the system can automatically generate an audio identifier for it.
[0026] Exemplarily, the method for determining the first fingerprint information is not limited. It can be pre-generated and directly obtained in this step, or generated currently based on the first audio resource. In the embodiments of the present invention, the generation method and the manifestation form of the fingerprint information are not limited. For example, a series of hash values that can represent its acoustic characteristics can be calculated based on the audio resource, and this series of hash values can be used as the fingerprint information.
[0027] Optionally, the fingerprint information can be generated in the following manner: divide the target audio data corresponding to the first audio resource into multiple frames of audio signals; convert the audio signals into spectrograms by means of Discrete Fourier Transform (DFT), short-time Fourier transform (STFT), etc.; traverse the data points representing peaks on the spectrogram as peak points; extract the feature information of the peak points, where the feature information can include the characteristics of the peak points themselves or the characteristics between the peak points, etc.; combine the peak points according to the feature information to obtain multiple peak pairs; encode each peak pair to obtain the fingerprint information. Optionally, for each peak pair, a hash coding method can be used to generate the hash value corresponding to the current peak pair, and this hash value can be used as the sub-fingerprint data corresponding to the current peak pair. The coding basis of a hash value can include the first time when the first peak point in the peak pair appears in the audio data, the first frequency of the first peak point and the second frequency of the second peak point in the peak pair, and the difference between the second time and the first time, where the second time is the time when the second peak point appears in the audio data.
[0028] Step 102: Match the first fingerprint information with each fingerprint information in the preset fingerprint database.
[0029] Exemplarily, the similarity between the first fingerprint information and each fingerprint information in the preset fingerprint database can be calculated respectively, and it is determined whether the two fingerprint information match successfully according to the similarity.
[0030] Exemplarily, the fingerprint information includes multiple sub-fingerprint data. As described above, a peak pair can correspond to a sub-fingerprint data. For audio, it usually has copyright attributes. The audio resources corresponding to the two fingerprint information to be matched may have the same copyright information. For example, they may come from the whole and the chorus part of the same song respectively. In this case, both fingerprint information contains sub-fingerprint data of the chorus part, and the arrangement order of each sub-fingerprint data in the chorus part and the time interval between every two sub-fingerprint data are the same.
[0031] In the embodiments of the present invention, it is possible to determine whether two fingerprint information match successfully based on the matching duration and the number of matching sub-fingerprint data. Exemplarily, the fingerprint information currently required to be matched with the first fingerprint information in the preset fingerprint library is denoted as the fingerprint information to be matched. The consistent sub-fingerprint data in the first fingerprint information and the fingerprint information to be matched are found, and a first sub-fingerprint sequence (corresponding to the first fingerprint information) and a second sub-fingerprint sequence (corresponding to the fingerprint information to be matched) are respectively formed. The number of sub-fingerprint data included in the first sub-fingerprint sequence and the second sub-fingerprint sequence (i.e., the number of matches) is the same, and they are arranged in the order in the corresponding fingerprint information. For the adjacent first sub-fingerprint data and second sub-fingerprint data in the first sub-fingerprint sequence, the first time interval between the two is calculated, and the third sub-fingerprint data consistent with the first sub-fingerprint data and the fourth sub-fingerprint data consistent with the second sub-fingerprint data are found in the second sub-fingerprint sequence, and the second time interval between the two is calculated. If the first time interval and the second time interval are equal, it is counted into the matching duration. Each pair of adjacent sub-fingerprint data in the first sub-fingerprint sequence is traversed, and it is determined whether to count into the matching duration based on the above rules, so as to determine the matching duration corresponding to the first fingerprint information and the fingerprint information to be matched.
[0032] Step 103, if there is a second fingerprint information that matches the first fingerprint information successfully, prohibit the execution of the warehousing operation of the first fingerprint information, and update the matching times of the library resource identifier associated with the second fingerprint information in the preset warehousing resource information library.
[0033] Optionally, if there is a second fingerprint information in the preset fingerprint library whose similarity to the first fingerprint information is greater than a preset threshold, it is determined that the second fingerprint information and the first fingerprint information match successfully.
[0034] Optionally, after obtaining the matching duration of each fingerprint information in the preset fingerprint library and the first fingerprint information, the maximum matching duration can be selected, and it is determined whether the maximum matching duration is greater than or equal to a preset duration threshold, and it is determined whether the corresponding number of matches is greater than or equal to a preset number threshold. If the maximum matching duration is greater than or equal to the preset duration threshold, and the corresponding number of matches is greater than or equal to the preset number threshold, the fingerprint information corresponding to the maximum matching duration is determined as the second fingerprint information that matches the first fingerprint information successfully. The advantage of such a setting is that the second fingerprint information closest to the first fingerprint information in the preset fingerprint library can be found quickly and accurately.
[0035] In an embodiment of the present invention, a preset incoming library resource information database can be pre-constructed to record relevant information of audio resources with fingerprint information stored in a preset fingerprint database. The relevant information may include an in-library resource identifier and the number of matches. The in-library resource identifier can be used to represent the audio category, and the fingerprint information that matches successfully has the same in-library resource identifier, thereby realizing the clustering of audio resources. For the first incoming audio resource, its audio identifier can be set as its corresponding in-library resource identifier, or its corresponding in-library resource identifier can be set in other ways, such as copyright information (such as song name and singer name) or storage location in a preset audio resource database, etc. Then, the fingerprint information of the first audio resource is stored in the preset fingerprint database, and the number of matches is set to 0 or 1 (for ease of description, the number of matches of the in-library resource identifier for the first incoming is uniformly recorded as 1), and the association relationship between the incoming fingerprint information and the in-library resource identifier is recorded.
[0036] Exemplarily, for the first fingerprint information, if there is a second fingerprint information that matches it successfully, it means that the first audio resource and the second audio resource to which the second fingerprint information belongs can be regarded as similar audio resources, for example, from the same song. Therefore, they have the same in-library resource identifier. At this time, the incoming operation of the first fingerprint information is prohibited, that is, the first fingerprint information is prohibited from being stored in the preset fingerprint database, saving the storage space of the preset fingerprint database. At the same time, in the preset incoming library resource information database, the number of matches of the in-library resource identifier associated with the second fingerprint information is incremented, such as adding 1, to update the number of matches. In this way, through the number of matches, the clustering situation of a certain type of audio resource can be intuitively understood, and this clustering situation can reflect the popularity of this type of audio resource, providing auxiliary information for service scenarios such as audio recommendation.
[0037] In the embodiment of the present invention, a processing method based on audio fingerprints is provided. The first fingerprint information of the first audio resource to be stored in the library is determined, and the first fingerprint information is matched with each fingerprint information in the preset fingerprint library. If there is a second fingerprint information that matches the first fingerprint information successfully, the operation of storing the first fingerprint information in the library is prohibited, and the matching times of the in-library resource identifier associated with the second fingerprint information are updated in the preset in-library resource information library. By adopting the above technical solution, when storing audio fingerprint information, it is first determined whether there is fingerprint information in the fingerprint library that matches it. If there is, the storage in the library is prohibited, and the matching times of the associated in-library resource identifier are updated in the resource information library. The audio fingerprints are clustered through the in-library resource identifier in the resource information library, thereby effectively saving the storage space of the fingerprint library. Moreover, for specific business scenarios such as music copyright identification tasks and audio recommendation tasks, effective auxiliary information can be provided, reducing the cost of querying copyright information by third-party services and improving the accuracy of providing audio recommendations. In addition, compared with clustering by calling third-party services, the number of audio resources in third-party services is limited, usually in the form of music, and new audio resources are not supported to be added. However, this solution also has a good clustering effect on non-music forms of audio such as rap and humming, as well as personalized audio recorded by users themselves.
[0038] Exemplarily, the music copyright identification task is mainly to determine the copyright information of the background music used in the audio or video uploaded by the user. Generally, it is necessary to call a third-party service for query and pay fees according to the call volume. After using the processing solution based on audio fingerprints provided in the embodiment of the present invention, clustering of the audio to be queried can be realized to achieve the purpose of deduplication. For audio of the same category, the third-party service needs to be called only once to query the copyright information, which can greatly reduce the call volume and save costs. For example, record the audio identifier sets corresponding to each in-library resource identifier in the preset in-library resource information library. The number of audio identifiers in the audio identifier set generally coincides with the matching times. By calling the third-party service to query the audio resource to which any audio identifier in the audio identifier set belongs once, the effect of batch query can be achieved. Optionally, when the in-library resource identifier is the same as the audio identifier of the first audio resource stored in the same audio category, the corresponding audio resource can also be found according to the in-library resource identifier, and then the copyright query can be performed.
[0039] Exemplarily, in the short video recommendation task, it is necessary to find some of the most used background music (such as selected from the music library or uploaded locally) in the video uploaded by the user. After using the processing solution based on audio fingerprints provided in the embodiment of the present invention, the category number (in-library resource identifier) to which the music belongs can be returned. Through certain statistical analysis of the clustering results, the categories with a larger aggregation quantity (matching times) are recommended to the user as popular music for video recording.
[0040] In some embodiments, after matching the first fingerprint information with each fingerprint information in the preset fingerprint database, the following steps are further included: If there is no second fingerprint information that matches the first fingerprint information successfully, a library resource identifier associated with the first fingerprint information is newly created in the preset in-storage resource information database, the corresponding number of matching times is initialized, and the in-storage operation of the first fingerprint information is executed. The advantage of this setting is that when it is determined that there is no similar audio, a new audio category is created and the number of matching times is initialized, ensuring that more audio types can be covered in the preset fingerprint database.
[0041] Exemplarily, executing the in-storage operation of the first fingerprint information can specifically be to store the first fingerprint information into the preset fingerprint database according to the storage method of the fingerprint information in the preset fingerprint database. In the preset in-storage resource information database, a first library resource identifier can be newly created according to the first audio identifier of the first audio resource. For example, the first library resource identifier is the same as the first audio identifier, and the number of matching times can be initialized to 1.
[0042] In some embodiments, the fingerprint information includes a plurality of sub-fingerprint data; the preset fingerprint database includes a first preset number of first storage positions, and each first storage position corresponds to a second preset number of second storage positions; the first storage position is used to store the sub-fingerprint data; the second storage position is used to store the associated storage information of the sub-fingerprint data on the corresponding first storage position, where the associated storage information includes the library resource identifier associated with the sub-fingerprint data and the position information of the sub-fingerprint data in the corresponding audio resource. The advantage of this setting is that the storage space in the preset fingerprint database can be rationally utilized.
[0043] Exemplarily, the above storage method can be regarded as a key-value storage method. The content stored in the first storage position can be understood as the key, and the content stored in the second storage position can be understood as the value. Assume that the size of the preset fingerprint database is a*b, where a represents the number of keys (for example, 2^n, n is an integer greater than 1), and b represents the number of values corresponding to each key (for example, 1024). The position information of the sub-fingerprint data in the corresponding audio resource can be understood as the appearance time of the sub-fingerprint data in the corresponding audio resource, such as a few minutes and seconds in the entire duration of the audio resource, or can be understood as the serial number of the sub-fingerprint data in the corresponding fingerprint information, such as the serial number of the sub-fingerprint data that appears first. Using the position information, the corresponding sub-fingerprint data can be sorted and spliced to obtain the complete fingerprint information.
[0044] In the related art, the situation of fingerprint storage when the fingerprint database is full is not fully considered. If the fingerprint is not stored, it will lead to the lack of fingerprints of the latest audio in the database, resulting in poor real-time performance. If the method of randomly overwriting the sub-fingerprints in the database is adopted, the fingerprints of the audio in the database will be incomplete and difficult to match again. In the scenario of processing a large amount of audio, the retrieval accuracy will be greatly reduced, and the application scenario is limited.
[0045] In the embodiments of the present invention, for the situation where the fingerprint database is full, optimization is carried out. A certain strategy is adopted to perform the operation of removing the complete fingerprint information from the database, avoiding the incomplete fingerprint information of the audio caused by randomly overwriting the sub-fingerprint data, and improving the utilization rate of each sub-fingerprint data in the fingerprint database.
[0046] In some embodiments, the first fingerprint information includes a plurality of first sub-fingerprint data. The operation of storing the first fingerprint information includes: during the process of sequentially performing the operation of storing the plurality of first sub-fingerprint data, determining whether there is a corresponding idle second storage location for the current first sub-fingerprint data in the preset fingerprint database; if not, screening the resource identifier in the target database by using the preset incoming storage resource information database based on a preset screening strategy, where the target fingerprint information associated with the resource identifier in the target database includes at least one of the current first sub-fingerprint data; performing the operation of removing the target fingerprint information associated with the resource identifier in the target database; and storing the associated storage information corresponding to the current first sub-fingerprint data into the preset fingerprint database. The advantage of this setting is that each sub-fingerprint data of the first audio resource is stored one by one. When it is found that the second storage location corresponding to the current sub-fingerprint data is already occupied, a certain strategy is adopted to determine the resource identifier in the target database, and the associated target fingerprint information is completely removed from the database. While freeing up the storage location for the associated storage information of the first sub-fingerprint data that needs to be stored currently, it avoids the partial residue of the target fingerprint information corresponding to the resource identifier in the target database, and the incomplete fingerprint information is difficult to be used for fingerprint information matching. The residue in the preset fingerprint database will occupy valuable storage space. The embodiments of the present invention can effectively improve the utilization rate of the storage space in the preset fingerprint database.
[0047] In some embodiments, the screening of the resource identifier in the target database by using the preset incoming storage resource information database based on a preset screening strategy includes: determining the resource identifier in the database stored at the second storage location corresponding to the current first sub-fingerprint data as the alternative resource identifier in the database; sorting the alternative resource identifiers in the database by using the preset incoming storage resource information database based on a preset sorting strategy, and screening the resource identifier in the target database according to the sorting result. The advantage of this setting is that the resource identifier in the target database can be quickly and accurately screened out.
[0048] Exemplarily, the library resource identifiers included in the associated storage information stored at the second storage location corresponding to the current first sub-fingerprint data can be queried in sequence, the queried library resource identifiers can be determined as candidate library resource identifiers, and the candidate library resource identifiers can be sorted according to the relevant information of the candidate library resource identifiers stored in the preset incoming library resource information library, and a suitable candidate library resource identifier can be selected as the target library resource identifier.
[0049] In some embodiments, the preset sorting strategy is determined based on at least one of the matching times corresponding to the candidate library resource identifiers, the incoming time of the fingerprint information associated with the candidate library resource identifiers, and the historical playback information of the audio resources corresponding to the candidate library resource identifiers, wherein the incoming time and the historical playback information are stored in the preset incoming library resource information library. The advantage of such a setting is that the candidate library resource identifiers can be reasonably sorted, so as to perform the outbound processing on the corresponding fingerprint information.
[0050] Exemplarily, the higher the matching times, the higher the popularity of the corresponding audio type, and the greater the possibility of the subsequent increase in the matching times. Generally, it needs to be retained. The candidate library resource identifiers can be sorted in ascending order according to the matching times, and the candidate library resource identifier ranked first can be determined as the target library resource identifier.
[0051] Exemplarily, the earlier the incoming time, the older the corresponding audio type. The candidate library resource identifiers can be sorted in ascending order according to the incoming time, and the candidate library resource identifier ranked first can be determined as the target library resource identifier.
[0052] Exemplarily, the historical playback information can include the historical playback volume or the playback volume within a preset historical period, etc. The higher the playback volume, the higher the popularity of the corresponding audio type. The candidate library resource identifiers can be sorted in ascending order according to the playback volume, and the candidate library resource identifier ranked first can be determined as the target library resource identifier.
[0053] Optionally, the above items can be combined for comprehensive sorting. Taking the matching times and the incoming time as an example, first sort in ascending order according to the matching times. If the matching times of multiple candidate library resource identifiers ranked first are the same, then sort in ascending order according to the incoming time, and then determine the final target library resource identifier.
[0054] Optionally, a sorting function can be set based on the above items. The sorting function can be a weighted sum of the values of at least two items or other forms, which is not specifically limited, so as to more accurately obtain the comprehensive scores corresponding to each candidate library resource identifier, and sort according to the comprehensive scores, and determine the one with the lowest comprehensive score as the target library resource identifier.
[0055] In some embodiments, performing an outbound operation on the target fingerprint information associated with the resource identifier in the target library includes: querying the target sub - fingerprint data in the target fingerprint information associated with the resource identifier in the target library from the preset sub - fingerprint records in the preset inbound resource information library, where when creating the resource identifier in the target library, the target sub - fingerprint data is associated with the resource identifier in the target library in the preset sub - fingerprint records; locating the target sub - fingerprint data in the preset fingerprint library, and searching for the resource identifier in the target library at the corresponding second storage location of the target sub - fingerprint data, and deleting the associated storage information to which the resource identifier in the target library belongs from the corresponding second storage location; deleting the resource identifier in the target library from the preset inbound resource information library. The advantage of such a setting is that there are preset sub - fingerprint records in the preset inbound resource information library, which facilitates quickly finding the associated storage information to be deleted when outbound processing of fingerprint information is required, avoiding traversing the entire preset fingerprint library and improving the outbound efficiency. Among them, when deleting the resource identifier in the target library from the preset inbound resource information library, relevant information of the resource identifier in the target library in the preset inbound resource information library, such as the matching times and the inbound time, can be synchronously deleted.
[0056] In some embodiments, before determining the first fingerprint information of the first audio resource to be inbound, it further includes: obtaining the first audio identifier of the first audio resource to be inbound; determining whether there is a resource identifier in the preset inbound resource information library that is the same as the first audio identifier; if so, updating the matching times of the resource identifier in the preset inbound resource information library that is the same as the first audio identifier. The advantage of such a setting is that storing the audio identifier as the resource identifier in the library can quickly query whether the fingerprint information corresponding to the first audio resource has been processed for inbound according to the audio identifier. If so, it is not necessary to determine the first fingerprint information again, and directly update the matching times, improving the audio clustering efficiency.
[0057] In some embodiments, after determining whether there is an in-library resource identifier in the preset incoming library resource information library that is consistent with the audio identifier of the first audio resource, the method further includes: if not, determining whether the first audio identifier exists in a preset matching list, where the preset matching list records the audio identifiers of audio resources that have undergone fingerprint information matching operations and the in-library resource identifiers associated with the corresponding successfully matched fingerprint information; if so, updating the matching times of the corresponding in-library resource identifier in the preset incoming library resource information library in the preset matching list. The advantage of this setting is that if there is no consistent in-library resource identifier, there may be a situation where the in-library resource identifier is named after the audio identifier of other similar audio. At this time, it is possible to determine whether the first audio resource has undergone the fingerprint information matching process by querying the preset matching list. If so, there is no need to determine the first fingerprint information or perform fingerprint matching again, improving the audio processing efficiency.
[0058] In some embodiments, performing the incoming operation of the first fingerprint information includes: determining a target time period in the current incoming cycle at the current moment, where the incoming cycle includes an incoming time period and a prohibited incoming time period; if the target time period is the incoming time period, performing the incoming operation of the first fingerprint information; if the target time period is the prohibited incoming time period, prohibiting the performance of the incoming operation of the first fingerprint information. The advantage of this setting is that a periodic prohibited incoming time period is set, taking into account both the processing efficiency of audio fingerprints and the diversity of fingerprint information in the preset fingerprint library.
[0059] Exemplarily, as the amount of audio processing increases, the preset fingerprint library may become slower and slower. When the fingerprint information of a new audio resource needs to be stored in the library, the number of audio that needs to be removed from the library will also increase, and the processing efficiency will decrease. At this time, an incoming and outgoing library cycle can be set. For example, the time of one cycle is T. In the first t1 time, the normal incoming and outgoing library scheme is adopted, and in the latter T - t1 time, the scheme of only matching without incoming (that is, if the match is successful, the matching times are updated, and if the match is unsuccessful, the incoming operation is prohibited) is adopted. While ensuring a certain deduplication rate, the processing efficiency can be greatly improved.
[0060] Figure 2 FIG. is a flowchart of another processing method based on audio fingerprints provided by an embodiment of the present invention, which is optimized on the basis of the above optional embodiments. As Figure 2 shown, the method may include:
[0061] Step 201, obtaining a first audio identifier of a first audio resource to be stored in the library.
[0062] Exemplarily, the audio identifier can be understood as the unique identity identifier of the audio resource in the preset audio resource library, denoted as music_id.
[0063] Step 202: Determine whether there is a library resource identifier in the preset warehousing resource information library that is the same as the first audio identifier. If so, execute Step 211; otherwise, execute Step 203.
[0064] Exemplarily, the preset warehousing resource information library stores relevant information of the fingerprint information of the audio saved in the preset fingerprint library, including the library resource identifier (which can be denoted as class_id), the warehousing time, the number of matches, and the corresponding sub-fingerprint data (which can be denoted as hash_id). class_id can be understood as a category identifier, and its value can be the audio identifier of the first audio resource warehoused in this category.
[0065] Exemplarily, for the sake of illustration, the following takes the size of the preset fingerprint library as 2^2*3 as an example for schematic illustration. The first storage location is used to store hash_id, the first preset quantity is 4, and the corresponding second storage location existing in the first storage location is used to store class_id (the value is music_id) and location information (denoted as info), and the second preset quantity is 3. Assume that the preset fingerprint library currently stores the fingerprint information of three types of audio resources. Table 1 shows the preset fingerprint library, and Table 2 shows the preset warehousing resource information library.
[0066] Table 1 Preset Fingerprint Library
[0067] hash_id1 music_idA info1 music_idB info4 music_idC info5 hash_id2 music_idB info3 music_idC info6 hash_id3 music_idA info2 hash_id4
[0068] Table 2 Preset Warehousing Resource Information Library
[0069]
[0070]
[0071] As shown in Table 1 and Table 2, 4 sub-fingerprint data are stored in the first storage location (the first column) of the current preset fingerprint database, namely hash_id1, hash_id2, hash_id3, and hash_id4; the class_ids of the three types of audio resources currently stored are music_idA, music_idB, and music_idC respectively; the fingerprint information corresponding to music_idA includes hash_id1 (info1, location information is marked here for easy distinction from the sub-fingerprint data in other fingerprint information) and hash_id3 (info2), the fingerprint information corresponding to music_idB includes hash_id1 (info4) and hash_id2 (info3), and the fingerprint information corresponding to music_idC includes hash_id1 (info5) and hash_id2 (info6); the storage times of music_idA, music_idB, and music_idC are represented by serial numbers, denoted as 1, 2, and 3 respectively; the matching times of music_idA, music_idB, and music_idC are 20, 50, and 20 respectively.
[0072] Assume that the music_id of the first audio resource to be stored in the database currently is D, and music_idD does not exist in the preset information database of resources to be stored, then step 203 can be continued.
[0073] Step 203: Determine whether there is a first audio identifier in the preset matching list. If so, execute step 210; otherwise, execute step 204.
[0074] Exemplarily, assume that music_idD is not included in the preset matching list either, then step 204 can be continued.
[0075] Step 204: Calculate the first fingerprint information of the first audio resource.
[0076] Exemplarily, the fingerprint information contains multiple sub-fingerprint data, and each sub-fingerprint data is a hash value (i.e., the above-mentioned hash_id). The specific calculation method can refer to the relevant description above and will not be elaborated here.
[0077] Assume that the sub-fingerprint data included in the first fingerprint information of music_idD calculated are hash_id1 (info7) and hash_id4 (info8) respectively.
[0078] Step 205: Match the first fingerprint information with each fingerprint information in the preset fingerprint database, and determine the matching duration corresponding to each fingerprint information and the matching quantity of the corresponding sub-fingerprint data.
[0079] Step 206: Determine whether the maximum matching duration is greater than or equal to a preset duration threshold and the corresponding matching count is greater than or equal to a preset count threshold. If so, execute Step 207; otherwise, execute Step 208.
[0080] Step 207: Determine the fingerprint information corresponding to the maximum matching duration as the second fingerprint information that successfully matches the first fingerprint information, prohibit the execution of the warehousing operation for the first fingerprint information, and update the matching count of the in-warehouse resource identifier associated with the second fingerprint information in the preset warehousing resource information database.
[0081] Exemplarily, assume that the matching duration between the first fingerprint information and the fingerprint information of music_idA is the longest, greater than or equal to the preset duration threshold, and the corresponding matching count is greater than or equal to the preset count threshold. Then, there is no need to store hash_id1(info7) and hash_id4(info8), and the matching count corresponding to music_idA is incremented by 1 to become 21.
[0082] Step 208: Create a new in-warehouse resource identifier associated with the first fingerprint information in the preset warehousing resource information database and initialize the corresponding matching count.
[0083] Exemplarily, assume that the first fingerprint information fails to match all fingerprint information in the preset fingerprint database. Then, a new class_id is added to the preset warehousing resource information database, with the value of music_idD, and the matching count is initialized to 1.
[0084] Step 209: Execute the warehousing operation for the first fingerprint information.
[0085] Exemplarily, the warehousing operation is sequentially executed for multiple first sub-fingerprint data in the first fingerprint information. If there is no corresponding idle second storage location for the current first sub-fingerprint data in the preset fingerprint database, the in-warehouse resource identifier stored in the second storage location corresponding to the current first sub-fingerprint data is determined as the alternative in-warehouse resource identifier. The alternative in-warehouse resource identifiers are sorted using the preset warehousing resource information database based on the preset sorting strategy. The target in-warehouse resource identifier is selected according to the sorting result. The target sub-fingerprint data in the target fingerprint information associated with the target in-warehouse resource identifier is queried in the preset sub-fingerprint record in the preset warehousing resource information database. The target sub-fingerprint data is located in the preset fingerprint database, and the target in-warehouse resource identifier is searched for in the second storage location corresponding to the target sub-fingerprint data. The associated storage information belonging to the target in-warehouse resource identifier is deleted from the corresponding second storage location. The associated storage information corresponding to the current first sub-fingerprint data is stored in the preset fingerprint database, and the target in-warehouse resource identifier is deleted from the preset warehousing resource information database.
[0086] Specifically, as in the above example, the current first sub-fingerprint data is hash_id1(info7), that is, an operation of storing it in the database needs to be performed on hash_id1(info7). It is found that the 3 second storage positions corresponding to hash_id1 are all occupied. Then, the in-library resource identifiers music_idA, music_idB, and music_idC stored in these 3 second storage positions are determined as alternative in-library resource identifiers. First, they are sorted in ascending order according to the corresponding matching times. It is found that the matching times of music_idA and music_idC are tied for the smallest. Then, the storage times are compared, and it is found that the storage time of music_idA is earlier than that of music_idC. Therefore, music_idA is determined as the target in-library resource identifier. Query the target sub-fingerprint data in the target fingerprint information associated with music_idA in the preset in-library resource information database (the last column in Table 2), and obtain hash_id1 and hash_id3. Find hash_id1 and hash_id3 in the preset fingerprint database, and search for music_idA in the storage positions corresponding to hash_id1 and hash_id3. Delete the associated storage information music_idA info1 and music_idAinfo2 containing music_idA. In this way, the corresponding second storage position is vacated, and music_idD info7 can be stored. Subsequently, when the current first sub-fingerprint data is hash_id4(info8), music_idD info8 can be directly stored in the second storage position corresponding to hash_id4. In addition, delete the data stored in the row where music_idA is located in the preset in-library resource information database, including music_idA, storage time 1, matching times 20, and the corresponding hash_id1 and hash_id3.
[0087] Exemplarily, Table 3 shows the processed preset fingerprint database, and Table 4 shows the processed preset in-library resource information database.
[0088] Table 3 Processed Preset Fingerprint Database
[0089]
[0090]
[0091] Table 4 Processed Preset In-library Resource Information Database
[0092] music_id Storage time (sequence number) Matching times hash_id Music_idB 2 50 hash_id1 hash_id2 Music_idC 3 20 hash_id1 hash_id2 Music_idD 4 1 hash_id1 hash_id4
[0093] Step 210: Update the matching times of the corresponding in-library resource identifiers in the preset matching list in the preset in-library resource information database.
[0094] Step 211: Update the matching times of the in-library resource identifier that is the same as the first audio identifier in the preset in-library resource information database.
[0095] In the processing method based on audio fingerprint provided in the embodiment of the present invention, first, it is determined whether the audio resource to be stored in the library has been stored in the library or has been matched according to the audio identifier. If so, the calculation time and matching time of the fingerprint information can be saved, and the processing efficiency can be improved. After determining that it needs to be stored in the library, the fingerprint information is calculated and matched with each fingerprint information in the preset fingerprint library. If the match is successful, there is no need to store it in the library, and the matching times are updated to achieve audio clustering and save the storage space of the fingerprint library. If the match fails and the expected storage location is already occupied, the in-library resource identifier of the type of audio resource to be taken out of the library is quickly and reasonably determined according to the matching times and the storage duration, etc., and the corresponding fingerprint information is quickly and completely cleared to ensure the integrity of the fingerprints in the library, and the fingerprints of the audio with relatively high popularity are retained, alleviating the problem of the decrease in the recognition accuracy caused by the large amount of random coverage and replacement of fingerprint data due to the limited size of the fingerprint library and the increase in the number of audio identifications. The withdrawal scheme can effectively ensure the integrity of the fingerprints in the library and the retention and matching performance of popular audio will not decrease with the increase in the number of audio. Moreover, for specific business scenarios such as music copyright identification tasks and audio recommendation tasks, effective auxiliary information can be provided, reducing the cost of querying copyright information by third-party services and improving the accuracy of providing audio recommendations, etc.
[0096] Figure 3 It is a structural block diagram of a processing device based on audio fingerprint provided in the embodiment of the present invention. This device can be implemented by software and / or hardware, and is generally integrated in a computer device, and can perform audio fingerprint processing by executing the processing method based on audio fingerprint. As Figure 3 shown, this device includes:
[0097] A fingerprint information determination module 301, configured to determine the first fingerprint information of the first audio resource to be stored in the library;
[0098] A fingerprint information matching module 302, configured to match the first fingerprint information with each fingerprint information in the preset fingerprint library;
[0099] An in-library control module 303, configured to, if there is a second fingerprint information that matches the first fingerprint information successfully, prohibit the execution of the in-library operation of the first fingerprint information, and update the matching times of the in-library resource identifier associated with the second fingerprint information in the preset in-library resource information database.
[0100] In the embodiment of the present invention, the processing device based on audio fingerprint, when it is necessary to store audio fingerprint information in the library, first determines whether there is matching fingerprint information in the fingerprint library. If there is, it prohibits storage and updates the matching times of the associated in-library resource identifier in the resource information library, and clusters the audio fingerprints through the in-library resource identifier in the resource information library, thus effectively saving the storage space of the fingerprint library. Moreover, for specific business scenarios such as music copyright identification tasks and audio recommendation tasks, it can provide effective auxiliary information, reduce the cost of querying copyright information by third-party services, and improve the accuracy of providing audio recommendations, etc.
[0101] The embodiment of the present invention provides a computer device, and the processing device based on audio fingerprint provided by the embodiment of the present invention can be integrated in this computer device. Figure 4 It is a structural block diagram of a computer device provided by the embodiment of the present invention. The computer device 400 includes a memory 401, a processor 402, and a computer program stored on the memory 401 and executable on the processor 402. When the processor 402 executes the computer program, it implements the processing method based on audio fingerprint provided by the embodiment of the present invention.
[0102] The embodiment of the present invention also provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute the processing method based on audio fingerprint provided by the embodiment of the present invention when executed by a computer processor.
[0103] The processing device, device, and storage medium based on audio fingerprint provided in the above embodiments can execute the processing method based on audio fingerprint provided by any embodiment of the present invention, and have the corresponding functional modules and beneficial effects for executing this method. For the technical details not described in detail in the above embodiments, reference can be made to the processing method based on audio fingerprint provided by any embodiment of the present invention.
[0104] Note that the above is only a preferred embodiment of the present invention. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described here, and various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, it can also include more other equivalent embodiments, and the scope of the present invention is determined by the scope of the claims.
Claims
1. A processing method based on audio fingerprints, characterized in that, Including: Determine the first fingerprint information of the first audio resource to be stored in the warehouse; Match the first fingerprint information with each fingerprint information in the preset fingerprint database; If there is a second fingerprint information that matches the first fingerprint information successfully, prohibit the execution of the warehousing operation of the first fingerprint information, and update the matching times of the in-warehouse resource identifiers associated with the second fingerprint information in the preset in-warehouse resource information database; Wherein, the fingerprint information includes a plurality of sub-fingerprint data; the preset fingerprint database includes a first preset number of first storage locations, and each first storage location corresponds to a second preset number of second storage locations; the first storage location is used to store sub-fingerprint data; the second storage location is used to store the associated storage information of the sub-fingerprint data on the corresponding first storage location; Wherein, the first fingerprint information includes a plurality of first sub-fingerprint data, and the execution of the warehousing operation of the first fingerprint information includes: During the process of sequentially executing the warehousing operation for the plurality of first sub-fingerprint data, determine whether there is a corresponding idle second storage location for the current first sub-fingerprint data in the preset fingerprint database; If not, screen the target in-warehouse resource identifier by using the preset in-warehouse resource information database based on a preset screening strategy, wherein the target fingerprint information associated with the target in-warehouse resource identifier includes the current first sub-fingerprint data; Execute the out-of-warehouse operation of the target fingerprint information; Store the associated storage information corresponding to the current first sub-fingerprint data into the preset fingerprint database.
2. The method according to claim 1, characterized in that, After the first fingerprint information is matched with each fingerprint information in the preset fingerprint database, it further includes: If there is no second fingerprint information that matches the first fingerprint information successfully, create a new in-warehouse resource identifier associated with the first fingerprint information in the preset in-warehouse resource information database, initialize the corresponding matching times, and execute the warehousing operation of the first fingerprint information.
3. The method according to claim 1, characterized in that, The fingerprint information includes a plurality of sub-fingerprint data, and the matching of the first fingerprint information with each fingerprint information in the preset fingerprint database includes: Match the first fingerprint information with each fingerprint information in the preset fingerprint database, and determine the matching duration corresponding to each fingerprint information and the matching quantity of the corresponding sub-fingerprint data; The existence of a second fingerprint information that matches the first fingerprint information successfully includes: If the maximum matching duration is greater than or equal to a preset duration threshold, and the corresponding matching quantity is greater than or equal to a preset quantity threshold, determine the fingerprint information corresponding to the maximum matching duration as the second fingerprint information that matches the first fingerprint information successfully; Wherein, the matching duration is determined by the following method: Record the fingerprint information that currently needs to be matched with the first fingerprint information in the preset fingerprint database as the fingerprint information to be matched; Find the consistent sub-fingerprint data in the first fingerprint information and the fingerprint information to be matched, and respectively form a first sub-fingerprint sequence corresponding to the first fingerprint information and a second sub-fingerprint sequence corresponding to the fingerprint information to be matched, wherein the number of sub-fingerprint data included in the first sub-fingerprint sequence and the second sub-fingerprint sequence is the same, and they are arranged in the order in the corresponding fingerprint information; Traverse every two adjacent sub - fingerprint data in the first sub - fingerprint sequence, and determine whether to include the matching duration based on the following rules to determine the matching duration corresponding to the first fingerprint information and the fingerprint information to be matched: For the adjacent first sub - fingerprint data and second sub - fingerprint data in the first sub - fingerprint sequence, calculate the first time interval between the two; In the second sub - fingerprint sequence, find the third sub - fingerprint data that is the same as the first sub - fingerprint data, and find the fourth sub - fingerprint data that is the same as the second sub - fingerprint data, and calculate the second time interval between the two; If the first time interval is equal to the second time interval, then include the first time interval or the second time interval in the matching duration.
4. The method according to claim 1, characterized in that, The associated storage information includes the in - library resource identifier associated with the sub - fingerprint data, and the position information corresponding to the sub - fingerprint data in the audio resource to which it belongs.
5. The method according to claim 4, characterized in that, Performing the out - of - library operation on the target fingerprint information includes: Query the target sub - fingerprint data in the target fingerprint information in the preset sub - fingerprint records in the preset in - library resource information library. When creating the new target in - library resource identifier, associate the target sub - fingerprint data with the target in - library resource identifier in the preset sub - fingerprint records; Locate the target sub - fingerprint data in the preset fingerprint library, and find the target in - library resource identifier in the second storage location corresponding to the target sub - fingerprint data, and delete the associated storage information to which the target in - library resource identifier belongs from the corresponding second storage location; Delete the target in - library resource identifier from the preset in - library resource information library.
6. The method according to claim 1, characterized in that, Before determining the first fingerprint information of the first audio resource to be stored in the library, it further includes: Obtain the first audio identifier of the first audio resource to be stored in the library; Judge whether there is an in - library resource identifier in the preset in - library resource information library that is the same as the first audio identifier; If it exists, update the matching times of the in - library resource identifier that is the same as the first audio identifier in the preset in - library resource information library.
7. The method according to claim 6, characterized in that, After judging whether there is an in - library resource identifier in the preset in - library resource information library that is the same as the audio identifier of the first audio resource, it further includes: If it does not exist, judge whether the first audio identifier exists in the preset matching list, where the preset matching list records the audio identifiers of the audio resources that have undergone fingerprint information matching operations and the in - library resource identifiers associated with the corresponding successfully matched fingerprint information; If so, update the matching times of the corresponding in - library resource identifier in the preset in - library resource information library in the preset matching list.
8. The method according to claim 2, characterized in that, Performing the in - library operation on the first fingerprint information includes: Determine the target time period in the current storage - in cycle at the current moment, where the storage - in cycle includes a storage - in time period and a prohibited storage - in time period; If the target time period is the storage - in time period, perform the in - library operation on the first fingerprint information; if the target time period is the prohibited storage - in time period, prohibit performing the in - library operation on the first fingerprint information.
9. A processing device based on audio fingerprint, characterized in that, It includes: A fingerprint information determination module for determining the first fingerprint information of the first audio resource to be stored in the library; A fingerprint information matching module, configured to match the first fingerprint information with each fingerprint information in a preset fingerprint database; An inbound storage control module, configured to, if there is second fingerprint information that successfully matches the first fingerprint information, prohibit the execution of the inbound storage operation of the first fingerprint information, and update the matching times of the library resource identifier associated with the second fingerprint information in a preset inbound storage resource information database; Wherein, the fingerprint information includes a plurality of sub-fingerprint data; the preset fingerprint database includes a first preset number of first storage locations, and each first storage location corresponds to a second preset number of second storage locations; the first storage locations are used to store sub-fingerprint data; the second storage locations are used to store the associated storage information of the sub-fingerprint data on the corresponding first storage location; Wherein, the first fingerprint information includes a plurality of first sub-fingerprint data, and the execution of the inbound storage operation of the first fingerprint information includes: During the process of sequentially executing the inbound storage operation for the plurality of first sub-fingerprint data, determining whether there is a corresponding idle second storage location for the current first sub-fingerprint data in the preset fingerprint database; If not, screening a target library resource identifier by using the preset inbound storage resource information database based on a preset screening strategy, wherein the target fingerprint information associated with the target library resource identifier includes the current first sub-fingerprint data; Executing the outbound storage operation of the target fingerprint information; Storing the associated storage information corresponding to the current first sub-fingerprint data into the preset fingerprint database.
10. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the method according to any one of claims 1-8 when executing the computer program.
11. A computer-readable storage medium, on which a computer program is stored, characterized in that, The program implements the method according to any one of claims 1-8 when executed by the processor.
Citation Information
Patent Citations
Dynamic facial image warehousing method and apparatus, electronic device, medium, and program
CN110799972A
Audio fingerprint processing method and device, computer equipment and storage medium
CN112784100A