Audio Retrieval Method, Apparatus and Electronic Device

By extracting audio clips and determining their spectrum fingerprint characteristics, detecting and updating the matching degree count value, the accuracy and timeliness of audio retrieval in streaming audio and video scenes are solved, and the robustness and processing speed of audio retrieval is improved.

CN114398513BActive Publication Date: 2025-05-27GUANGZHOU HUYA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210041171.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-14
Publication Date
2025-05-27
Estimated Expiration
2042-01-14

AI Technical Summary

Technical Problem

The existing audio fingerprint retrieval algorithm is difficult to adapt effectively in streaming audio and video scenarios, especially when facing frame loss or out of order caused by transmission delay, noise interference and encoding losses, the accuracy and timeliness of the retrieval cannot be guaranteed.

Method used

By extracting the audio clip from the target audio and determining its spectral fingerprint characteristics, matching candidate audio clips are detected from the known audio, and position offset is calculated based on the timing position, the matching degree count value is updated to determine the starting position of the target audio in the known audio.

Benefits of technology

It improves the robustness of audio retrieval, and can maintain the accuracy of the matching degree count value and the simplicity of calculation processing when individual audio clips in the target audio are lost or out of order, and adapt to streaming audio retrieval scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114398513B_ABST
    Figure CN114398513B_ABST
Patent Text Reader

Abstract

The present application provides an audio retrieval method, apparatus and electronic device, which can relatively independently extract the fingerprint features of each first audio segment in the target audio, and relatively independently update the matching degree count value of the second audio segment of the known audio based on the fingerprint features, and finally determine the known audio that matches the target audio according to the matching degree count value. In this way, when the number of valid audio segments is sufficient, even if individual audio segments in the target audio are lost or out of order, it will not have a great impact on the statistical result of the matching degree count value, improving the robustness of audio retrieval. Moreover, the calculation processing logic is simple, the retrieval processing speed is fast, and it can better adapt to the retrieval scenario of streaming audio.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of audio - video data processing. Specifically, it relates to an audio retrieval method, apparatus, and electronic device. Background Art

[0002] Audio fingerprint retrieval is a technology for retrieving target audio from a known audio library. Usually, the retrieval result includes the known audio retrieved based on the target audio and the position of the target audio in the known audio. Some existing audio fingerprint retrieval algorithms are usually designed for offline scenarios. To ensure retrieval accuracy, they have high requirements for the integrity and temporal accuracy of audio data frames. However, in streaming audio - video scenarios, data may experience problems such as frame loss or out - of - order due to transmission delay, noise interference, coding loss, etc. Moreover, for some streaming audio - video scenarios (such as live broadcast scenarios), there are also high requirements for the timeliness of audio data retrieval. Therefore, the existing audio fingerprint retrieval methods cannot well adapt to the streaming audio retrieval scenario. Summary of the Invention

[0003] To overcome the above - mentioned deficiencies in the prior art, the purpose of this application is to provide an audio retrieval method, and the method includes:

[0004] Extract a first audio segment from the target audio, and determine the fingerprint feature of the first audio segment according to the spectrum of the first audio segment;

[0005] Detect candidate audio segments with the same fingerprint feature as the first audio segment from the second audio segments of each known audio;

[0006] Determine a position offset according to the temporal position of the first audio segment in the target audio;

[0007] For each of the candidate audio segments, determine a candidate start position in the known audio corresponding to the candidate audio segment according to the position offset, and increase the matching degree count value of the second audio segment corresponding to the candidate start position;

[0008] Based on the matching degree count values of the second audio segments of each known audio obtained from multiple first audio segments, determine the known audio that matches the target audio retrieval, and determine the audio start position corresponding to the target audio in the known audio.

[0009] In a possible implementation manner, the step of determining the fingerprint feature of the first audio segment according to the spectrum of the first audio segment includes:

[0010] Perform a Fourier transform on the first audio segment to obtain spectral data of a preset dimension of the first audio segment; the spectral data includes amplitude values corresponding to multiple frequency identification values;

[0011] Divide the spectral data into multiple spectral sub-bands by partitioning intervals according to a set frequency identification value;

[0012] For each of the spectral sub-bands, use the frequency identification value with the largest amplitude value in the spectral sub-band as the eigenvalue of the spectral sub-band;

[0013] Use the set of eigenvalues of the multiple spectral sub-bands corresponding to the first audio segment as the fingerprint feature of the first audio segment.

[0014] In a possible implementation manner, the step of determining a position offset according to the temporal position of the first audio segment in the target audio includes:

[0015] Determine the offset length between the first audio segment and the start position of the target audio in terms of time sequence as the position offset;

[0016] The step of, for each of the candidate audio segments, determining a candidate start position in the known audio corresponding to the candidate audio segment according to the position offset includes:

[0017] For each of the candidate audio segments, determine the position at the offset length before the candidate audio segment in terms of time sequence in the known audio where the candidate audio segment is located as the candidate start position.

[0018] In a possible implementation manner, the step of determining a known audio that is retrieved and matched with the target audio based on the match count values of the second audio segments of the respective known audios obtained from multiple first audio segments, and determining the audio start position corresponding to the target audio in the known audio includes:

[0019] Detect whether the distribution of the match count values of the second audio segments of the respective known audios is a normal distribution;

[0020] Determine the known audio with the match count value distribution being a normal distribution as the known audio retrieved and matched with the target audio, and determine the corresponding position of the second audio segment with the highest match count value as the audio start position.

[0021] In a possible implementation manner, the step of determining a known audio that is retrieved and matched with the target audio based on the match count values of the second audio segments of the respective known audios obtained from multiple first audio segments, and determining the start position corresponding to the target audio in the known audio includes:

[0022] Check whether the matching degree count value of the second audio segment of each known audio reaches a preset threshold;

[0023] Determine the known audio corresponding to the second audio segment whose matching degree count value first reaches the preset threshold as the known audio retrieved for the target audio, and determine the corresponding position of this second audio segment as the audio start position.

[0024] In a possible implementation manner, the step of extracting the first audio segment from the target audio includes:

[0025] Perform sliding window data extraction on the target audio according to a set window length and a set step length to obtain a plurality of the first audio segments; wherein, the set step length is less than the set window length;

[0026] The method further includes:

[0027] For each of the known audios, perform sliding window data extraction on the known audio according to the set window length and the set step length to obtain a plurality of the second audio segments.

[0028] In a possible implementation manner, the sampling frequency of each of the known audios is a preset sampling frequency; the method further includes:

[0029] Perform downsampling on the target audio to convert the sampling frequency of the target audio to the preset sampling frequency.

[0030] Another object of the present application is to provide an audio retrieval device, and the device includes:

[0031] A data acquisition module, configured to extract a first audio segment from a target audio and determine a fingerprint feature of the first audio segment according to the spectrum of the first audio segment;

[0032] A feature extraction module, configured to detect a candidate audio segment having the same fingerprint feature as the first audio segment from the second audio segments of each known audio;

[0033] An offset calculation module, configured to determine a position offset according to the timing position of the first audio segment in the target audio;

[0034] A matching degree calculation module, configured to determine a candidate start position in the known audio corresponding to the segment and increase the matching degree count value of the second audio segment corresponding to the candidate start position;

[0035] A result determination module, configured to determine a known audio that matches the target audio retrieval based on the match degree count values of the second audio segments of each known audio obtained from multiple said first audio segments, and determine the audio start position corresponding to the target audio in the known audio.

[0036] Another object of the present application is to provide an electronic device, including a processor and a machine-readable storage medium, where the machine-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are executed by the processor, the audio retrieval method provided by the present application is implemented.

[0037] Another object of the present application is to provide a machine-readable storage medium, where the machine-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are executed by one or more processors, the audio retrieval method provided by the present application is implemented.

[0038] Compared with the prior art, an audio retrieval method, device and electronic device provided by the present application can relatively independently extract the fingerprint features of each first audio segment in the target audio, and relatively independently update the match degree count values of the second audio segments of the known audio based on the fingerprint features, and finally determine the known audio that matches the target audio according to the match degree count values. In this way, when the number of effective audio segments is sufficient, even if individual audio segments in the target audio are lost or out of order, it will not have a great impact on the statistical result of the match degree count value, improving the robustness of audio retrieval, and the calculation processing logic is simple and the retrieval processing speed is fast, which can better adapt to the streaming audio retrieval scenario. Description of the Drawings

[0039] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, so they should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0040] Figure 1 It is a schematic flowchart of the steps of the audio retrieval method provided by the embodiment of the present application;

[0041] Figure 2 It is a schematic diagram of the update process of the match degree count value provided by the embodiment of the present application;

[0042] Figure 3 It is a schematic flowchart of the sub-steps of step S120;

[0043] Figure 4 It is a schematic diagram of the electronic device provided by the embodiment of the present application;

[0044] Figure 5 It is a schematic diagram of the functional modules of the audio retrieval device provided by the embodiment of the present application. Specific embodiments

[0045] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. Usually, the components of the embodiments of the present application described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.

[0046] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application claimed, but merely represents selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the scope of protection of the present application.

[0047] It should be noted that: Similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0048] In the description of the present application, it should be noted that the terms "first", "second", "third", etc. are only used for descriptive distinction and cannot be construed as indicating or implying relative importance.

[0049] In the description of the present application, it should also be noted that unless otherwise clearly defined and limited, the terms "set", "installed", "connected", and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific situations.

[0050] Please refer to Figure 1 , Figure 1 It is a flowchart of an audio retrieval method provided for this embodiment. The following will elaborate on each step included in the method in detail.

[0051] Step S110: Extract a first audio segment from the target audio, and determine the fingerprint feature of the first audio segment according to the spectrum of the first audio segment.

[0052] In this embodiment, a first audio segment can be extracted from the target audio according to a preset window length (duration). For example, a first audio segment is extracted from the target audio with a window length of 800 ms. The fingerprint feature can be the spectral feature of the first audio segment.

[0053] Among them, for non-streaming scenarios, the stored target audio can be split or extracted to obtain multiple first audio segments, and then the subsequent steps S120 to S140 are executed for each of the first audio segments. For streaming scenarios, during the continuous acquisition of the target audio, the newly obtained first audio segments can also be continuously and sequentially extracted, and the subsequent steps S120 to S140 are executed.

[0054] Step S120: Detect candidate audio segments with the same fingerprint feature as the first audio segment from the second audio segments of each known audio.

[0055] In this embodiment, a known audio library can be pre-configured. The known audio library stores known audio and the fingerprint features of the second audio segments extracted from each of the known audio. Among them, the known audio and the target audio can adopt the same audio segment division method and fingerprint feature extraction method.

[0056] In step S120, the fingerprint features of the first audio segment can be compared with each second audio segment of each known audio. If a certain second audio segment has the same fingerprint feature as the first audio segment, it means that the content of the second audio segment is the same as or very similar to the content of the first audio segment, and this second audio segment can be marked as the candidate audio segment of the first audio segment.

[0057] It can be understood that in this embodiment, the same first audio segment may have candidate audio segments in multiple different known audio, or the same first audio segment may also have multiple candidate audio segments in the same known audio.

[0058] Step S130: Determine a position offset according to the temporal position of the first audio segment in the target audio.

[0059] In this embodiment, the position offset can represent the position of the first audio segment in the target audio. In a possible implementation manner, the number of audio segments between the current first audio segment in time sequence and the start position of the target audio can be used as the position offset.

[0060] Step S140: For each of the candidate audio segments, determine a candidate start position in the known audio corresponding to the candidate audio segment according to the position offset, and increase the match count value of the second audio segment corresponding to the candidate start position.

[0061] In this embodiment, the candidate start position may represent the possible start position of the target audio in the known audio. For each of the candidate audio segments, the position in the known audio where the candidate audio segment is located and that is at a distance of the position offset from the candidate audio segment in the time sequence and before the candidate audio segment may be determined as the candidate start position.

[0062] For example, please refer to Figure 2 , when extracting audio segments from the target audio and the known audio, the extracted audio segments can be identified by numbers that increase sequentially from 0 according to the time sequence.

[0063] Suppose when performing fingerprint feature matching on the nth first audio segment of the target audio, the determined candidate audio segments include the ith second audio segment of the known audio A and the jth second audio segment of the known audio B.

[0064] Since the ith second audio segment in the known audio A matches the first audio segment, according to the position offset, the (i - n)th second audio segment in the known audio A may correspond to the candidate start position A of the target audio, so the match count value of the (i - n)th second audio segment in the known audio A is incremented by 1.

[0065] Since the jth second audio segment in the known audio B matches the first audio segment, according to the position offset, the (j - n)th second audio segment in the known audio B may correspond to the candidate start position B of the target audio, so the match count value of the (j - n)th second audio segment in the known audio B is incremented by 1.

[0066] Similarly, suppose when performing fingerprint feature matching on the (n + 1)th first audio segment of the target audio, the determined candidate audio segments include the (i + 1)th second audio segment of the known audio A and the kth second audio segment of the known audio C. Then the match count value of the (i - n)th second audio segment in the known audio A is incremented by 1 again, and the match count value of the (k - n - 1)th second audio segment in the known audio C is incremented by 1.

[0067] Step S150: Based on the match count values of the second audio segments of each known audio obtained from multiple first audio segments, determine the known audio that retrieves and matches the target audio, and determine the audio start position corresponding to the target audio in the known audio.

[0068] It can be understood that in this embodiment, for each first audio segment, the actions of performing fingerprint feature extraction, matching, and increasing the match count value are relatively independent. Therefore, even if there are a few dropped frames or out-of-order situations of the first audio segments in the target audio, it will not affect the actions of increasing the match count value performed on other first audio segments. For example, referring to the above example, when a large number of first audio segments in the target audio are all matched with the second audio segments in the known audio A, the match count value of the i-nth second audio segment in the known audio A will be significantly higher than that of other second audio segments.

[0069] Therefore, based on the match count values of the second audio segments of each known audio obtained from multiple said first audio segments, the known audio where the second audio segment with a match count value significantly greater than that of other audio segments is located can be determined as the matching audio that matches the target audio, and the position where the second audio segment with a match count value significantly greater than that of other audio segments is located can be determined as the corresponding audio start position of the target audio in this matching audio.

[0070] Based on the above design, in the audio retrieval method provided by this application, by relatively independently performing fingerprint feature extraction, matching, and increasing the match count value of the first audio segment, and finally determining the known audio that matches the target audio according to the match count value, the influence of dropped frames or out-of-order problems in the target audio on the matching result can be reduced, and the robustness of audio retrieval is improved. And the update calculation processing logic of the match count value is relatively simple, the match count value can be gradually and synchronously executed during the gradual acquisition process of the streaming audio data, the matching processing speed is fast, and it can better adapt to the streaming audio retrieval scenario.

[0071] In a possible implementation manner, as Figure 3 shown, step S120 may include the sub-steps described in S121 - S124 below.

[0072] Step S121, perform a Fourier transform on the first audio segment to obtain the spectral data of the first audio segment in a preset dimension, where the spectral data includes amplitude values corresponding to multiple frequency identification values.

[0073] In this embodiment, the first audio segment can be converted into 1024-dimensional spectral data through the Fast Fourier Transform (FFT). This spectral data includes frequency identification values in the range from 1 to 1024 and the amplitude value corresponding to each frequency.

[0074] Step S122, divide the interval according to the set frequency identification value, and divide the spectral data into multiple spectral sub-bands.

[0075] For example, in this embodiment, the 1024-dimensional spectrum data can be divided into 6 spectrum sub-bands, and the corresponding frequency identification values are: 1 - 100, 101 - 200, 201 - 300, 301 - 400, 401 - 500, 501 - 1024.

[0076] Step S123, for each of the spectrum sub-bands, use the frequency identification value with the largest amplitude value in the spectrum sub-band as the eigenvalue of the spectrum sub-band.

[0077] For example, in the spectrum sub-band with frequency identification values of 1 - 100, if the amplitude value corresponding to the frequency identification value 5 is the largest, then 5 is used as the eigenvalue of this spectrum sub-band.

[0078] Step S124, use the set of eigenvalues of the multiple spectrum sub-bands corresponding to the first audio segment as the fingerprint feature of the first audio segment.

[0079] According to the above division method of 6 spectrum sub-bands, the fingerprint feature of the first audio segment can be a 6-dimensional vector. For example, if the eigenvalues of the 6 spectrum sub-bands corresponding to the first audio segment are 5, 123, 201, 370, 423, and 599 respectively, then the fingerprint feature of the first audio segment is [5, 123, 201, 370, 423, 599].

[0080] Based on the above design, by dividing the spectrum sub-bands and using the frequency identification value with the largest amplitude value as the fingerprint feature, it is possible to simplify the data structure while being able to represent the spectrum feature of the first audio segment, thereby effectively reducing the computational complexity of subsequent comparison operations and improving the efficiency of comparison and retrieval.

[0081] Through the inventor's research, it is found that even if there are a small number of missing frames or out-of-order of the first audio segment in the target audio, in the case of valid first audio segments, among the known audios that match the target audio, the matching degree count values of each second audio segment usually show a relatively obvious normal distribution, and the matching degree count value at the middle position of the normal distribution is significantly higher. While among the known audios that do not match the target audio, the distribution of the matching degree count values of each second audio segment is usually relatively discrete, and the overall matching degree count values are relatively low.

[0082] Therefore, in a possible implementation, in step S140, it is possible to detect whether the distribution of the matching degree count values of the second audio segments of each known audio is a normal distribution, and determine the known audio with the normal distribution of the matching degree count values as the known audio retrieved for the target audio, and determine the corresponding position of the second audio segment with the highest matching degree count value as the audio start position.

[0083] In another possible implementation, considering the timeliness of streaming audio processing, in step S140, it is possible to detect whether the match count value of the second audio segment of each known audio reaches a preset threshold, and determine the known audio corresponding to the second audio segment whose match count value first reaches the preset threshold as the known audio retrieved for the target audio, and determine the corresponding position of this second audio segment as the audio start position.

[0084] For example, as the streaming target audio data is gradually acquired and processed, the match count values of each of the second audio segments are continuously updated, and it is detected whether there is a match count value greater than the preset threshold. When a match count value greater than the preset threshold is detected, the known audio where the corresponding second audio segment is located is determined as the known audio retrieved for the target audio, and the corresponding position of this second audio segment is determined as the audio start position.

[0085] In a possible implementation, in order to ensure the consistency of the fingerprint feature extraction actions for the target audio data and the known audio data, for the known audio data, after downsampling the known audio data to convert the sampling frequency of the known audio to a preset sampling frequency, the actions of audio segment extraction and fingerprint feature extraction can be performed on the known audio data. For the target audio data, before step S110, the target audio can be downsampled first to convert the sampling frequency of the target audio to the preset sampling frequency, and then the subsequent audio segment extraction and fingerprint feature extraction actions can be performed. For example, usually the sampling rate of most music is 44100Hz. In this embodiment, the target audio and the known audio can be uniformly downsampled to 8000Hz and then the subsequent processing actions can be performed.

[0086] In a possible implementation, in step S110, the target audio can be extracted with sliding window data according to a set window length and a set step length to obtain a plurality of the first audio segments. Wherein, the set step length is less than the set window length. In this way, there is partial overlap between each of the first audio segments, so that when matching according to the fingerprint features of each of the first audio segments, the robustness of the overall recognition result of the target audio can be improved.

[0087] Correspondingly, for each of the known audios, sliding window data extraction is also performed on the known audio according to the set window length and the set step length to obtain a plurality of the second audio segments. Thus, the consistency of the audio segment division method and the fingerprint feature extraction result between the known audio data and the target audio data is ensured.

[0088] Based on the same inventive concept, this embodiment also provides an electronic device for executing the above audio retrieval method. Please refer toFigure 4 , Figure 4 1 is a block diagram of an electronic device 100 provided in this embodiment. The electronic device 100 includes an audio search device 110 , a machine-readable storage medium 120 , and a processor 130 .

[0089] The components of the machine-readable storage medium 120 and the processor 130 are electrically connected to each other directly or indirectly to realize data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines. The audio retrieval device 110 includes at least one software function module that can be stored in the machine-readable storage medium 120 in the form of software or firmware or solidified in the operating system (OS) of the electronic device 100. The processor 130 is used to execute the executable modules stored in the machine-readable storage medium 120, such as the software function modules and computer programs included in the audio retrieval device 110.

[0090] The machine-readable storage medium 120 may be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable read-only memory (EEPROM), etc. The machine-readable storage medium 120 is used to store a program, and the processor 130 executes the program after receiving an execution instruction.

[0091] The processor 130 may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application may be implemented or executed. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0092] Please refer toFigure 5 In this embodiment, an audio retrieval device 110 is further provided. The audio retrieval device 110 includes at least one functional module that can be stored in a machine-readable storage medium 120 in software form. Functionally, the audio retrieval device 110 may include a data acquisition module 111, a feature extraction module 112, an offset calculation module 113, a matching degree calculation module 114, and a result determination module 115.

[0093] The data acquisition module 111 is configured to extract a first audio segment from a target audio and determine a fingerprint feature of the first audio segment according to the spectrum of the first audio segment.

[0094] In this embodiment, the data acquisition module 111 can be used to execute Figure 1 the step S110 shown. For a specific description of the data acquisition module 111, reference can be made to the description of the step S110.

[0095] The feature extraction module 112 is configured to detect a candidate audio segment with the same fingerprint feature as the first audio segment from second audio segments of each known audio.

[0096] In this embodiment, the feature extraction module 112 can be used to execute Figure 1 the step S120 shown. For a specific description of the feature extraction module 112, reference can be made to the description of the step S120.

[0097] The offset calculation module 113 is configured to determine a position offset according to the timing position of the first audio segment in the target audio.

[0098] In this embodiment, the offset calculation module 113 can be used to execute Figure 1 the step S130 shown. For a specific description of the offset calculation module 113, reference can be made to the description of the step S130.

[0099] The matching degree calculation module 114 is configured to determine a candidate start position in the known audio corresponding to the segment and increase the matching degree count value of the second audio segment corresponding to the candidate start position.

[0100] In this embodiment, the matching degree calculation module 114 can be used to execute Figure 1 the step S140 shown. For a specific description of the matching degree calculation module 114, reference can be made to the description of the step S140.

[0101] The result determination module 115 is configured to determine the known audio that matches the target audio retrieval based on the matching degree count values of the second audio segments of each known audio derived from multiple first audio segments of the target audio, and determine the audio start position corresponding to the target audio in the known audio.

[0102] In this embodiment, the result determination module 115 can be used to execute Figure 1 the steps S150 shown. For the specific description of the result determination module 115, reference can be made to the description of the steps S150.

[0103] In summary, an audio retrieval method, apparatus, and electronic device provided by this application can relatively independently extract the fingerprint features of each first audio segment in the target audio, and relatively independently execute the update of the matching degree count values of the second audio segments of the known audio based on the fingerprint features. Finally, the known audio that matches the target audio is determined according to the matching degree count values. In this way, when the number of valid audio segments is sufficient, even if individual audio segments in the target audio are lost or out of order, it will not have a great impact on the statistical result of the matching degree count values, improving the robustness of audio retrieval. Moreover, the calculation processing logic is simple, the retrieval processing speed is relatively fast, and it can better adapt to the streaming audio retrieval scenario.

[0104] In the embodiments provided by this application, it should be understood that the disclosed apparatus and method can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the drawings show the possible architectures, functions, and operations of apparatuses, methods, and computer program products according to multiple embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0105] In addition, the functional modules in each embodiment of this application can be integrated together to form an independent part, or each module can exist separately, or two or more modules can be integrated to form an independent part.

[0106] When the above-mentioned functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0107] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0108] As described above, these are only various implementation manners of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. An audio retrieval method, characterized in that, the method includes: extracting a first audio segment from a target audio, and determining fingerprint features of the first audio segment according to the spectrum of the first audio segment; detecting candidate audio segments with the same fingerprint features as the first audio segment from second audio segments of each known audio; determining a position offset according to the temporal position of the first audio segment in the target audio; for each of the candidate audio segments, determining a candidate start position in the known audio corresponding to the candidate audio segment according to the position offset, and increasing the matching degree count value of the second audio segment corresponding to the candidate start position; based on the matching degree count values of the second audio segments of each known audio obtained from multiple first audio segments, determining a known audio that matches the target audio retrieval, and determining the audio start position corresponding to the target audio in the known audio.

2. The method according to claim 1, characterized in that, the step of determining the fingerprint features of the first audio segment according to the spectrum of the first audio segment includes: performing a Fourier transform on the first audio segment to obtain spectrum data of the first audio segment in a preset dimension; the spectrum data includes amplitude values corresponding to multiple frequency identification values; dividing the spectrum data into multiple spectrum sub-bands according to a set frequency identification value; for each of the spectrum sub-bands, using the frequency identification value with the largest amplitude value in the spectrum sub-band as the characteristic value of the spectrum sub-band; using the set of characteristic values of the multiple spectrum sub-bands corresponding to the first audio segment as the fingerprint features of the first audio segment.

3. The method according to claim 1, characterized in that, the step of determining a position offset according to the temporal position of the first audio segment in the target audio includes: determining the offset length between the first audio segment and the start position of the target audio in time sequence as the position offset; the step of determining a candidate start position in the known audio corresponding to each candidate audio segment according to the position offset includes: for each of the candidate audio segments, determining the position in the known audio where the candidate audio segment is located and having an offset length before the candidate audio segment in time sequence as the candidate start position.

4. The method according to claim 1, characterized in that, the step of determining a known audio that matches the target audio retrieval and determining the audio start position corresponding to the target audio in the known audio based on the matching degree count values of the second audio segments of each known audio obtained from multiple first audio segments includes: detecting whether the distribution of the matching degree count values of the second audio segments of each known audio is a normal distribution; determining the known audio with the matching degree count value distribution being a normal distribution as the known audio retrieved for the target audio, and determining the corresponding position of the second audio segment with the highest matching degree count value as the audio start position.

5. The method according to claim 1, characterized in that, The step of determining the known audio that matches the target audio retrieval and determining the starting position of the target audio corresponding to the known audio based on the matching degree count value of the second audio segment of each known audio derived from multiple first audio segments includes: Detecting whether the matching degree count value of the second audio segment of each known audio reaches a preset threshold; Determining the known audio corresponding to the second audio segment whose matching degree count value first reaches the preset threshold as the known audio retrieved for the target audio, and determining the corresponding position of the second audio segment as the audio starting position.

6. The method according to claim 1, wherein, the step of extracting the first audio segment from the target audio includes: Performing sliding window data extraction on the target audio according to a set window length and a set step length to obtain a plurality of the first audio segments; wherein, the set step length is less than the set window length; The method further includes: For each known audio, performing sliding window data extraction on the known audio according to the set window length and the set step length to obtain a plurality of the second audio segments.

7. The method according to claim 1, wherein, the sampling frequency of each known audio is a preset sampling frequency; the method further includes: Performing downsampling on the target audio to convert the sampling frequency of the target audio to the preset sampling frequency.

8. An audio retrieval device, wherein, the device includes: A data acquisition module, configured to extract a first audio segment from a target audio and determine the fingerprint feature of the first audio segment according to the spectrum of the first audio segment; A feature extraction module, configured to detect a candidate audio segment having the same fingerprint feature as the first audio segment from the second audio segments of each known audio; An offset calculation module, configured to determine a position offset according to the temporal position of the first audio segment in the target audio; A matching degree calculation module, configured to determine a candidate starting position in the known audio corresponding to the segment and increase the matching degree count value of the second audio segment corresponding to the candidate starting position; A result determination module, configured to determine the known audio that matches the target audio retrieval and determine the audio starting position of the target audio corresponding to the known audio based on the matching degree count value of the second audio segment of each known audio derived from multiple first audio segments.

9. An electronic device, wherein, including a processor and a machine-readable storage medium, the machine-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are executed by the processor, the method according to any one of claims 1-7 is implemented.

10. A machine-readable storage medium, wherein, the machine-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are executed by one or more processors, the method according to any one of claims 1-7 is implemented.

Citation Information

Patent Citations

  • Audio processing method and device

    CN105825850A

  • Audio matching method, electronic equipment and storage medium

    CN111599378A