Audio matching method and device, electronic equipment and computer readable storage medium

CN115221351BActive Publication Date: 2026-05-29TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2022-07-26
Publication Date
2026-05-29

Smart Images

  • Figure CN115221351B_ABST
    Figure CN115221351B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose an audio matching method and device, electronic equipment and a computer readable storage medium. In the embodiments of the present application, a to-be-matched audio and an audio set are obtained, the audio set comprising an audio and a numerical matrix corresponding to an audio fingerprint of the audio; a fingerprint feature of the to-be-matched audio is extracted to obtain a target audio fingerprint corresponding to the to-be-matched audio, the target audio fingerprint comprising a plurality of fingerprint elements; each of the fingerprint elements in the target audio fingerprint is mapped to a target numerical value in a preset numerical interval to obtain a target numerical matrix corresponding to the target audio fingerprint, a difference between the target numerical value and an endpoint value of the preset numerical interval being within a preset range; and an audio matching the to-be-matched audio is searched from the audio set according to the target numerical matrix and the numerical matrix corresponding to the audio fingerprint. The embodiments of the present application can improve the speed of audio matching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio technology, specifically to an audio matching method, apparatus, electronic device, and computer-readable storage medium. Background Technology

[0002] With the development of society, there is an increasing amount of audio online. After a user obtains an audio clip, they can use the audio fingerprint in the clip to find audio that matches it.

[0003] However, when matching with audio fingerprints extracted using current audio fingerprinting methods, the matching speed is slow. Summary of the Invention

[0004] This application provides an audio matching method, apparatus, electronic device, and computer-readable storage medium, which can solve the technical problem of slow audio matching speed.

[0005] An audio matching method, comprising:

[0006] Obtain the audio to be matched and the audio set, wherein the audio set includes the audio and the numerical matrix corresponding to the audio fingerprint of the audio;

[0007] Fingerprint features are extracted from the above-mentioned audio to be matched to obtain the target audio fingerprint corresponding to the above-mentioned audio to be matched. The target audio fingerprint includes multiple fingerprint elements.

[0008] Each of the fingerprint elements in the target audio fingerprint is mapped to a target value within a preset value range to obtain the target value matrix corresponding to the target audio fingerprint. The difference between the target value and the endpoint value of the preset value range is within a preset range.

[0009] Based on the target numerical matrix and the numerical matrix corresponding to the audio fingerprint, search for audio that matches the audio to be matched from the audio set.

[0010] Accordingly, embodiments of this application provide an audio matching device, including:

[0011] The acquisition module is used to acquire the audio to be matched and the audio set, wherein the audio set includes the audio and the numerical matrix corresponding to the audio fingerprint of the audio;

[0012] The extraction module is used to extract fingerprint features from the audio to be matched to obtain the target audio fingerprint corresponding to the audio to be matched. The target audio fingerprint includes multiple fingerprint elements.

[0013] The mapping module is used to map each of the fingerprint elements in the target audio fingerprint to a target value within a preset value range, thereby obtaining a target value matrix corresponding to the target audio fingerprint. The difference between the target value and the endpoint value of the preset value range is within a preset range.

[0014] The search module is used to search for audio that matches the audio to be matched from the audio set based on the target numerical matrix and the numerical matrix corresponding to the audio fingerprint.

[0015] The numerical matrix corresponding to the audio fingerprint of the above audio includes the binary matrix corresponding to the audio fingerprint of the above audio.

[0016] Accordingly, the lookup module is specifically used to perform:

[0017] The above target numerical matrix is ​​subjected to binary mapping processing to obtain the target binary matrix corresponding to the above target audio fingerprint;

[0018] Based on the aforementioned target binary matrix and the binary matrix corresponding to the aforementioned audio fingerprint, search for audio that matches the aforementioned audio to be matched from the aforementioned audio set.

[0019] Optionally, the extraction module is specifically used to perform:

[0020] Using a trained audio fingerprint model, fingerprint features are extracted from the audio to be matched to obtain the target audio fingerprint corresponding to the audio to be matched.

[0021] The mapping module is specifically used for execution:

[0022] Using the trained audio fingerprint model, each fingerprint element in the target audio fingerprint is mapped to a target value within a preset range.

[0023] Optionally, the audio matching device further includes:

[0024] The training module is used to perform:

[0025] Obtain the training sample set of the audio fingerprint model to be trained, and extract fingerprint features from the training samples in the training sample set to obtain the sample audio fingerprints corresponding to the training samples.

[0026] Each element in the above sample audio fingerprint is mapped to a sample value within the above preset value range to obtain the sample value matrix corresponding to the above sample audio fingerprint.

[0027] Based on the above sample numerical matrix and the above endpoint values ​​corresponding to the above preset numerical interval, determine the target loss value of the above audio fingerprint model to be trained;

[0028] The audio fingerprint model to be trained is trained based on the target loss value to make it converge to the endpoint value, thus obtaining the trained audio fingerprint model.

[0029] Optionally, the training module is specifically used to perform:

[0030] Perform a short-time Fourier transform on the training samples in the above training sample set to obtain the time-frequency feature vectors corresponding to the above training samples;

[0031] Fingerprint features are extracted from the above time-frequency feature vectors to obtain the sample audio fingerprints corresponding to the above training samples.

[0032] Optionally, fingerprint feature extraction includes temporal feature extraction and spatial feature extraction.

[0033] Accordingly, the training module is specifically used to perform:

[0034] Time feature extraction is performed on the above time-frequency feature vectors to obtain the time feature vectors corresponding to the above training samples;

[0035] Spatial features are extracted from the above temporal feature vectors to obtain the sample audio fingerprints corresponding to the above training samples.

[0036] Optionally, the training module is specifically used to perform:

[0037] Obtain the target audio segments;

[0038] The target audio is subjected to windowing and frame segmentation to obtain the corresponding sample speech segments.

[0039] Based on the above sample speech segments, the training sample set is determined.

[0040] Optionally, the training module is specifically used to perform:

[0041] Temporal data augmentation is performed on the above sample speech segments to obtain the positive samples corresponding to the above sample speech segments;

[0042] Based on the positive samples corresponding to the above sample speech segments and the above sample speech segments, the training sample set is determined.

[0043] Optionally, the training module is specifically used to perform:

[0044] The time-frequency feature vectors mentioned above are subjected to frequency domain data augmentation to obtain the sample time-frequency feature vectors corresponding to the training samples.

[0045] Fingerprint features are extracted from the time-frequency feature vectors of the above samples to obtain the sample audio fingerprints corresponding to the above training samples.

[0046] Optionally, the training module is specifically used to perform:

[0047] Based on the above sample audio fingerprints, determine the contrastive loss value of the audio fingerprint model to be trained.

[0048] Based on the above sample numerical matrix and the endpoint values ​​corresponding to the above preset numerical interval, the quantization loss value of the above audio fingerprint model to be trained is determined.

[0049] Based on the above comparative loss value and the above quantization loss value, the target loss value of the above audio fingerprint model to be trained is obtained.

[0050] Optionally, the training module is specifically used to perform:

[0051] The first sample audio fingerprint is selected from the above sample audio fingerprints;

[0052] Based on the first sample audio fingerprint and the second sample audio fingerprint corresponding to the first sample audio fingerprint in the sample audio fingerprint, the first inner product value corresponding to the first sample audio fingerprint is determined. When the first sample audio fingerprint is the sample audio fingerprint corresponding to the first sample speech segment, the second sample audio fingerprint is the sample audio fingerprint of the first positive sample corresponding to the first sample speech segment. Alternatively, when the first sample audio fingerprint is the sample audio fingerprint corresponding to the first positive sample, the second sample audio fingerprint is the sample audio fingerprint of the first sample speech segment corresponding to the first positive sample.

[0053] Based on the first sample audio fingerprint, the second sample audio fingerprint, and the third sample audio fingerprint, the second inner product value corresponding to the first sample audio fingerprint is determined, and the third sample audio fingerprint is the sample audio fingerprint other than the first sample audio fingerprint and the second sample audio fingerprint among the sample audio fingerprints.

[0054] Based on the first inner product value and the second inner product value, the contrastive loss value of the audio fingerprint model to be trained is determined.

[0055] Furthermore, this application also provides an electronic device, including a processor and a memory, wherein the memory stores a computer program, and the processor is used to run the computer program in the memory to implement the audio matching method provided in this application.

[0056] Furthermore, embodiments of this application also provide a computer-readable storage medium storing a computer program adapted for loading by a processor to execute any of the audio matching methods provided in embodiments of this application.

[0057] Furthermore, this application also provides a computer program product, including a computer program that, when executed by a processor, implements any of the audio matching methods provided in this application.

[0058] In this embodiment, the audio to be matched and an audio set are first obtained. The audio set includes the audio files and numerical matrices corresponding to the audio fingerprints of the audio files. Then, fingerprint feature extraction is performed on the audio files to be matched to obtain the target audio fingerprint, which includes multiple fingerprint elements. Next, each fingerprint element in the target audio fingerprint is mapped to a target value within a preset numerical range, resulting in a target numerical matrix corresponding to the target audio fingerprint. The difference between the target value and the endpoint values ​​of the preset numerical range is within a preset range. Finally, based on the target numerical matrix and the numerical matrix corresponding to the audio fingerprint, the audio file that matches the audio file is searched for in the audio set.

[0059] In this embodiment of the application, each fingerprint element in the target audio fingerprint is mapped to a target value within a preset value range to obtain a target value matrix corresponding to the target audio fingerprint. The difference between the target value and the endpoint value of the preset value range is within a preset range, so that the audio matching the audio to be matched can be found based on the target value in the target value matrix and the value in the value matrix, thereby reducing the time to find the audio matching the audio to be matched and improving the speed of finding the audio matching the audio to be matched. Attached Figure Description

[0060] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0061] Figure 1 This is a schematic diagram of the audio matching process provided in an embodiment of this application;

[0062] Figure 2 This is a flowchart illustrating the audio matching method provided in an embodiment of this application;

[0063] Figure 3 This is a schematic diagram of the model training process provided in the embodiments of this application;

[0064] Figure 4 This is a flowchart illustrating the model structure provided in the embodiments of this application;

[0065] Figure 5 This is a flowchart illustrating the application of the model provided in the embodiments of this application;

[0066] Figure 6 This is a schematic diagram of the structure of the audio matching device provided in the embodiments of this application;

[0067] Figure 7 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0068] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0069] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0070] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0071] Key technologies in speech technology include Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and Voiceprint Recognition. Enabling computers to hear, see, speak, and feel is the future direction of human-computer interaction, with speech emerging as one of the most promising methods.

[0072] Machine Learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instruction-based learning.

[0073] The solutions provided in this application involve artificial intelligence technologies such as speech recognition and machine learning, which are specifically illustrated through the following embodiments.

[0074] This application provides an audio matching method, apparatus, electronic device, and computer-readable storage medium. The audio matching apparatus can be integrated into an electronic device, which may be a server or a terminal, etc.

[0075] The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, network acceleration services (Content Delivery Network, CDN), as well as big data and artificial intelligence platforms.

[0076] Furthermore, multiple servers can form a blockchain, with the servers being nodes on the blockchain.

[0077] The terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. The terminal and the server can be connected directly or indirectly through wired or wireless communication, which is not limited herein.

[0078] For example, such as Figure 1As shown, the terminal can acquire the audio to be matched and send it to the server. The server then extracts fingerprint features from the audio to be matched, obtaining the target audio fingerprint corresponding to the audio to be matched. The target audio fingerprint includes multiple fingerprint elements, and each fingerprint element in the target audio fingerprint is mapped to a target value within a preset value range, resulting in a target value matrix corresponding to the target audio fingerprint. The difference between the target value and the endpoint values ​​of the preset value range is within a preset range. Next, based on the target value matrix and the value matrix corresponding to the audio fingerprints of the audio in the audio set, the server searches for the audio that matches the audio to be matched in the audio set and returns the audio that matches the audio to be matched to the terminal.

[0079] Furthermore, in the embodiments of this application, "multiple" refers to two or more. The terms "first" and "second," etc., in the embodiments of this application are used for distinguishing descriptions and should not be construed as implying relative importance.

[0080] The following sections provide detailed descriptions of each example. It should be noted that the order in which the embodiments are described is not intended to limit the preferred order of the embodiments.

[0081] In this embodiment, the description will be from the perspective of the audio matching device. In order to facilitate the explanation of the audio matching method of this application, the following will describe the audio matching device integrated into the terminal in detail, that is, the terminal will be used as the execution subject for detailed explanation.

[0082] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of an audio matching method provided in this application. The audio matching method may include:

[0083] S201. Obtain the audio to be matched and the audio set. The audio set includes the audio and the numerical matrix corresponding to the audio fingerprint of the audio.

[0084] Upon receiving a acquisition command, the terminal can acquire the audio to be matched via its own microphone. Alternatively, it can acquire the audio to be matched via the microphone of another terminal, which then sends the audio to the terminal, thus acquiring the audio to be matched. The method by which the terminal acquires the audio to be matched can be selected according to the actual situation, and this embodiment does not limit it.

[0085] The terminal can send a request command to the server when acquiring the audio to be matched. The server then returns an audio set to the terminal based on the request, allowing the terminal to obtain the audio set. Alternatively, the terminal can send a request command to the server before acquiring the audio to be matched. The server then returns an audio set to the terminal based on the request, allowing the terminal to obtain the audio set and store it locally. The terminal then retrieves the audio set from local storage, ensuring that even when the terminal is offline, it can find the audio that matches the audio to be matched. This embodiment does not limit the method by which the terminal acquires the audio set.

[0086] It should be understood that the audio set sent by the server to the terminal may only include audio. After the terminal obtains the audio set, it performs feature extraction and mapping processing on the audio in the audio set to obtain the numerical matrix corresponding to the audio fingerprint of the audio in the audio set, and stores the numerical matrix corresponding to the audio fingerprint in the audio set.

[0087] Alternatively, the server can first perform feature extraction and mapping processing on the audio in the audio set to obtain the numerical matrix corresponding to the audio fingerprint of the audio in the audio set, store the numerical matrix corresponding to the audio fingerprint in the audio set, and then send the audio set to the terminal.

[0088] The process of feature extraction and mapping of audio in the audio set can be referred to as the process of fingerprint feature extraction and mapping of the audio to be matched, which will not be described in detail here.

[0089] S202. Extract fingerprint features from the audio to be matched to obtain the target audio fingerprint corresponding to the audio to be matched. The target audio fingerprint includes multiple fingerprint elements.

[0090] An audio fingerprint is a set of unique identifiers (such as symbols or numbers) obtained by extracting features from audio. A fingerprint element refers to a unique identifier within a set of unique identifiers.

[0091] The method for extracting fingerprint features from the audio to be matched to obtain the target audio fingerprint can be selected according to the actual situation. For example, a trained audio fingerprint model can be used to extract fingerprint features from the audio to be matched to obtain the target audio fingerprint. This embodiment does not limit this.

[0092] In some embodiments, fingerprint feature extraction is performed on the audio to be matched to obtain the target audio fingerprint corresponding to the audio to be matched, including:

[0093] Perform a short-time Fourier transform on the audio to be matched to obtain the target time-frequency feature vector of the audio to be matched;

[0094] Fingerprint features are extracted from the target time-frequency feature vector to obtain the target audio fingerprint corresponding to the audio to be matched.

[0095] In this embodiment, a Short-time Fourier Transform (STFT) is first performed on the audio to be matched to obtain the target time-frequency feature vector of the audio to be matched. This reduces the computational load of fingerprint feature extraction, improves the speed of fingerprint feature extraction, and simultaneously obtains the time-domain and frequency-domain information of the audio to be matched. This reduces the loss of the audio to be matched during the fingerprint feature extraction process, thereby improving the accuracy of the target audio.

[0096] In other embodiments, fingerprint feature extraction is performed on the target time-frequency feature vector to obtain the target audio fingerprint corresponding to the audio to be matched, including:

[0097] Time features are extracted from the target time-frequency feature vector to obtain the target time feature vector corresponding to the audio to be matched;

[0098] Spatial features are extracted from the target temporal feature vector to obtain the target audio fingerprint corresponding to the audio to be matched.

[0099] The target time-frequency feature vector includes the time-domain and frequency-domain information of the audio to be matched. For example, each row element of the target time-frequency feature vector includes the time-domain information of the audio to be matched, and each column element of the target time-frequency feature vector includes the frequency-domain information of the audio to be matched. Time feature extraction of the target time-frequency feature vector can be understood as extracting the features of each row element of the target time-frequency feature vector to obtain the target time feature vector. Spatial feature extraction of the target time feature vector can be understood as extracting the features of each column element of the target time feature vector.

[0100] The methods for extracting temporal features from the target time-frequency feature vector and extracting spatial features from the target time-frequency feature vector can be selected according to the actual situation. For example, temporal features can be extracted from the target time-frequency feature vector and spatial features can be extracted from the target time-frequency feature vector using a spatially separable convolutional neural network (SSCNN) in a trained audio fingerprint model, or temporal features can be extracted from the target time-frequency feature vector using a temporal convolutional layer and spatial features can be extracted from the target time-frequency feature vector using a spatial convolutional layer. This embodiment does not limit the scope of the extraction.

[0101] In this embodiment, the temporal and spatial features of the audio to be matched are extracted simultaneously, so that the target audio fingerprint includes both the temporal and frequency domain information of the audio to be matched, thereby improving the accuracy of the target audio fingerprint and thus improving the accuracy of the audio that matches the audio to be matched.

[0102] It should be understood that the process of extracting temporal and spatial features can be performed once or multiple times, and this embodiment does not limit it.

[0103] S203. Map each fingerprint element in the target audio fingerprint to a target value within a preset value range to obtain the target value matrix corresponding to the target audio fingerprint. The difference between the target value and the endpoint value of the preset value range is within a preset range.

[0104] The terminal can use a trained audio fingerprint model to map each fingerprint element in the target audio fingerprint to a target value within a preset numerical range, obtaining the target numerical matrix corresponding to the target audio fingerprint. Alternatively, the terminal can use a preset mapping function (which can be selected based on a preset numerical range; for example, when the preset numerical range is (-1, 1), the preset mapping function can be the hyperbolic tangent function tanh) to map each fingerprint element in the target audio fingerprint to a target value within a preset numerical range, obtaining the target numerical matrix corresponding to the target audio fingerprint. Alternatively, the terminal can also use a preset mapping matrix to map each fingerprint element in the target audio fingerprint to a target value within a preset numerical range.

[0105] The preset value range can be selected according to the actual situation. For example, the preset value range can be set to (-1, 1) or (0, 1). This embodiment does not limit it here.

[0106] The preset range can be set according to the actual situation. For example, the preset range can be [0, 1] or [0, 0.5]. This embodiment does not limit it here.

[0107] In this embodiment, each fingerprint element in the target audio fingerprint is mapped to a target value within a preset value range to obtain a target value matrix corresponding to the target audio fingerprint. The difference between the target value and the endpoint value of the preset value range is within a preset range, thereby limiting the fingerprint elements in the target audio fingerprint to target values ​​whose difference with the endpoint value of the preset value range is within a preset range, so as to speed up the search for audio that matches the audio to be matched.

[0108] S204. Based on the target numerical matrix and the numerical matrix corresponding to the audio fingerprint, find the audio that matches the audio to be matched from the audio set.

[0109] After obtaining the target numerical matrix and the numerical matrix, the terminal can match the target numerical matrix and the numerical matrix, and then use the audio corresponding to the numerical matrix that matches the target numerical matrix as the audio to be matched.

[0110] In this embodiment, since the difference between the target value in the target value matrix and the endpoint value of the preset value interval is within a preset range, and the difference between the value in the value matrix and the endpoint value of the preset value interval is also within a preset range, the audio that matches the audio to be matched can be found more quickly from the audio set based on the target value matrix and the value matrix corresponding to the audio fingerprint.

[0111] In some embodiments, in order to find the audio that matches the audio to be matched more quickly from the audio set, the numerical matrix corresponding to the audio fingerprint of the audio includes the binary matrix corresponding to the audio fingerprint of the audio.

[0112] Based on the target numerical matrix and the numerical matrix corresponding to the audio fingerprint, search the audio set for audio that matches the audio to be matched, including:

[0113] The target numerical matrix is ​​subjected to binary mapping to obtain the target binary matrix corresponding to the target audio fingerprint;

[0114] Based on the target binary matrix and the binary matrix corresponding to the audio fingerprint, find the audio that matches the audio to be matched from the audio set.

[0115] Binary mapping refers to using two values ​​to represent the target value in the target numerical matrix. For example, using 0 and 1 to represent the target value in the target numerical matrix, that is, the elements in the target binary matrix are 0 or 1.

[0116] The method for performing binary mapping on the target numerical matrix to obtain the target binary matrix corresponding to the target audio fingerprint can be selected according to the actual situation. For example, the target values ​​in the target numerical matrix can be input into the sign function for binary mapping to obtain the target binary matrix.

[0117]

[0118] Where t represents the target value.

[0119] Alternatively, the target numerical values ​​in the target numerical matrix can be input into a piecewise function for binary mapping processing to obtain the target binary matrix. This embodiment does not impose any limitations on this.

[0120] In this embodiment, since the target numerical matrix is ​​mapped to the target binary matrix, the audio that matches the audio to be matched can be found more quickly from the audio set based on the target binary matrix and the binary matrix corresponding to the audio fingerprint.

[0121] In other embodiments, to more accurately find audio that matches the audio to be matched, fingerprint feature extraction is performed on the audio to be matched to obtain the target audio fingerprint corresponding to the audio to be matched, including:

[0122] The audio to be matched is windowed and framed to obtain each target speech segment of the audio to be matched.

[0123] Fingerprint features are extracted from the target speech segment to obtain the target sub-audio fingerprint corresponding to the target speech segment.

[0124] Each fingerprint element in the target audio fingerprint is mapped to a target value within a preset value range, resulting in a target value matrix corresponding to the target audio fingerprint, including:

[0125] Each fingerprint element in the target sub-audio fingerprint is mapped to a target value within a preset value range to obtain the target sub-value matrix corresponding to the target sub-audio fingerprint.

[0126] Based on the target numerical matrix and the numerical matrix corresponding to the audio fingerprint, search the audio set for audio that matches the audio to be matched, including:

[0127] Based on the target sub-numerical matrix and the numerical matrix corresponding to the audio fingerprint, search for audio that matches the audio to be matched from the audio set.

[0128] At this point, the terminal can fuse the target sub-numerical matrix to obtain the target numerical matrix, and then search for the audio that matches the audio to be matched from the audio set based on the target numerical matrix and the numerical matrix corresponding to the audio fingerprint.

[0129] Alternatively, the audio set includes a numerical matrix corresponding to the audio fingerprint and a numerical submatrix corresponding to the speech segment. The terminal can match the target numerical submatrix with the numerical submatrix, and then determine the audio that matches the audio to be matched based on the speech segment corresponding to the numerical submatrix that matches the target numerical submatrix.

[0130] Extracting fingerprint features from a target speech segment to obtain the target sub-audio fingerprint corresponding to the target speech segment may include:

[0131] Perform a short-time Fourier transform on the target speech segment to obtain the target sub-time-frequency feature vector of the target speech segment;

[0132] Fingerprint features are extracted from the target sub-time-frequency feature vector to obtain the target sub-audio fingerprint corresponding to the target speech segment.

[0133] Optionally, fingerprint feature extraction is performed on the target sub-time-frequency feature vector to obtain the target sub-audio fingerprint corresponding to the target speech segment, which may include:

[0134] Temporal features are extracted from the target sub-time-frequency feature vector to obtain the target sub-time feature vector corresponding to the target speech segment;

[0135] Spatial features are extracted from the target sub-temporal feature vector to obtain the target sub-audio fingerprint corresponding to the target speech segment.

[0136] When extracting fingerprint features from the audio to be matched using a trained audio fingerprint model to obtain the target audio fingerprint corresponding to the audio to be matched, and mapping each fingerprint element in the target audio fingerprint to a target value within a preset numerical range using the trained audio fingerprint model, the process before extracting fingerprint features from the audio to be matched using the trained audio fingerprint model to obtain the target audio fingerprint corresponding to the audio to be matched also includes:

[0137] Obtain the training sample set of the audio fingerprint model to be trained, and extract fingerprint features from the training samples in the training sample set to obtain the sample audio fingerprints corresponding to the training samples.

[0138] Each element in the sample audio fingerprint is mapped to a sample value within a preset range to obtain the sample value matrix corresponding to the sample audio fingerprint.

[0139] The target loss value of the audio fingerprint model to be trained is determined based on the sample numerical matrix and the endpoint values ​​corresponding to the preset numerical range.

[0140] The audio fingerprint model to be trained is trained according to the target loss value so that the audio fingerprint model to be trained converges to the endpoint value, thus obtaining the trained audio fingerprint model.

[0141] In this process, the sample values ​​in the sample value matrix and the endpoint values ​​corresponding to the preset value interval can be substituted into the loss function to obtain the target loss value of the audio fingerprint model to be trained. The type of loss function can be selected according to the actual situation. For example, the loss function can be a quantization loss function or a cross-entropy loss function. This embodiment does not limit it here.

[0142] If the target loss value meets the preset condition, it means that the audio fingerprint model to be trained converges to the endpoint value. That is, it means that through the audio fingerprint model to be trained, each fingerprint element in the target audio fingerprint can be mapped to the target value within the preset value range. If the difference between the target value and the endpoint value of the preset value range is within the preset range, then the audio fingerprint model to be trained is regarded as the trained audio fingerprint model.

[0143] If the target loss value does not meet the preset conditions, the network parameters of the audio fingerprint model to be trained are updated according to the target loss value, and the process returns to the step of extracting fingerprint features from the training samples in the training sample set to obtain the sample audio fingerprints corresponding to the training samples.

[0144] In this embodiment, when training the audio fingerprint model to be trained, each element in the sample audio fingerprint is first mapped to a sample value within a preset value range to obtain a sample value matrix corresponding to the sample audio fingerprint. Then, based on the sample value matrix and the endpoint values ​​corresponding to the preset value range, the target loss value of the audio fingerprint model to be trained is determined. Finally, the audio fingerprint model to be trained is trained based on the target loss value so that the audio fingerprint model to be trained converges to the endpoint value, that is, the sample values ​​move closer to the endpoint value. This allows the trained audio fingerprint model to map each fingerprint element in the target audio fingerprint to a target value within a preset value range, obtaining a target value matrix corresponding to the target audio fingerprint. The difference between the target value and the endpoint value of the preset value range is within a preset range, thereby improving the speed of finding audio that matches the audio to be matched.

[0145] In other embodiments, fingerprint feature extraction is performed on training samples in the training sample set to obtain sample audio fingerprints corresponding to the training samples, including:

[0146] Perform a short-time Fourier transform on the training samples in the training sample set to obtain the time-frequency feature vectors corresponding to the training samples;

[0147] Fingerprint features are extracted from the time-frequency feature vector to obtain the sample audio fingerprints corresponding to the training samples.

[0148] For example, in the process of performing short-time Fourier transform on the training samples in the training sample set, the window length of the short-time Fourier transform is set to 1024 sampling points, the window shift is set to 256 sampling points, and the feature dimension is set to 256.

[0149] In this embodiment, a short-time Fourier transform is first performed on the training samples to obtain the time-frequency feature vector corresponding to the training samples, which replaces the spectral features. This improves the training speed of the audio fingerprint model to be trained, reduces the loss of training samples, and improves the matching effect of the trained audio fingerprint model on the audio to be matched.

[0150] It should be understood that a short-time Fourier transform can be performed on the training samples in the training sample set through the audio fingerprint model to be trained, or a short-time Fourier transform can be performed on the training samples in the training sample set through a convolutional layer, which may not be located in the audio fingerprint model to be trained. Then, the time-frequency feature vector is extracted by the audio fingerprint model to be trained to obtain the sample audio fingerprint corresponding to the training sample.

[0151] In other embodiments, fingerprint feature extraction includes temporal feature extraction and spatial feature extraction. Fingerprint feature extraction is performed on the time-frequency feature vector to obtain the sample audio fingerprint corresponding to the training sample, including:

[0152] Time features are extracted from the time-frequency feature vector to obtain the time feature vector corresponding to the training sample;

[0153] Spatial features are extracted from the temporal feature vector to obtain the sample audio fingerprint corresponding to the training sample.

[0154] When training the audio fingerprint model, the training process involves extracting temporal features from the time-frequency feature vector and spatial features from the time-frequency feature vector. This reduces the dimensionality of the time-frequency feature vector while preserving the temporal information of the training samples. As a result, the trained audio fingerprint model can extract temporal features from the target time-frequency feature vector and spatial features from the target time-frequency feature vector, improving the accuracy of finding audio that matches the target audio. Consequently, when performing audio matching using the trained audio fingerprint model, both matching speed and accuracy are improved.

[0155] It should be understood that spatial feature extraction can be performed first, followed by temporal feature extraction. That is, spatial feature extraction can be performed on the time-frequency feature vector first to obtain the spatial feature vector, and then temporal feature extraction can be performed on the spatial feature vector to obtain the sample audio fingerprint corresponding to the training sample.

[0156] In other embodiments, obtaining the training sample set includes:

[0157] Obtain the target audio segments;

[0158] Windowing and frame segmentation are performed on the target audio to obtain the corresponding sample speech segments;

[0159] The training sample set is determined based on the sample speech segments.

[0160] In this embodiment, after obtaining each target audio segment, the target audio is not directly combined into a training sample set. Instead, the target audio is windowed and framed to obtain each sample speech segment, and then a training sample set is formed based on the sample speech segments.

[0161] Then, during the training process of the audio fingerprint model to be trained based on the training sample set, sample speech segments of the same target audio and sample speech segments of different target audio are randomly selected from the training sample set to form the same batch of training samples, thereby interrupting the training samples. Finally, the audio fingerprint model to be trained is trained each time based on the same batch of training samples.

[0162] For example, during the windowing and framing process, the terminal converts each target audio segment into a WAV format with a sampling rate of 8000, mono, and a quantization bit depth of 16. Then, it uses a window with a timing length of 1 second, moving once every half second, to obtain sample speech segments.

[0163] To improve the robustness of the trained audio fingerprint model, a training sample set is determined based on sample speech segments, including:

[0164] Temporal data augmentation is performed on the sample speech segments to obtain the positive samples corresponding to the sample speech segments;

[0165] The training sample set is determined based on the positive samples corresponding to the sample speech segments and the sample speech segments themselves.

[0166] Among them, performing time-domain data augmentation on sample speech segments can refer to adding different types of noise, changing the speed of sound, or changing the pitch of sample speech segments. Noise can include background music, human voices, reverberation, and echoes, and the noise level can be between -10dB and 30dB.

[0167] Temporal data augmentation is performed on the sample speech segments to obtain the positive samples corresponding to the sample speech segments. Then, based on the positive samples corresponding to the sample speech segments and the sample speech segments, the training sample set is determined, thereby making the trained audio fingerprint model trained based on the training sample set more robust.

[0168] To further improve the robustness of the trained audio fingerprint model, fingerprint features are extracted from the time-frequency feature vector to obtain the audio fingerprints corresponding to the training samples, including:

[0169] The time-frequency feature vectors are augmented in the frequency domain to obtain the sample time-frequency feature vectors corresponding to the training samples.

[0170] Fingerprint features are extracted from the time-frequency feature vectors of the samples to obtain the audio fingerprints of the training samples.

[0171] The method of frequency domain data augmentation can be selected according to the actual situation. For example, the terminal can randomly set the elements in the time-frequency feature vector to zero, or the method of frequency domain data augmentation can also be frequency domain masking, which refers to a strong pure tone masking a weak pure tone that is emitted in the vicinity at the same time. This embodiment does not limit it.

[0172] In this embodiment, not only is data augmentation processing performed on the sample speech segments in the time domain, but also on the time-frequency feature vectors in the frequency domain, thereby further improving the robustness of the trained audio fingerprint model.

[0173] To further improve the robustness of the trained audio fingerprint model, before determining the target loss value of the audio fingerprint model to be trained based on the endpoint values ​​corresponding to the sample numerical matrix and the preset numerical range, the following steps are also included:

[0174] Based on the sample audio fingerprints, determine the contrastive loss value of the audio fingerprint model to be trained;

[0175] Based on the sample numerical matrix and the endpoint values ​​corresponding to the preset numerical intervals, the target loss value of the audio fingerprint model to be trained is determined, including:

[0176] The quantization loss value of the audio fingerprint model to be trained is determined based on the sample numerical matrix and the endpoint values ​​corresponding to the preset numerical range.

[0177] The target loss value for the audio fingerprint model to be trained is determined by comparing the loss value and quantizing the loss value.

[0178] In this embodiment, instead of directly using the quantization loss value as the target loss value, the target loss value is determined based on the comparison loss value and the quantization loss value, thereby improving the robustness of the trained audio fingerprint model.

[0179] The quantization loss value of the audio fingerprint model to be trained can be obtained by substituting the sample numerical matrix and the endpoint values ​​corresponding to the preset numerical interval into the following formula:

[0180]

[0181] L q Let z represent the quantization loss value, N represent the number of training samples in a batch, S represent the set of training samples in a batch, and z represent the training loss value. i The sample audio fingerprint, z, represents the sample speech segment. j The sample audio fingerprint represents the positive sample of a sample speech segment, tanh(z). i ) represents the sample audio fingerprint z i The corresponding sample numerical matrix, tanh(z) j ) represents the numerical matrix corresponding to the sample audio fingerprint of the positive samples of the sample speech segment. ||tanh(z i )-1||1 means first solve tanh(z) i The L1 norm of ) is then subtracted by 1, that is, || ||1 represents the quantization function and 1 represents the endpoint value.

[0182] The target loss value can be obtained by substituting the contrast loss value and the quantization loss value into the following formula:

[0183] L=αL c +(1-α)L q

[0184] L represents the target loss value, L c This represents the comparison loss value, and α represents the weighting coefficient, which can be set according to the actual situation.

[0185] Based on the sample audio fingerprints, determine the contrastive loss value for the audio fingerprint model to be trained, including:

[0186] The first sample audio fingerprint is selected from the sample audio fingerprints;

[0187] Based on the first sample audio fingerprint and the second sample audio fingerprint corresponding to the first sample audio fingerprint in the sample audio fingerprint, the first inner product value corresponding to the first sample audio fingerprint is determined. When the first sample audio fingerprint is the sample audio fingerprint corresponding to the first sample speech segment, the second sample audio fingerprint is the sample audio fingerprint of the first positive sample corresponding to the first sample speech segment. Alternatively, when the first sample audio fingerprint is the sample audio fingerprint corresponding to the first positive sample, the second sample audio fingerprint is the sample audio fingerprint of the first sample speech segment corresponding to the first positive sample.

[0188] Based on the first sample audio fingerprint, the second sample audio fingerprint, and the third sample audio fingerprint, determine the second inner product value corresponding to the first sample audio fingerprint. The third sample audio fingerprint is the sample audio fingerprint other than the first and second sample audio fingerprints.

[0189] The contrastive loss value of the audio fingerprint model to be trained is determined based on the first inner product value and the second inner product value.

[0190] Specifically, when the first sample audio fingerprint is the sample audio fingerprint corresponding to the first sample speech segment, the first sample audio fingerprint and the second sample audio fingerprint can be inner-producted to obtain the inner-product value corresponding to the first sample audio fingerprint, and the first sample audio fingerprint and the third sample audio fingerprint can be inner-producted to obtain the inner-product value corresponding to the third sample audio fingerprint.

[0191] When the first sample audio fingerprint is the sample audio fingerprint corresponding to the first positive sample, the second sample audio fingerprint is inner-producted with the first sample audio fingerprint to obtain the inner-product value corresponding to the first sample audio fingerprint. The second sample audio fingerprint is inner-producted with the third sample audio fingerprint to obtain the inner-product value corresponding to the third sample audio fingerprint.

[0192] Then, the inner product value corresponding to the first sample audio fingerprint is divided by a preset constant to obtain the first quotient. The inner product value corresponding to each third sample audio fingerprint is divided by the preset constant to obtain the second quotient.

[0193] Next, exponential operations are performed on both the first and second quotients to obtain the first exponential result corresponding to the first quotient and the second exponential result corresponding to the second quotient. The first exponential result is the first inner product value of the first sample audio fingerprint. The first exponential result and the second exponential result are then summed to obtain the second inner product value of the first sample audio fingerprint.

[0194] Finally, the first exponential result is divided by the second inner product value to obtain the division result. A logarithmic operation is then performed on the division result to obtain the first contrastive loss value of the first sample audio fingerprint. The process returns to the step of filtering out the first sample audio fingerprint from the sample audio fingerprints, thus obtaining the first contrastive loss value of each first sample audio fingerprint. These first contrastive loss values ​​are then summed to obtain the contrastive loss value of the audio fingerprint model to be trained.

[0195] In other words, by substituting the first inner product value and the second inner product value into the following contrastive learning function formula, we obtain the contrastive loss value of the audio fingerprint model to be trained:

[0196]

[0197]

[0198] Where i = 2^k - 1, j = 2^k, l(i, j) represents the first contrastive loss value of each first sample audio fingerprint. τ represents a preset constant, i represents the first sample speech segment, j represents the first positive sample, y represents the first positive sample, the second sample speech segment, or a positive sample representing the second sample speech segment, and a ij Let a represent the inner product value of the first sample audio fingerprint. When y = j, a iy Let a represent the inner product value of the first sample audio fingerprint. When y ≠ j, a iy This represents the inner product value corresponding to the third sample audio fingerprint.

[0199] exp() represents an exponential function with base e, exp(a ij / τ) represents the first inner product value, exp(a iy / τ) represents the first exponent result or the second exponent result. This represents the second inner product value of the first sample audio fingerprint.

[0200] At this time, the above z i This can be used to represent the sample audio fingerprint corresponding to the first sample speech segment, as mentioned above. j It can represent the sample audio fingerprint of the first positive sample.

[0201] It should be understood that since the training samples in the training sample set can be either sample speech segments or positive samples of sample speech segments, the third sample audio fingerprint can be either the sample audio fingerprint corresponding to the second sample speech segment or the sample audio fingerprint of the positive sample corresponding to the second sample speech segment. The second sample speech segment is any sample speech segment in the training sample set other than the first sample speech segment.

[0202] The first sample speech segment can also be called the reference sample, and the second sample speech segment and the positive sample corresponding to the second sample speech segment can also be called the negative sample.

[0203] In this embodiment, the first sample audio fingerprint is the sample audio fingerprint corresponding to the first sample speech segment, that is, the first sample audio fingerprint is the sample audio fingerprint corresponding to the sample speech segment, or the first sample audio fingerprint can also be the sample audio fingerprint corresponding to the first positive sample, that is, the first sample audio fingerprint is the sample audio fingerprint corresponding to the positive sample.

[0204] To further improve the audio matching performance of the trained audio fingerprint model, a first inner product value corresponding to the first sample audio fingerprint is determined based on the first sample audio fingerprint and the second sample audio fingerprint corresponding to the first sample audio fingerprint in the sample audio fingerprint. A second inner product value corresponding to the first sample audio fingerprint is determined based on the first sample audio fingerprint, the second sample audio fingerprint, and the third sample audio fingerprint, including:

[0205] Based on the first sample audio fingerprint and the second sample audio fingerprint corresponding to the first sample audio fingerprint in the sample audio fingerprint, determine the first inner product value and the third inner product value corresponding to the first sample audio fingerprint;

[0206] Based on the first sample audio fingerprint, the second sample audio fingerprint, and the third sample audio fingerprint, determine the second inner product value and the fourth inner product value corresponding to the first sample audio fingerprint.

[0207] Based on the first inner product value and the second inner product value, the contrastive loss value of the audio fingerprint model to be trained is determined, including:

[0208] The contrastive loss value of the audio fingerprint model to be trained is determined based on the first inner product value, the second inner product value, the third inner product value, and the fourth inner product value.

[0209] Specifically, when the first sample audio fingerprint is the sample audio fingerprint corresponding to the first sample speech segment, and the second sample audio fingerprint is the sample audio fingerprint of the first positive sample corresponding to the first sample speech segment, the first sample audio fingerprint and the second sample audio fingerprint can be inner-producted to obtain the first inner-product value of the first sample audio fingerprint. The inner-product of the second sample audio fingerprint and the first sample audio fingerprint can then be inner-producted to obtain the third inner-product value of the first sample audio fingerprint. The result of the inner-product of the first and third sample audio fingerprints is then added to the first inner-product value to obtain the second inner-product value corresponding to the first sample audio fingerprint. Finally, the result of the inner-product of the second and third sample audio fingerprints is added to the third inner-product value to obtain the fourth inner-product value corresponding to the first sample audio fingerprint.

[0210] When the first sample audio fingerprint is the sample audio fingerprint corresponding to the first positive sample, and the second sample audio fingerprint is the sample audio fingerprint of the first sample speech segment corresponding to the first positive sample, the first sample audio fingerprint and the second sample audio fingerprint can be inner-producted to obtain the third inner-product value of the first sample audio fingerprint. The second sample audio fingerprint and the first sample audio fingerprint can be inner-producted to obtain the first inner-product value of the first sample audio fingerprint. The result of the inner product of the first sample audio fingerprint and the third sample audio fingerprint is added to the third inner-product value to obtain the fourth inner-product value corresponding to the first sample audio fingerprint. The result of the inner product of the second sample audio fingerprint and the third sample audio fingerprint is added to the first inner-product value to obtain the second inner-product value corresponding to the first sample audio fingerprint.

[0211] In other words, the first inner product is obtained by taking the inner product of the sample audio fingerprint of the first sample speech segment and the sample audio fingerprint of the first positive sample. The second inner product is obtained by taking the inner product of the sample audio fingerprint of the first positive sample and the sample audio fingerprint of the first sample speech segment. The third inner product is obtained by adding the result of taking the inner product of the sample audio fingerprint of the first sample speech segment and the third sample audio fingerprint to the first inner product value. The fourth inner product is obtained by adding the result of taking the inner product of the sample audio fingerprint of the first positive sample and the third sample audio fingerprint to the third inner product value.

[0212] The calculation methods for the first inner product value, the second inner product value, the third inner product value, and the fourth inner product value in this embodiment can be referred to the calculation process of the first inner product value and the second inner product value described above. This embodiment will not repeat the details here.

[0213] Alternatively, one can first determine the first contrast loss value of the first sample speech segment based on the first inner product value and the second inner product value, then determine the first sample contrast loss value of the first positive sample based on the third inner product value and the fourth inner product value, and finally determine the contrast loss value based on the first contrast loss value and the first sample contrast loss value.

[0214] At this point, the comparison loss value can be expressed using the following formula:

[0215]

[0216] i = 2k-1, j = 2k, l(j, i) represents the first sample contrast loss value.

[0217] The method for calculating the first sample contrast loss value of the first positive sample corresponding to the first sample speech segment can be referred to in the process of calculating the first contrast loss value of the first sample audio fingerprint, which will not be repeated here in this embodiment.

[0218] It should be understood that in this embodiment, the first sample speech segment and the second sample speech segment are relative concepts, as are the first sample audio fingerprint and the third sample audio fingerprint. For example, the training sample set includes sample speech segment 1, positive sample 2 corresponding to sample speech segment 1, sample speech segment 3, and positive sample 4 corresponding to sample speech segment 2. The sample audio fingerprints include sample audio fingerprint a corresponding to sample speech segment 1, sample audio fingerprint b corresponding to positive sample 2, sample audio fingerprint c corresponding to sample speech segment 3, and sample audio fingerprint d corresponding to positive sample 4.

[0219] When the first sample audio fingerprint is sample audio fingerprint a, the first sample speech segment is sample speech segment 1, the second sample audio fingerprint is sample audio fingerprint b, the second sample speech segment is sample speech segment 3, and the third sample audio fingerprint is sample audio fingerprint c and sample audio fingerprint d.

[0220] When the first sample audio fingerprint is sample audio fingerprint b, the first sample speech segment is sample speech segment 1, the second sample audio fingerprint is sample audio fingerprint a, the second sample speech segment is sample speech segment 3, and the third sample audio fingerprint is sample audio fingerprint c and sample audio fingerprint d.

[0221] When the first sample audio fingerprint is sample audio fingerprint c, the first sample speech segment is sample speech segment 3, the second sample audio fingerprint is sample audio fingerprint d, the second sample speech segment is sample speech segment 1, the second sample audio fingerprint is sample audio fingerprint a and sample audio fingerprint b.

[0222] When the first sample audio fingerprint is sample audio fingerprint d, the first sample speech segment is sample speech segment 1, the second sample audio fingerprint is sample audio fingerprint c, the second sample speech segment is sample speech segment 3, and the third sample audio fingerprint is sample audio fingerprint a and sample audio fingerprint b.

[0223] As described above, in this embodiment, the audio to be matched and an audio set are first obtained. The audio set includes the audio files and the numerical matrix corresponding to the audio fingerprints of the audio files. Then, fingerprint features are extracted from the audio files to be matched to obtain the target audio fingerprint corresponding to the audio files to be matched. The target audio fingerprint includes multiple fingerprint elements. Next, each fingerprint element in the target audio fingerprint is mapped to a target value within a preset numerical range to obtain the target numerical matrix corresponding to the target audio fingerprint. The difference between the target value and the endpoint values ​​of the preset numerical range is within a preset range. Finally, based on the target numerical matrix and the numerical matrix corresponding to the audio fingerprint, the audio files that match the audio files to be matched are searched from the audio set.

[0224] In this embodiment of the application, each fingerprint element in the target audio fingerprint is mapped to a target value within a preset value range to obtain a target value matrix corresponding to the target audio fingerprint. The difference between the target value and the endpoint value of the preset value range is within a preset range, so that the audio matching the audio to be matched can be found based on the target value in the target value matrix and the value in the value matrix, thereby reducing the time to find the audio matching the audio to be matched and improving the speed of finding the audio matching the audio to be matched.

[0225] Based on the methods described in the above embodiments, the following examples will provide further detailed explanations.

[0226] Please see Figure 3 , Figure 3 This is a schematic flowchart illustrating the model training method provided in an embodiment of this application. The model training method may include:

[0227] S301. The terminal acquires each segment of target audio and performs windowing and frame-segmentation processing on the target audio to obtain each sample speech segment corresponding to each segment of target audio.

[0228] S302. The terminal performs time-domain data augmentation on the sample speech segments to obtain the positive samples corresponding to the sample speech segments, and determines the training sample set based on the positive samples corresponding to the sample speech segments and the sample speech segments.

[0229] S303. The terminal performs a short-time Fourier transform on the training samples in the training sample set to obtain the time-frequency feature vectors corresponding to the training samples.

[0230] S304. The terminal performs frequency domain data augmentation on the time-frequency feature vector to obtain the sample time-frequency feature vector corresponding to the training sample.

[0231] For example, such as Figure 4As shown, after obtaining the target audio, the target audio is windowed and framed to obtain sample speech segments. The samples are then shuffled so that the training samples in the same batch include both sample speech segments of the same target audio and sample speech segments of different target audios.

[0232] S305. The terminal extracts temporal features from the sample time-frequency feature vector through the spatially separable convolutional layer in the audio fingerprint model to be trained, thereby obtaining the temporal feature vector corresponding to the training sample, and extracts spatial features from the temporal feature vector to obtain the initial sample audio fingerprint corresponding to the training sample.

[0233] Spatially separable convolutional layers consist of stacked temporal and spatial convolutions. For example, in an audio fingerprint model to be trained, spatially separable convolutional layers can be stacked as follows: Figure 4 As shown ( Figure 4 In this case, Groupnorm represents group normalization. At this time, the spatially separable convolutional layer can perform multiple temporal and spatial feature extractions. That is, after spatial feature extraction is performed on the temporal feature vector, a spatial feature vector is obtained, and then temporal and spatial feature extractions are performed on the spatial feature vector again.

[0234] It should be understood that during temporal and spatial feature extraction, the elements in the matrix can be mapped to values ​​within a preset range, thereby further improving the accuracy of the trained audio fingerprint model. That is, the elements in the temporal feature vector are mapped to values ​​within a preset range to obtain a temporal numerical matrix. Then, spatial features are extracted from the temporal numerical matrix to obtain the initial sample audio fingerprint. Next, the elements in the initial sample audio matrix are mapped to values ​​within a preset range to obtain the spatial numerical matrix. Finally, the spatial numerical matrix is ​​input into the fully connected layer.

[0235] S306. The terminal performs dimensionality reduction on the initial sample audio fingerprint through the fully connected layer in the audio fingerprint model to be trained, and obtains the sample audio fingerprint.

[0236] By reducing the dimensionality of the initial sample audio fingerprints, the computational cost of the audio fingerprint model to be trained is reduced, while maintaining the independence between sample audio fingerprints.

[0237] Optionally, the initial sample audio fingerprint can be split first using the Split function in the fully connected layer to obtain the splitting result, and then the splitting result can be input into the neurons in the fully connected layer.

[0238] S307. The terminal filters out the first sample audio fingerprint from the sample audio fingerprints, and determines the first inner product value and the third inner product value corresponding to the first sample audio fingerprint based on the first sample audio fingerprint and the second sample audio fingerprint corresponding to the first sample audio fingerprint in the sample audio fingerprint. When the first sample audio fingerprint is the sample audio fingerprint corresponding to the first sample speech segment, the second sample audio fingerprint is the sample audio fingerprint of the first positive sample corresponding to the first sample speech segment. Alternatively, when the first sample audio fingerprint is the sample audio fingerprint corresponding to the first positive sample, the second sample audio fingerprint is the sample audio fingerprint of the first sample speech segment corresponding to the first positive sample.

[0239] S308. The terminal determines the second inner product value and the fourth inner product value corresponding to the first sample audio fingerprint based on the first sample audio fingerprint, the second sample audio fingerprint and the third sample audio fingerprint. The third sample audio fingerprint is the sample audio fingerprint other than the first sample audio fingerprint and the second sample audio fingerprint.

[0240] S309. The terminal determines the contrast loss value of the audio fingerprint model to be trained based on the first inner product value, the second inner product value, the third inner product value, and the fourth inner product value.

[0241] S3010. The terminal maps the elements in the sample audio fingerprint to sample values ​​within a preset value range, obtains the sample value matrix corresponding to the sample audio fingerprint, and determines the quantization loss value of the audio fingerprint model to be trained based on the sample value matrix and the endpoint values ​​corresponding to the preset value range.

[0242] S3011. The terminal determines the target loss value of the audio fingerprint model to be trained based on the comparison loss value and the quantization loss value.

[0243] S3012. If the target loss value meets the preset conditions, the terminal will use the audio fingerprint model to be trained as the already trained audio fingerprint model. If the target loss value does not meet the preset conditions, the network parameters of the audio fingerprint model to be trained will be updated according to the target loss value, and the process will return to step S303.

[0244] In this embodiment, when training the audio fingerprint model to be trained, time-domain data augmentation and frequency-domain data augmentation are performed on the sample speech segments simultaneously, thereby improving the coverage of the trained audio fingerprint model and enhancing its robustness to complex application scenarios.

[0245] By performing a short-time Fourier transform on the training samples, the time-frequency feature vector corresponding to the training samples can be obtained to replace the spectral features. This improves the training speed of the audio fingerprint model and reduces the loss of training samples, thereby improving the matching effect of the trained audio fingerprint model on the audio to be matched.

[0246] By extracting temporal and spatial features from the training samples, sample audio fingerprints are obtained. These fingerprints are then transformed into sample values ​​within a preset range, and the sample values ​​converge to the endpoints of this range. This improves the accuracy of finding audio that matches the audio to be matched, thus ensuring accuracy while increasing matching speed when using the trained audio fingerprint model for audio matching.

[0247] The specific implementation method and corresponding beneficial effects in this embodiment can be referred to the above audio matching method embodiment, which will not be repeated here.

[0248] Please see Figure 5 , Figure 5 This is a flowchart illustrating the model application method provided in an embodiment of this application. The model application method may include:

[0249] S501, The terminal obtains the audio to be matched and the audio set, the audio set including the audio and the binary matrix corresponding to the audio fingerprint of the audio.

[0250] S502. The terminal performs a short-time Fourier transform on the audio to be matched to obtain the target time-frequency feature vector of the audio to be matched.

[0251] S503. The terminal extracts temporal features from the target time-frequency feature vector through the spatially separable convolutional layer in the trained audio fingerprint model to obtain the target time feature vector corresponding to the audio to be matched, and extracts spatial features from the target time feature vector to obtain the audio fingerprint before dimensionality reduction corresponding to the audio to be matched.

[0252] It should be understood that during temporal and spatial feature extraction, the elements in the matrix can be mapped to values ​​within a preset range, thereby further improving the accuracy of audio matching. That is, spatial feature extraction is performed on the target temporal feature vector to obtain the undiminished audio fingerprint corresponding to the audio to be matched, including:

[0253] The elements in the target time feature vector are mapped to values ​​within a preset range to obtain the target time value matrix;

[0254] Spatial features are extracted from the target time-value matrix to obtain the spatial sample audio fingerprint;

[0255] The elements in the spatial sample audio matrix are mapped to values ​​within a preset range to obtain the audio fingerprint before dimensionality reduction corresponding to the audio to be matched.

[0256] S504. The terminal performs dimensionality reduction processing on the audio fingerprint before dimensionality reduction through the fully connected layer in the trained audio fingerprint model to obtain the target audio fingerprint.

[0257] S505. The terminal uses the trained audio fingerprint model to map each fingerprint element in the target audio fingerprint to a target value within a preset value range, thereby obtaining the target value matrix corresponding to the target audio fingerprint. The difference between the target value and the endpoint value of the preset value range is within a preset range.

[0258] S506. The terminal performs binary mapping processing on the target numerical matrix to obtain the target binary matrix corresponding to the target audio fingerprint.

[0259] S507. The terminal searches for audio that matches the audio to be matched from the audio set based on the target binary matrix and the binary matrix corresponding to the audio fingerprint.

[0260] In this embodiment, since the trained audio fingerprint model can extract both temporal and spatial features, the audio matching accuracy of the trained audio fingerprint model in this application embodiment is higher than that of audio fingerprint extraction methods in related technologies.

[0261] Since the trained audio fingerprint model can map each fingerprint element in the target audio fingerprint to a target value within a preset value range, thus obtaining the target value matrix corresponding to the target audio fingerprint, and the difference between the target value and the endpoint value of the preset value range is within a preset range, and the target value matrix is ​​subjected to binary mapping processing to obtain the target binary matrix corresponding to the target audio fingerprint, the trained audio fingerprint model in this embodiment has a faster audio matching speed than the audio fingerprint extraction method in related technologies.

[0262] The following describes the effect of the trained audio fingerprint model in the embodiments of this application.

[0263] Ten thousand songs were randomly selected from the FMA open-source music dataset as the training set. Two thousand songs were then randomly perturbed, and these perturbed songs were used as the test set. Accuracy and Real-Time Factor (RTF) were used as evaluation metrics; higher accuracy indicates better performance, and lower RTF indicates faster matching speed. The results of testing the trained audio fingerprint model and related audio fingerprint extraction and matching methods (such as the band energy method, Landmark method, Nowplaying model, and the Sequence-to-Sequence Autoencoder Model for Audio Fingerprinting (SAMAF)) on the test set are shown in the table below.

[0264] method accuracy Real-time rate Bandwidth energy 50.25% 0.03 Landmark 53.90% 0.03 Nowplaying 65.70% 0.15 SAMAF 70.30% 0.15 Trained audio fingerprint model 80.10% 0.05

[0265] As can be seen from the table above, compared with the spectrum-based audio fingerprint extraction and matching method (which includes frequency band energy and Landmark), the accuracy and speed of the neural network model-based audio matching method (including Nowplaying, SAMAF, and trained audio fingerprint models) are significantly improved, indicating that the neural network model-based audio matching method is faster and better than the spectrum-based audio fingerprint extraction and matching method.

[0266] In the audio matching method based on neural network models, the trained audio fingerprint model in this embodiment improves accuracy by 10% and matching speed by 2 times compared to Nowplaying and SAMAF. This indicates that the trained audio fingerprint model in this embodiment performs better and is faster.

[0267] The specific implementation method and corresponding beneficial effects in this embodiment can be referred to the above audio matching method embodiment, which will not be repeated here.

[0268] To facilitate better implementation of the audio matching method provided in this application, this application also provides an apparatus based on the above-described audio matching method. The meanings of the terms used are the same as in the audio matching method described above, and specific implementation details can be found in the descriptions within the method embodiments.

[0269] For example, such as Figure 6 As shown, the audio matching device may include:

[0270] The acquisition module 601 is used to acquire the audio to be matched and the audio set, the audio set including the audio and the numerical matrix corresponding to the audio fingerprint of the audio.

[0271] The extraction module 602 is used to extract fingerprint features from the audio to be matched, and obtain the target audio fingerprint corresponding to the audio to be matched. The target audio fingerprint includes multiple fingerprint elements.

[0272] The mapping module 603 is used to map each fingerprint element in the target audio fingerprint to a target value within a preset value range, thereby obtaining a target value matrix corresponding to the target audio fingerprint. The difference between the target value and the endpoint value of the preset value range is within a preset range.

[0273] The lookup module 604 is used to find the audio that matches the audio to be matched from the audio set based on the target numerical matrix and the numerical matrix corresponding to the audio fingerprint.

[0274] Optionally, the numerical matrix corresponding to the audio fingerprint of the audio includes the binary matrix corresponding to the audio fingerprint of the audio.

[0275] Accordingly, the lookup module 604 is specifically used to perform:

[0276] The target numerical matrix is ​​subjected to binary mapping to obtain the target binary matrix corresponding to the target audio fingerprint;

[0277] Based on the target binary matrix and the binary matrix corresponding to the audio fingerprint, find the audio that matches the audio to be matched from the audio set.

[0278] Optionally, the extraction module 602 is specifically used to perform:

[0279] Using a trained audio fingerprint model, fingerprint features are extracted from the audio to be matched to obtain the target audio fingerprint corresponding to the audio to be matched.

[0280] Mapping module 603 is specifically used for execution:

[0281] Using a trained audio fingerprint model, each fingerprint element in the target audio fingerprint is mapped to a target value within a preset numerical range.

[0282] Optionally, the audio matching device further includes:

[0283] The training module is used to perform:

[0284] Obtain the training sample set of the audio fingerprint model to be trained, and extract fingerprint features from the training samples in the training sample set to obtain the sample audio fingerprints corresponding to the training samples.

[0285] Each element in the sample audio fingerprint is mapped to a sample value within a preset range to obtain the sample value matrix corresponding to the sample audio fingerprint.

[0286] The target loss value of the audio fingerprint model to be trained is determined based on the sample numerical matrix and the endpoint values ​​corresponding to the preset numerical range.

[0287] The audio fingerprint model to be trained is trained according to the target loss value so that the audio fingerprint model to be trained converges to the endpoint value, thus obtaining the trained audio fingerprint model.

[0288] Optionally, the training module is specifically used to perform:

[0289] Perform a short-time Fourier transform on the training samples in the training sample set to obtain the time-frequency feature vectors corresponding to the training samples;

[0290] Fingerprint features are extracted from the time-frequency feature vector to obtain the sample audio fingerprints corresponding to the training samples.

[0291] Optionally, fingerprint feature extraction includes temporal feature extraction and spatial feature extraction.

[0292] Accordingly, the training module is specifically used to perform:

[0293] Time features are extracted from the time-frequency feature vector to obtain the time feature vector corresponding to the training sample;

[0294] Spatial features are extracted from the temporal feature vector to obtain the sample audio fingerprint corresponding to the training sample.

[0295] Optionally, the training module is specifically used to perform:

[0296] Obtain the target audio segments;

[0297] Windowing and frame segmentation are performed on the target audio to obtain the corresponding sample speech segments;

[0298] The training sample set is determined based on the sample speech segments.

[0299] Optionally, the training module is specifically used to perform:

[0300] Temporal data augmentation is performed on the sample speech segments to obtain the positive samples corresponding to the sample speech segments;

[0301] The training sample set is determined based on the positive samples and sample speech segments corresponding to the sample speech segments.

[0302] Optionally, the training module is specifically used to perform:

[0303] The time-frequency feature vectors are augmented in the frequency domain to obtain the sample time-frequency feature vectors corresponding to the training samples.

[0304] Fingerprint features are extracted from the time-frequency feature vectors of the samples to obtain the audio fingerprints of the training samples.

[0305] Optionally, the training module is specifically used to perform:

[0306] Based on the sample audio fingerprints, determine the contrastive loss value of the audio fingerprint model to be trained;

[0307] The quantization loss value of the audio fingerprint model to be trained is determined based on the sample numerical matrix and the endpoint values ​​corresponding to the preset numerical range.

[0308] The target loss value for the audio fingerprint model to be trained is determined by comparing the loss value and quantizing the loss value.

[0309] Optionally, the training module is specifically used to perform:

[0310] The first sample audio fingerprint is selected from the sample audio fingerprints;

[0311] Based on the first sample audio fingerprint and the second sample audio fingerprint corresponding to the first sample audio fingerprint in the sample audio fingerprint, the first inner product value corresponding to the first sample audio fingerprint is determined. When the first sample audio fingerprint is the sample audio fingerprint corresponding to the first sample speech segment, the second sample audio fingerprint is the sample audio fingerprint of the first positive sample corresponding to the first sample speech segment. Alternatively, when the first sample audio fingerprint is the sample audio fingerprint corresponding to the first positive sample, the second sample audio fingerprint is the sample audio fingerprint of the first sample speech segment corresponding to the first positive sample.

[0312] Based on the first sample audio fingerprint, the second sample audio fingerprint, and the third sample audio fingerprint, determine the second inner product value corresponding to the first sample audio fingerprint. The third sample audio fingerprint is the sample audio fingerprint other than the first and second sample audio fingerprints.

[0313] The contrastive loss value of the audio fingerprint model to be trained is determined based on the first inner product value and the second inner product value.

[0314] In practice, each of the above modules can be implemented as an independent entity or can be combined arbitrarily to be implemented as the same or several entities. For the specific implementation methods and corresponding beneficial effects of each of the above modules, please refer to the previous method embodiments, which will not be repeated here.

[0315] This application also provides an electronic device, which may be a server or a terminal, etc. Figure 7 As shown, it illustrates a structural schematic diagram of the electronic device involved in the embodiments of this application, specifically:

[0316] The electronic device may include components such as a processor 701 with one or more processing cores, a memory 702 with one or more computer-readable storage media, a power supply 703, and an input unit 704. Those skilled in the art will understand that... Figure 7 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:

[0317] The processor 701 is the control center of the electronic device, connecting various parts of the device via various interfaces and lines. It executes computer programs and / or modules stored in the memory 702, and calls data stored in the memory 702 to perform various functions and process data. Optionally, the processor 701 may include one or more processing cores; preferably, the processor 701 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 701.

[0318] The memory 702 can be used to store computer programs and modules. The processor 701 executes various functional applications and data processing by running the computer programs and modules stored in the memory 702. The memory 702 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, computer programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 702 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 702 may also include a memory controller to provide the processor 701 with access to the memory 702.

[0319] The electronic device also includes a power supply 703 that supplies power to the various components. Preferably, the power supply 703 can be logically connected to the processor 701 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 703 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0320] The electronic device may also include an input unit 704, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0321] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 701 in the electronic device loads the executable files corresponding to the processes of one or more computer programs into the memory 702 according to the following instructions, and the processor 701 runs the computer programs stored in the memory 702 to realize various functions, such as:

[0322] Obtain the audio to be matched and the audio set, where the audio set includes the audio and the numerical matrix corresponding to the audio fingerprint of the audio;

[0323] Fingerprint features are extracted from the audio to be matched to obtain the target audio fingerprint, which includes multiple fingerprint elements.

[0324] Each fingerprint element in the target audio fingerprint is mapped to a target value within a preset value range to obtain the target value matrix corresponding to the target audio fingerprint. The difference between the target value and the endpoint value of the preset value range is within a preset range.

[0325] Based on the target numerical matrix and the numerical matrix corresponding to the audio fingerprint, search for audio that matches the audio to be matched from the audio set.

[0326] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by a computer program, or by a computer program controlling related hardware. The computer program can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0327] Therefore, embodiments of this application provide a computer-readable storage medium storing a computer program that can be loaded by a processor to execute the steps of any of the audio matching methods provided in embodiments of this application. For example, the computer program can execute the following steps:

[0328] Obtain the audio to be matched and the audio set, where the audio set includes the audio and the numerical matrix corresponding to the audio fingerprint of the audio;

[0329] Fingerprint features are extracted from the audio to be matched to obtain the target audio fingerprint, which includes multiple fingerprint elements.

[0330] Each fingerprint element in the target audio fingerprint is mapped to a target value within a preset value range to obtain the target value matrix corresponding to the target audio fingerprint. The difference between the target value and the endpoint value of the preset value range is within a preset range.

[0331] Based on the target numerical matrix and the numerical matrix corresponding to the audio fingerprint, search for audio that matches the audio to be matched from the audio set.

[0332] For details on the specific implementation methods and corresponding beneficial effects of the above operations, please refer to the previous embodiments, which will not be repeated here.

[0333] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0334] Since the computer program stored in the computer-readable storage medium can execute the steps of any of the audio matching methods provided in the embodiments of this application, the beneficial effects that any of the audio matching methods provided in the embodiments of this application can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.

[0335] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the aforementioned audio matching method.

[0336] The above provides a detailed description of an audio matching method, apparatus, electronic device, and computer-readable storage medium provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. An audio matching method, characterized in that, include: Obtain the audio to be matched and the audio set, wherein the audio set includes the audio and the numerical matrix corresponding to the audio fingerprint of the audio, and the numerical matrix corresponding to the audio fingerprint of the audio includes the binary matrix corresponding to the audio fingerprint of the audio; Fingerprint features are extracted from the audio to be matched to obtain the target audio fingerprint corresponding to the audio to be matched, and the target audio fingerprint includes multiple fingerprint elements; Each fingerprint element in the target audio fingerprint is mapped to a target value within a preset value range to obtain a target value matrix corresponding to the target audio fingerprint. The difference between the target value and the endpoint value of the preset value range is within a preset range. Based on the target numerical matrix and the numerical matrix corresponding to the audio fingerprint, the audio that matches the audio to be matched is searched from the audio set, including: performing binary mapping processing on the target numerical matrix to obtain the target binary matrix corresponding to the target audio fingerprint; Based on the target binary matrix and the binary matrix corresponding to the audio fingerprint, find the audio that matches the audio to be matched from the audio set.

2. The audio matching method as described in claim 1, characterized in that, The fingerprint features of the audio to be matched are extracted to obtain the target audio fingerprint corresponding to the audio to be matched; Mapping each fingerprint element in the target audio fingerprint to a target value within a preset value range includes: The fingerprint features of the audio to be matched are extracted using a trained audio fingerprint model to obtain the target audio fingerprint corresponding to the audio to be matched. The trained audio fingerprint model maps each fingerprint element in the target audio fingerprint to a target value within a preset range.

3. The audio matching method as described in claim 2, characterized in that, Before extracting fingerprint features from the audio to be matched using a trained audio fingerprint model to obtain the target audio fingerprint corresponding to the audio to be matched, the process further includes: Obtain the training sample set of the audio fingerprint model to be trained, and extract fingerprint features from the training samples in the training sample set to obtain the sample audio fingerprint corresponding to the training samples. Each element in the sample audio fingerprint is mapped to a sample value within the preset value range to obtain the sample value matrix corresponding to the sample audio fingerprint. The target loss value of the audio fingerprint model to be trained is determined based on the sample numerical matrix and the endpoint values ​​corresponding to the preset numerical interval. The audio fingerprint model to be trained is trained according to the target loss value so that the audio fingerprint model to be trained converges to the endpoint value, thereby obtaining the trained audio fingerprint model.

4. The audio matching method according to claim 3, characterized in that, The step of extracting fingerprint features from training samples in the training sample set to obtain the sample audio fingerprint corresponding to the training sample includes: Perform a short-time Fourier transform on the training samples in the training sample set to obtain the time-frequency feature vector corresponding to the training samples; Fingerprint features are extracted from the time-frequency feature vector to obtain the sample audio fingerprint corresponding to the training sample.

5. The audio matching method according to claim 4, characterized in that, The fingerprint feature extraction includes temporal feature extraction and spatial feature extraction; The step of extracting fingerprint features from the time-frequency feature vector to obtain the sample audio fingerprint corresponding to the training sample includes: Time features are extracted from the time-frequency feature vector to obtain the time feature vector corresponding to the training sample; Spatial features are extracted from the temporal feature vector to obtain the sample audio fingerprint corresponding to the training sample.

6. The audio matching method according to claim 3, characterized in that, The process of obtaining the training sample set for the audio fingerprint model to be trained includes: Obtain the target audio segments; The target audio is subjected to windowing and frame segmentation to obtain each sample speech segment corresponding to the target audio. Based on the sample speech segments, a training sample set is determined.

7. The audio matching method according to claim 6, characterized in that, The step of determining the training sample set based on the sample speech segments includes: The sample speech segment is subjected to time-domain data augmentation to obtain the positive sample corresponding to the sample speech segment; The training sample set is determined based on the positive samples corresponding to the sample speech segments and the sample speech segments themselves.

8. The audio matching method according to claim 4, characterized in that, The step of extracting fingerprint features from the time-frequency feature vector to obtain the sample audio fingerprint corresponding to the training sample includes: The time-frequency feature vector is subjected to frequency domain data augmentation to obtain the sample time-frequency feature vector corresponding to the training sample; Fingerprint features are extracted from the time-frequency feature vector of the sample to obtain the sample audio fingerprint corresponding to the training sample.

9. The audio matching method according to claim 8, characterized in that, Before determining the target loss value of the audio fingerprint model to be trained based on the endpoint values ​​corresponding to the sample numerical matrix and the preset numerical interval, the method further includes: Based on the sample audio fingerprint, determine the contrastive loss value of the audio fingerprint model to be trained; The step of determining the target loss value of the audio fingerprint model to be trained based on the sample numerical matrix and the endpoint values ​​corresponding to the preset numerical interval includes: The quantization loss value of the audio fingerprint model to be trained is determined based on the sample numerical matrix and the endpoint values ​​corresponding to the preset numerical range. The target loss value of the audio fingerprint model to be trained is determined based on the contrast loss value and the quantization loss value.

10. The audio matching method according to claim 9, characterized in that, The step of determining the contrastive loss value of the audio fingerprint model to be trained based on the sample audio fingerprint includes: The first sample audio fingerprint is selected from the sample audio fingerprints; Based on the first sample audio fingerprint and the second sample audio fingerprint corresponding to the first sample audio fingerprint in the sample audio fingerprint, a first inner product value corresponding to the first sample audio fingerprint is determined. When the first sample audio fingerprint is the sample audio fingerprint corresponding to the first sample speech segment, the second sample audio fingerprint is the sample audio fingerprint of the first positive sample corresponding to the first sample speech segment. Alternatively, when the first sample audio fingerprint is the sample audio fingerprint corresponding to the first positive sample, the second sample audio fingerprint is the sample audio fingerprint of the first sample speech segment corresponding to the first positive sample. Based on the first sample audio fingerprint, the second sample audio fingerprint, and the third sample audio fingerprint, determine the second inner product value corresponding to the first sample audio fingerprint, wherein the third sample audio fingerprint is the sample audio fingerprint other than the first sample audio fingerprint and the second sample audio fingerprint. The contrastive loss value of the audio fingerprint model to be trained is determined based on the first inner product value and the second inner product value.

11. An audio matching device, characterized in that, include: The acquisition module is used to acquire the audio to be matched and the audio set, wherein the audio set includes the audio and the numerical matrix corresponding to the audio fingerprint of the audio, and the numerical matrix corresponding to the audio fingerprint of the audio includes the binary matrix corresponding to the audio fingerprint of the audio. The extraction module is used to extract fingerprint features from the audio to be matched to obtain the target audio fingerprint corresponding to the audio to be matched, wherein the target audio fingerprint includes multiple fingerprint elements; The mapping module is used to map each fingerprint element in the target audio fingerprint to a target value within a preset value range, thereby obtaining a target value matrix corresponding to the target audio fingerprint, wherein the difference between the target value and the endpoint value of the preset value range is within a preset range. The search module is used to search for an audio that matches the audio to be matched from the audio set based on the target numerical matrix and the numerical matrix corresponding to the audio fingerprint, including: performing a binary mapping process on the target numerical matrix to obtain a target binary matrix corresponding to the target audio fingerprint; Based on the target binary matrix and the binary matrix corresponding to the audio fingerprint, find the audio that matches the audio to be matched from the audio set.

12. The apparatus according to claim 11, characterized in that, The extraction module is specifically used to perform: The fingerprint features of the audio to be matched are extracted using a trained audio fingerprint model to obtain the target audio fingerprint corresponding to the audio to be matched. The mapping module is specifically used to execute: The trained audio fingerprint model maps each fingerprint element in the target audio fingerprint to a target value within a preset range.

13. The apparatus according to claim 12, characterized in that, The audio matching device further includes: The training module is used to perform: Obtain the training sample set of the audio fingerprint model to be trained, and extract fingerprint features from the training samples in the training sample set to obtain the sample audio fingerprint corresponding to the training samples. Each element in the sample audio fingerprint is mapped to a sample value within the preset value range to obtain the sample value matrix corresponding to the sample audio fingerprint. The target loss value of the audio fingerprint model to be trained is determined based on the sample numerical matrix and the endpoint values ​​corresponding to the preset numerical interval. The audio fingerprint model to be trained is trained according to the target loss value so that the audio fingerprint model to be trained converges to the endpoint value, thereby obtaining the trained audio fingerprint model.

14. The apparatus according to claim 13, characterized in that, The training module is specifically used to perform: Perform a short-time Fourier transform on the training samples in the training sample set to obtain the time-frequency feature vector corresponding to the training samples; Fingerprint features are extracted from the time-frequency feature vector to obtain the sample audio fingerprint corresponding to the training sample.

15. The apparatus according to claim 14, characterized in that, The fingerprint feature extraction includes temporal feature extraction and spatial feature extraction; The training module is specifically used to execute: Time features are extracted from the time-frequency feature vector to obtain the time feature vector corresponding to the training sample; Spatial features are extracted from the temporal feature vector to obtain the sample audio fingerprint corresponding to the training sample.

16. The apparatus according to claim 13, characterized in that, The training module is specifically used to perform: Obtain the target audio segments; The target audio is subjected to windowing and frame segmentation to obtain each sample speech segment corresponding to the target audio. Based on the sample speech segments, a training sample set is determined.

17. The apparatus according to claim 16, characterized in that, The training module is specifically used to perform: The sample speech segment is subjected to time-domain data augmentation to obtain the positive sample corresponding to the sample speech segment; The training sample set is determined based on the positive samples corresponding to the sample speech segments and the sample speech segments themselves.

18. The apparatus according to claim 14, characterized in that, The training module is specifically used to perform: The time-frequency feature vector is subjected to frequency domain data augmentation to obtain the sample time-frequency feature vector corresponding to the training sample; Fingerprint features are extracted from the time-frequency feature vector of the sample to obtain the sample audio fingerprint corresponding to the training sample.

19. The apparatus according to claim 18, characterized in that, The training module is specifically used to perform: Based on the sample audio fingerprint, determine the contrastive loss value of the audio fingerprint model to be trained; The quantization loss value of the audio fingerprint model to be trained is determined based on the sample numerical matrix and the endpoint values ​​corresponding to the preset numerical range. The target loss value of the audio fingerprint model to be trained is determined based on the contrast loss value and the quantization loss value.

20. The apparatus according to claim 19, characterized in that, The training module is specifically used to perform: The first sample audio fingerprint is selected from the sample audio fingerprints; Based on the first sample audio fingerprint and the second sample audio fingerprint corresponding to the first sample audio fingerprint in the sample audio fingerprint, a first inner product value corresponding to the first sample audio fingerprint is determined. When the first sample audio fingerprint is the sample audio fingerprint corresponding to the first sample speech segment, the second sample audio fingerprint is the sample audio fingerprint of the first positive sample corresponding to the first sample speech segment. Alternatively, when the first sample audio fingerprint is the sample audio fingerprint corresponding to the first positive sample, the second sample audio fingerprint is the sample audio fingerprint of the first sample speech segment corresponding to the first positive sample. Based on the first sample audio fingerprint, the second sample audio fingerprint, and the third sample audio fingerprint, determine the second inner product value corresponding to the first sample audio fingerprint, wherein the third sample audio fingerprint is the sample audio fingerprint other than the first sample audio fingerprint and the second sample audio fingerprint. The contrastive loss value of the audio fingerprint model to be trained is determined based on the first inner product value and the second inner product value.

21. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a computer program, and the processor running the computer program in the memory to perform the audio matching method according to any one of claims 1-10.

22. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted for loading by a processor to perform the audio matching method according to any one of claims 1-10.

23. A computer program product, characterized in that, The computer program product stores a computer program adapted for loading by a processor to execute the audio matching method according to any one of claims 1-10.