Audio Recognition Method, Apparatus, Computer Device, and Storage Medium
By analyzing the playback timing of audio clips through timing correlation, the problem of low audio recognition accuracy in the prior art is solved, and more accurate audio file repetition recognition is achieved, and the recognition effect is maintained in a noisy environment.
Patent Information
- Application Number
- CN202110827135.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-21
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-07-21
AI Technical Summary
In the prior art, the audio recognition method has low accuracy and is susceptible to noise, slight distortion and voice change processing, resulting in identification errors.
By performing timing analysis of the first audio file and the second audio file based on the playback timing of the audio clip pair of feature matching, the correlation information is obtained, which is used to indicate the degree of correlation between the feature change trends between the two, thereby determining whether the audio file is repeated.
It improves the accuracy of audio recognition, can more accurately define the timing relationship between audio files, clearly represent the continuous characteristics of the audio, and maintain the recognition effect in a noisy environment.
Smart Images

Figure CN113823320B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical fields of artificial intelligence, multimedia data, audio processing, etc. The present application relates to an audio recognition method, device, computer device and storage medium. Background Art
[0002] With the development of network technology, audio recognition technology is widely used. In many scenarios, it is necessary to recognize duplicate audio. For example, audio copyright protection, information flow recommendation, etc.
[0003] In the related art, the audio recognition method may include: the terminal searches for segments with similar audio features in two audio files based on the audio features of the two audio files, and determines whether the two audio files are duplicates based on the proportion of the similar segments in the audio file. If the proportion exceeds the threshold, it is determined that the two audio files are duplicates; otherwise, they are not duplicates; for example, if 90% of the segments in Audio A are repeated with Audio B, then Audio A and Audio B are repeated.
[0004] The above process performs recognition based on the proportion of similar segments. However, the proportion of similar segments is extremely easy to change. For example, adding noise, slightly distorting, or changing the voice in the audio makes it difficult to detect the segments with similar audio features. Even if the threshold is adjusted appropriately, it is very easy to make misidentifications, resulting in a low accuracy rate of audio recognition. Summary of the Invention
[0005] The present application provides an audio recognition method, device, computer device and storage medium, which can solve the problem of low accuracy rate of audio recognition in the related art. The technical solutions are as follows:
[0006] On the one hand, an audio recognition method is provided, and the method includes:
[0007] Determine audio segment pairs with matching features in the first audio file and the second audio file based on the audio features of the first audio file and the second audio file;
[0008] Perform timing correlation analysis on the first audio file and the second audio file based on the playback timings of the audio segment pairs to obtain correlation information, where the correlation information is used to indicate the correlation degree between the feature change trend of the first audio file based on the playback timing and the feature change trend of the second audio file based on the playback timing;
[0009] Determine the recognition results of the first audio file and the second audio file based on the correlation information, where the recognition results are at least used to indicate whether there are duplicates between the first audio file and the second audio file.
[0010] On the other hand, an audio recognition device is provided, and the device includes:
[0011] A determination module, configured to determine pairs of audio segments with matching features in the first audio file and the second audio file based on the audio features of the first audio file and the audio features of the second audio file;
[0012] An analysis module, configured to perform timing correlation analysis on the first audio file and the second audio file based on the playing timings of the pairs of audio segments, to obtain correlation information, where the correlation information is used to indicate the degree of correlation between the feature change trend of the first audio file based on the playing timing and the feature change trend of the second audio file based on the playing timing;
[0013] An identification module, configured to determine identification results of the first audio file and the second audio file based on the correlation information, where the identification results are at least used to indicate whether the first audio file and the second audio file are duplicates.
[0014] On the other hand, a computer device is provided, where the computer device includes:
[0015] One or more processors;
[0016] A memory;
[0017] One or more computer programs, where the one or more computer programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs are configured to: execute the above audio recognition method.
[0018] On the other hand, a computer-readable storage medium is provided, where the computer storage medium is used to store computer instructions, and when the computer instructions are run on a computer device, the computer device can execute the above audio recognition method.
[0019] On the other hand, a computer program is provided, where the computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium, a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above audio recognition method.
[0020] The beneficial effects brought by the technical solution provided in this application are:
[0021] By performing timing correlation analysis on the first audio file and the second audio file based on the playing timing of the audio segment pairs matched by features, correlation information is obtained, and this correlation information can indicate the degree of correlation between the feature change trends of the two audio files based on the playing timing respectively, so as to accurately define the relationship in timing between the two audio files. Further, the characteristics of the continuity of the audio are clearly represented by the feature change trends in timing; based on the correlation information between the feature change trends, it is possible to accurately identify whether there is duplication between the two audio files, improving the accuracy of audio recognition. Description of the Drawings
[0022] To more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for describing the embodiments of the present application will be briefly introduced below.
[0023] Figure 1 Schematic diagram of the implementation environment of an audio recognition method provided by the present application;
[0024] Figure 2 Schematic flowchart of an audio recognition method provided by an embodiment of the present application;
[0025] Figure 3 Schematic flowchart of a method for extracting audio features provided by an embodiment of the present application;
[0026] Figure 4 Schematic diagram of the coordinate system of a regression analysis provided by an embodiment of the present application;
[0027] Figure 5 Schematic diagram of the visualization result of performing regression analysis using audio fingerprint features and obtaining regression coefficients provided by an embodiment of the present application;
[0028] Figure 6 Schematic diagram of the position of the timing pair obtained using MFCC features in the coordinate system provided by an embodiment of the present application;
[0029] Figure 7 Schematic diagram of the visualization result of a regression analysis provided by an embodiment of the present application;
[0030] Figure 8 Schematic diagram of the structure of a server provided by an embodiment of the present application;
[0031] Figure 9 Schematic diagram of an audio compilation provided by an embodiment of the present application;
[0032] Figure 10 Schematic diagram of an image in a video file provided by an embodiment of the present application;
[0033] Figure 11Structural schematic diagram of an audio recognition device provided by an embodiment of the present application;
[0034] Figure 12 Structural schematic diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0035] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals indicate the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be construed as a limitation to the present invention.
[0036] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application means the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.
[0037] To make the objectives, technical solutions, and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.
[0038] Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.
[0039] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields involved, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0040] Computer Vision Technology (CV) Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to using cameras and computers to replace human eyes for tasks such as object recognition and measurement in machine vision, and further performing graphic processing to make the computer-processed images more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0041] The key technologies of speech technology include automatic speech recognition technology (ASR), text-to-speech technology (TTS), and voiceprint recognition technology. Enabling computers to listen, see, speak, and feel is the future development direction of human-computer interaction, and among them, speech has become one of the most promising human-computer interaction methods in the future.
[0042] Machine Learning (ML) is an interdisciplinary subject across multiple fields, involving multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills, and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0043] The audio recognition method provided by this application relates to the above-mentioned artificial intelligence technology. Exemplarily, the big data processing technology in artificial intelligence technology can be used to process a large amount of audio data to identify whether there are duplicates between different audios. Of course, the above-mentioned machine learning technology can also be used to perform reinforcement learning on the above audio recognition process to identify a large number of audio files. For example, the audio data processing results can be input into a neural network, and the neural network outputs the recognition results. Exemplarily, the audio recognition method of this application can be applied in the field of computer vision technology. For example, when performing video recognition, the audio recognition method provided by this application can be used to recognize the audio in the video, and the audio recognition results are input into a machine learning model to identify duplicate videos.
[0044] Figure 1 It is a schematic diagram of the implementation environment of an audio recognition method provided by this application. This implementation environment includes a computer device, which can recognize an audio file based on the audio characteristics of the audio file. The computer device can be a server or any processor, data processing device, etc. with data processing capabilities to recognize audio. Exemplarily, the following will take a server as an example for illustration. As Figure 1 shown, this implementation environment includes a server 101, and the server 101 can be the background server of an application. Exemplarily, the server 101 can use the audio recognition method of this application to identify duplicate audio files among multiple audio files. Or, the server 101 can also recognize the audio included in a video file and identify video files with duplicate corresponding audio among multiple video files.
[0045] Exemplarily, the application can include an information interaction platform that provides an information stream to users. The content of the information stream can include audio, video, etc., and of course, it can also include text, images, animations, etc. The server 101 performs processing such as classification, screening, and deduplication on the information stream based on the audio recognition results; or, the server 101 can also perform personalized push of information streams for different users or different user groups based on the audio recognition results and the processed information stream.
[0046] In a possible application scenario, the server 101 can identify the audio in an audio file or a video file and process it based on the identification result. Exemplarily, taking a video file as an example, the server 101 combines the audio and the image in the video file to identify the video file, so as to improve the accuracy of identification. For example, the server 101 can identify the audio files in multiple video files and identify the images included in the video files, and combine the identification results of the audio files and the identification results of the images to determine that there is a duplication between two video files with duplicate images and duplicate audio files. The server 101 can perform operations such as deleting or marking the duplicate video files. For example, for a self-created video file newly uploaded by a user account, the server 101 can identify the newly uploaded video file based on the existing video files. When the newly uploaded video file is duplicate with the existing video files, the server 101 can reject this upload and prompt the user account that the video is duplicate.
[0047] In a possible application scenario, the implementation environment may further include a terminal 102. The terminal 102 is installed with an application program, and the terminal 102 and the server 101 can perform data interaction based on the application program. The server 101 can push an audio file or a video file to the user based on the identification result and the data interaction with the user's terminal 102. Exemplarily, the server 101 can aggregate multiple video files with the same background music into a video set based on the identification result of the background music of the video file, and push the video set with the same background music to the user.
[0048] It should be noted that the audio content carried by the audio file can be user voice, background music, user original commentary, songs, etc.; of course, it can also be various types of sound effects such as game sound effects, transition sound effects, and funny sound effects. The application program can be any application with audio recognition requirements such as a content interaction application, a video application, a news and information application, an audio playback application, a social application, an information browsing application, and a game application. Among them, an information stream refers to a content organization form arranged vertically according to a feature specification style. From the perspective of display sorting, common ones include chronological order, popularity, algorithm sorting, etc. For example, short videos arranged by popularity shown on the user's home page of a content interaction application. The server 101 can be a single server or a server cluster integrated by multiple servers.
[0049] The technical solution of the present application will be described in detail below with specific embodiments.
[0050] Figure 2The figure is a flowchart of an audio recognition method provided in an embodiment of this application. This method is executed by a computer device, which can be any device such as a server or a processor or data processing device capable of audio recognition. In the embodiment of this application, a server is taken as an example for introduction, but the execution entity of this method is not specifically limited. As Figure 2 shown, the method includes the following steps:
[0051] Step 201, the server determines the audio features of the first audio file and the audio features of the second audio file.
[0052] The server can obtain the first audio file and the second audio file to be recognized, and extract the audio features of the first audio file and the audio features of the second audio file. Exemplarily, the audio features can be in the form of feature vectors. Then the server can extract the feature vectors of each audio segment in the first audio file and the feature vectors of each audio segment in the second audio file.
[0053] In a possible example, the audio features can be MFCC (Mel Frequency Cepstrum Coefficient); since the Mel frequency is proposed based on the human ear's auditory characteristics, it has a non-linear correspondence with the Hz frequency. MFCC is a frequency feature calculated using the non-linear correspondence between the Mel frequency and the Hz frequency. MFCC has been widely applied in technical fields such as speech recognition and audio processing. In another possible example, the audio features can also be audio fingerprint features, which can be obtained based on the Chromaprint algorithm (audio fingerprint algorithm). That is, the server can use the Chromaprint algorithm to extract the audio fingerprint features of audio segments from the audio file. The Chromaprint algorithm can also extract fixed features from the audio, and the similarity between different features can be calculated based on the Hamming distance. Of course, the audio features can also be any feature that can characterize audio characteristics such as FBank features (Filter bank), sound intensity, frequency, signal-to-noise ratio, etc. The embodiment of this application does not specifically limit the types of the audio features.
[0054] The server can collect the feature vectors of multiple audio segments to obtain a feature vector set of the audio file. Exemplarily, taking the first audio file as an example, the server can extract n feature vectors from n audio segments of the first audio file. The audio features of the multiple audio segments of the first audio file can be represented as a feature vector set composed of n feature vectors. The feature vector set X can be expressed as: X = [x1, x2, …, x i , …, x n , where n is a positive integer greater than 0. xi is the feature vector of the i-th audio segment among n audio segments. In a possible example, taking MFCC features as an example, the server can use an 8-dimensional vector to represent the MFCC feature vector of each audio segment; then the feature vector set X can be an n * 8-dimensional feature vector set, where x i can be an 8-dimensional vector, and this 8-dimensional vector can include 8 binary data, and the corresponding value range is: x i ∈[0, 255]. For example, the 8-dimensional vector can be (0, 1, 1, 0, 1, 1, 1, 1). In another possible example, taking audio fingerprint features as an example, the server can use a 32-dimensional vector to represent the audio fingerprint feature vector of each audio segment; then the feature vector set X can be an n * 32-dimensional feature vector set, where x i can be a 32-dimensional vector, and this 32-dimensional vector can include 32 binary data, and the corresponding value range can be: x i ∈[-2 31 , 2 31 -1]. Among them, the above is only illustrated by taking the first audio file as an example, and the feature representation form of the second audio file is the same as that of the first audio file, and will not be elaborated here one by one.
[0055] Exemplarily, the server can extract the audio features of each audio segment based on the sampling duration. The sampling duration refers to the duration of each audio segment, and the sampling duration can be set as needed. The embodiments of the present application do not make specific limitations on this. For example, the sampling duration can be 100 ms, that is, the duration of each audio segment can be 100 ms. Of course, the server can also extract audio features by combining a certain sampling frequency and sampling duration. For example, the sampling frequency can be 100 times per second, and the corresponding playing durations between adjacent audio segments can overlap or not overlap. Exemplarily, the server can obtain the corresponding audio file based on the link address of the audio file. The link address can be a local storage path or a network address. The server can obtain the first audio file and the second audio file from the local storage space, or the server can also obtain the first audio file and the second audio file from other devices based on the network address. In a possible example, the server can first detect whether the audio features of the first audio file and the second audio file have been extracted. Exemplarily, taking the first audio file as an example, the server can obtain the file identifier of the first audio file based on the link address and detect whether the audio features of the first audio file corresponding to the audio identifier have been extracted. If they have been extracted, there is no need to extract, and directly obtain the extracted audio features and execute step 202. If not, obtain the first audio file based on the link address and extract the audio features from the first audio file.
[0056] In a possible implementation, a part of the audio in the audio file can be focused on. Then, the server can intercept a part of the audio in the audio file for feature extraction. For example, the server can extract the features of the beginning and end of a segment of audio; alternatively, the server can also crop two segments of audio with the same playback duration from two audio files for feature extraction. Correspondingly, this step can include the following two methods.
[0057] The first method: The server determines the target segment of the first audio file and the target segment of the second audio file, and extracts the audio features of the target segment of the first audio file and the audio features of the target segment of the second audio file respectively.
[0058] The target segment is an audio segment in the corresponding audio file that is before the first time sequence threshold or after the second time sequence threshold, and the first time sequence threshold is greater than the second time sequence threshold; Exemplarily, taking the first audio file as an example, the server can obtain the beginning segment before the first time sequence threshold and the ending segment after the second time sequence threshold. The first time sequence threshold and the second time sequence threshold can be set according to needs. For example, the playback duration can be 20 seconds, the first time sequence threshold can be 5, and the second time sequence threshold can be 15. The beginning audio of the background sound and the ending audio of the background sound that are repeated can be extracted from the audio with a playback duration of 20 seconds. For example, the server identifies the audio segments at the beginning and end, and filters out the video files with repeated background sound at the beginning and end to push to the user. For example, in some lectures and explanatory videos, the video creator will use a consistent beginning and end. Then, lectures and explanatory videos of the same video creator can be recommended to the user.
[0059] The second method: The server crops out the first part from the first audio file respectively, and crops out the second part with the same playback duration as the first part from the second audio file; The server extracts the audio features of the first part and the audio features of the second part respectively.
[0060] The first part and the second part are obtained by cropping the first audio file and the second audio file respectively based on the target playback duration.
[0061] When the server detects that the first audio file and the second audio file meet the cropping conditions, the server can crop the first audio file and the second audio file based on the second method. The cropping conditions may include, but are not limited to: the playing durations of the first audio file and the second audio file are different, the playing duration of the first audio file or the second audio file exceeds the target playing duration, etc. Exemplarily, the server can crop the first part with the target playing duration from the first audio file and the second part with the target playing duration from the second audio file based on the target playing duration. For example, the server can also crop the first part with the target playing duration starting from the starting cropping position of the first audio file and the second part with the target playing duration starting from the starting cropping position of the second audio file based on the starting cropping position and the target playing duration. For example, the playing duration can be 8 minutes, and the server can crop an audio part with a duration of 6 minutes starting from the 5th second along the playing time sequence.
[0062] Exemplarily, the server may include an audio feature generation module, and step 201 may be executed by the audio feature generation module. Figure 3 A flowchart for extracting audio features provided by this application is as Figure 3 shown. The audio feature generation module can obtain the file identifiers of the first audio file and the second audio file based on the input audio file link, and detect whether the audio features of the first audio file and the second audio file have been extracted based on the file identifiers. If they have been extracted, the extracted audio features can be directly obtained and the process ends. If not, the first audio file and the second audio file are downloaded based on the audio file link; the server detects whether the first audio file and the second audio file meet the cropping conditions. If they meet the cropping conditions, the first audio file and the second audio file are cropped; if they do not meet the cropping conditions, the audio features are extracted using a specified method. For example, the audio fingerprint features are extracted using the Chromaprint algorithm, and the process ends.
[0063] Step 202: The server determines the pairs of audio segments with matching features in the first audio file and the second audio file based on the audio features of the first audio file and the audio features of the second audio file.
[0064] The audio clip pair includes a first clip of the first audio file and a second clip of the second audio file; the server can perform feature matching on the audio clips between the two audio files respectively based on the audio features of each audio clip in the audio files, obtain multiple audio clip pairs with feature matching, and acquire the playing timings of the audio clip pairs with feature matching; each audio clip pair includes a first clip and a second clip that matches the first clip. In a possible implementation, the server can first perform feature matching on the two audio files and obtain the corresponding timing pairs based on the matching audio clip pairs. Correspondingly, this step may include step 2021 and step 2022 below.
[0065] Step 2021: The server performs feature matching on the audio clips in the first audio file and the audio clips in the second audio file based on the audio features of the first audio file and the audio features of the second audio file, and obtains at least two first clips and the second clips respectively corresponding to them.
[0066] For each audio clip in the first audio file, the server can perform feature matching on each audio clip in the first audio file with each audio clip in the second audio file respectively based on the audio feature of the audio clip, and obtain the matching results between each audio clip in the first audio file and each audio clip in the second audio file; for each matching result, the server can determine the first clip in the first audio file and the second clip in the second audio file corresponding to the matching result that meets the matching condition as an audio clip pair with feature matching.
[0067] In a possible implementation, the server can perform feature matching based on feature vectors. The process may include: the server determines the feature distances between the feature vectors of each audio clip in the first audio file and the feature vectors of each audio clip in the second audio file, and determines the first clip and the second clip corresponding to it whose feature distances are within the second threshold range as an audio clip pair. The server can calculate the vector distances between each feature vector in the feature vector set of the first audio file and each feature vector in the feature vector set of the second audio file. Exemplarily, the feature distance can be Euclidean distance, Mahalanobis distance, Manhattan distance, etc. The second threshold range can be set as needed; the second threshold range can be a range less than a preset threshold or a range greater than the first threshold and less than the second threshold, etc., where the first threshold is less than the second threshold; for example, less than 8, or greater than 0 and less than 6, etc. The embodiments of the present application do not make specific limitations on the manifestation form of the feature distance and the second threshold range.
[0068] Exemplarily, the server can extract m audio clips from the first audio file, then the feature vector set of the first audio file can be expressed as X = [x1, x2, …, x i,…,x m , where m is a positive integer greater than 1, and x i is the feature vector of the i-th audio segment among m audio segments, and i is a positive integer greater than 0. The server can extract n audio segments from the second audio file, and the feature vector set of the second audio file can be expressed as Y = [y1, y2, …, y j ,…,y n , where n is a positive integer greater than 1, and y j is the feature vector of the j-th audio segment among n audio segments, and j is a positive integer greater than 0. The server calculates a distance matrix D including m * n feature distances, and the distance matrix D can be expressed as: Among them, the distance matrix D can be a matrix with m rows and n columns, and d mn represents the vector distance between the m-th feature vector in X and the n-th feature vector in Y. That is, in the distance matrix D, d ij represents the vector distance between the i-th feature vector in X and the j-th feature vector in Y. The server filters the distance matrix based on the second threshold range. For example, k vector distances less than the preset threshold can be filtered out from the distance matrix to obtain k pairs of audio segments corresponding to the k vector distances.
[0069] In a possible implementation, the server can, based on the two methods in step 201 above, intercept part of the audio in the audio file for feature extraction. Correspondingly, the server can obtain pairs of audio segments by using the following Method 1 or Method 2.
[0070] Method 1: The server determines pairs of audio segments with matching features in the target segment of the first audio file and the target segment of the second audio file.
[0071] Exemplarily, to distinguish the target segments of the two audio files, the target segment of the first audio file is hereinafter referred to as the first target segment, and correspondingly, the target segment of the second audio file is referred to as the second target segment. The server can perform feature matching between the audio features of each audio segment in the first target segment and the audio features of each audio segment in the second target segment to obtain pairs of audio segments with matching features in the first target segment and the second target segment. Each pair of audio segments includes the first segment of the first target segment and the second segment of the second target segment.
[0072] Method 2: The server determines pairs of audio segments with matching features in the first part of the first audio file and the second part of the second audio file.
[0073] The server can perform feature matching on the audio features of each audio segment in the first part with the audio features of each audio segment in the second part to obtain audio segment pairs with matching features in the first part and the second part. Each audio segment pair includes a first segment in the first part and a second segment in the second part.
[0074] Step 2022: For each audio segment pair, the server obtains the playback timing of the first segment of the audio segment pair and the playback timing of the corresponding second segment.
[0075] For each audio segment pair, a timing pair of the audio segment pair is composed of the playback timing of the first segment and the playback timing of the second segment of the audio segment pair.
[0076] The playback timing of an audio segment is used to indicate: the order of the playback time corresponding to the audio segment in the total playback duration of the audio file among multiple audio segments. For example, from a 30 - second audio file, 3 audio segments, namely segment a, segment b, and segment c, are extracted. The corresponding playback times of segment a, segment b, and segment c are from the 1st second to the 12th second, from the 10th second to the 22nd second, and from the 20th second to the 30th second respectively. Then the playback timings of segment a, segment b, and segment c are 1, 2, and 3 respectively. Of course, if the server samples multiple audio segments in sequence according to the playback timing of the audio file, the playback timing of the audio segment is the same as the numerical value of the sampling order of the audio segment.
[0077] In this step, the server composes the playback timing of the first segment included in each audio segment pair and the playback timing of the second segment pair into a timing pair corresponding to the audio segment pair. Exemplarily, the k timing pairs corresponding to the k audio segment pairs are represented as a point - pair set M, and the point - pair set M is: M = {[i,j]} k , where k is a positive integer greater than 0; in the timing pair [i,j], i represents the playback timing of the first segment, j represents the playback timing of the second segment that is feature - matched with the first segment, and both i and j are positive integers greater than 0.
[0078] In this step, by respectively performing feature matching on the audio segments based on the audio features of the two audio files, audio segment pairs with the same or similar features can be accurately obtained, improving the accuracy of subsequent processing. Moreover, by respectively performing feature matching on the audio features of each audio segment in the first audio file with the audio features of each audio segment in the second audio file, audio segment pairs in the two audio files can be matched one by one, exhausting the possibilities of audio segment pairs with feature matching as much as possible, further ensuring the accuracy and reliability of feature matching.
[0079] Moreover, it is also possible to intercept a part of the audio in the audio file for feature matching; for example, the beginning and ending parts, or two parts with the same playing duration in two audio files, etc., to meet various requirements, so that the audio recognition method of the present application can be applicable to various scenarios, thereby improving the applicability of the audio recognition method of the present application.
[0080] Step 203: The server performs timing correlation analysis on the first audio file and the second audio file based on the playing timing of the audio segment pair to obtain correlation information.
[0081] The correlation information is used to indicate the degree of correlation between the feature change trend of the first audio file based on the playing timing and the feature change trend of the second audio file based on the playing timing. The server can analyze the playing timing between two segments with feature matching according to the playing timing of the first segment and the playing timing of the second segment that matches the features of the first segment to obtain the correlation relationship.
[0082] In a possible implementation manner, the server can analyze the correlation between the feature change trends of the first audio file and the second audio file based on the change trend of the playing timing between the first segment and the second segment with feature matching. Then this step may include: The server analyzes the correlation information of the change trend of the playing timing of the first segment and the playing timing of the second segment in each audio segment pair based on the playing timing of the audio segment pair. Exemplarily, the correlation information reflects the correlation relationship between the change trends of multiple first segments and multiple second segments in terms of playing timing.
[0083] In a possible implementation manner, it is possible to focus on the degree of correlation between the change trends of the playing timing of multiple first segments and the change trends of the playing timing of multiple second segments, that is, the correlation information may include the degree of correlation between the two change trends; in another possible implementation manner, it is also possible to focus on the offset degree between the change trends of the playing timing of multiple first segments and the change trends of the playing timing of multiple second segments, then the correlation information may include the degree of correlation and the offset degree.
[0084] The server can use statistical analysis methods for analysis. In a possible example, the server can use regression analysis, and the correlation information may include the regression coefficient obtained from the regression analysis. In another possible example, the correlation information may include the regression coefficient of the regression analysis and the timing offset. Correspondingly, this step may include the following two methods.
[0085] The first method: The associated information includes a regression coefficient; this step includes: The server performs a regression analysis on the playing timings of the first segment and the second segment in the at least two timing pairs based on the at least two timing pairs of at least two audio segment pairs, and determines the regression coefficient obtained from the regression analysis as the associated information.
[0086] Among them, the regression coefficient is used to indicate the variation relationship of the playing timing of the second segment with respect to the playing timing of the first segment. The timing pair of each audio segment pair includes the playing order of the first segment and the playing timing of the second segment of the audio segment. Each timing pair represents the relationship between the two segments included in the corresponding audio segment pair in terms of playing timing; then multiple timing pairs can represent the relationship between the variation trends of the audio segments from two audio files in multiple audio segment pairs in terms of playing timing.
[0087] Exemplarily, the server can perform a regression analysis on multiple timing pairs. The server can use the playing timing of the first audio file as the X coordinate axis and the playing timing of the second audio file as the Y coordinate axis to establish a two-dimensional rectangular coordinate system; if the regression analysis result of the multiple timing pairs shows that the audio segments of the two audio files are linearly related in terms of playing timing, that is, it is a straight line in the rectangular coordinate system, then the regression coefficient obtained from the linear regression analysis is the slope of the straight line in the rectangular coordinate system.
[0088] For example, the server represents the k timing pairs corresponding to k audio segment pairs as a point pair set M = {[i, j]} k , and the server can perform a linear regression analysis on the k point pairs [i, j] in the point pair set M to obtain a regression coefficient. The regression coefficient represents the variation relationship between i and j in the k point pairs. For example, there is a linear relationship between i and j, that is, j changes linearly with the change of i.
[0089] Figure 4 It is a schematic diagram of the coordinate system for a regression analysis provided by this application. As Figure 4 shown, Figure 4 In the coordinate systems in the left and right figures in, both use the playing timing of the first audio file as the X coordinate axis and the playing timing of the second audio file as the Y coordinate axis to establish a two-dimensional rectangular coordinate system. Figure 4 In the left figure in, it is the corresponding positions of the k point pairs [i, j] in the coordinate system before the regression analysis. The right figure is the result obtained by performing a regression analysis on the k point pairs [i, j]. As Figure 4 shown in the right figure in, there is a linear relationship with a slope of 1 between i and j.
[0090] The second way: the association information includes a regression coefficient and a timing offset; this step includes: the server performs a regression analysis on the playing timings of the first segment and the second segment in the at least two timing pairs based on the at least two timing pairs of at least two audio segment pairs, to obtain a regression coefficient and a timing offset; the server determines the regression coefficient and the timing offset as the association information.
[0091] Wherein, the timing offset is used to indicate the offset distance of the playing timings between the repeated segments in the first audio file and the second audio file. Multiple timing pairs can represent the relationship between the changing trends of the audio segments from two audio files in multiple audio segment pairs in terms of playing timings. The server can establish a two-dimensional rectangular coordinate system in the same way as in the first way of step 202. If the regression analysis result of the multiple timing pairs shows that the audio segments of the two audio files are linearly related in terms of playing timings, that is, it is a straight line in the rectangular coordinate system, the regression coefficient is the slope of the straight line in the rectangular coordinate system, and the timing offset can be the intercept of the straight line. The intercept can be the Y-axis coordinate of the intersection point between the straight line and the Y-axis; for example, if the intercept is (0, 20), it can indicate that the second audio file starts from the audio segment with a playing timing of 20 and has repetition with the audio segment in the first audio file. That is, the offset distance of the playing timings between the repeated segments in the two audio files is 0; correspondingly, the audio segment with a playing timing of 20 in the second audio file repeats with the audio segment with a playing timing of 0 in the first audio file; the audio segment with a playing timing of 21 in the second audio file repeats with the audio segment with a playing timing of 1 in the first audio file; and so on.
[0092] It should be noted that the server can use a target regression algorithm to perform a regression analysis on the playing timings of the first segment and the second segment in the at least two timing pairs. Exemplarily, the target regression algorithm can be the RANSAC (Random Sample Consensus) algorithm. The RANSAC algorithm is a supervised algorithm with high robustness. Based on the RANSAC algorithm, the server can find the correct matching rule in a large amount of noise data and obtain a relatively accurate analysis result, that is, the regression coefficient, or the regression coefficient and the timing offset. Of course, the server can also use other algorithms for regression analysis, and the embodiments of this application do not make specific limitations in this regard. For example, the target regression algorithm can also be a robust regression method such as the Huber regression algorithm or the Theil Sen regression algorithm.
[0093] In this step, by performing a temporal correlation analysis on two audio files based on the playback timings of audio segments, the correlation degree between the characteristic change trends of the two audio files based on the playback timings is obtained, and the characteristic change situations of the two audio files are found in terms of timing, thereby determining whether there is a correlation between the two audio files in terms of timing, and improving the accuracy of audio recognition.
[0094] Moreover, by analyzing the correlation information of the change trends of the playback timings of the first segment and the second segment, the correlation degree between the characteristic change trends of the two audio files is transformed into the correlation relationship between the change trends of the timings of the feature-matching segments, so that the correlation between the characteristic change trends of the audio files can be accurately represented as the correlation relationship between the change trends of the timings, improving the accuracy of the correlation degree between the characteristic change trends and further ensuring the high accuracy of audio recognition.
[0095] Moreover, through regression analysis, information such as regression coefficients or regression coefficients and timing offsets is determined, so that the correlation relationship presented in terms of timing for segments with the same or similar characteristics in the two audio files can be accurately measured by the numerical values of the regression coefficients and timing offsets; and regression analysis can accurately find the correct matching rules in a large amount of noise data, ensuring the accuracy and precision of the determined correlation information, thereby further improving the accuracy of audio recognition.
[0096] Moreover, regression analysis is performed on the timing pairs, and only by using audio features to find the audio segment pairs with feature matching is required, thereby reducing the over-reliance of audio recognition on features, reducing the requirements for feature generalization ability, and accurate recognition results can be obtained regardless of what specific audio features the features are, the types and quantities of the features, reducing the requirements for feature design, further expanding the range of available features for audio recognition, and improving the practicality and reliability of audio recognition.
[0097] Step 204: The server determines the recognition results of the first audio file and the second audio file based on this correlation information.
[0098] The recognition result is at least used to indicate whether there is a duplicate between the first audio file and the second audio file. When the degree of association indicated by the association information exceeds the target degree threshold, the server determines that there is a duplicate between the first audio file and the second audio file. When the degree of association exceeds the target degree threshold, it means that the degree of association between the characteristic change trends of the two audio files based on the playback time sequence is relatively high. That is, the characteristic change of one audio file with the playback time sequence is relatively close to or the same as that of the other audio file, indicating that the two audio files are duplicates. The target degree threshold can be set as needed, and the embodiments of the present application do not make specific limitations in this regard. For example, the target degree threshold can be 0.8, 90%, etc.
[0099] In a possible implementation manner, the association information includes a regression coefficient. When the regression coefficient is within the first threshold range, it is determined that there is a duplicate between the first audio file and the second audio file. The first threshold range can be a numerical range close to 1. For example, the first threshold range is greater than 0.9 and less than 1.1. When the regression coefficient is close to 1 or equal to 1, it means that the time sequence change trend between the playback time sequence of the first segment and the playback time sequence of the second segment is the same or similar. That is, whenever the playback time sequence of the first segment changes by one time sequence unit, the playback time sequence of the second segment that matches the characteristics of the first segment also changes by one time sequence unit, indicating that the characteristic change trends between the two audio files are basically the same and can characterize the existence of a duplicate between the two audio files. Of course, the first threshold range can be set as needed. For example, if other regression algorithms are used, the corresponding first threshold range can be set based on other regression algorithms. In a possible example, the server can use a classifier to obtain the recognition result. For example, the server pre-trains a classifier based on the regression coefficient and the recognition label. Then the server can input the regression coefficient obtained from the current regression analysis into the classifier, and the classifier outputs the recognition result based on the trained classification logic. For example, it can output whether there is a duplicate. Of course, if the RANSAC algorithm is used, the RANSAC regression analysis result can be directly input into the classifier. If other regression algorithms are used, the regression analysis result of the corresponding algorithm can be used as the input of the classifier to determine whether there is a duplicate in the audio.
[0100] In another possible implementation, the recognition result is further used to indicate the starting playback positions corresponding to the duplicate segments in the first audio file and the second audio file; the association information includes a regression coefficient and a timing offset. When the regression coefficient is within a first threshold range, the server determines that there are duplicates in the first audio file and the second audio file; the server obtains the starting playback positions corresponding to the duplicate segments in the first audio file and the second audio file based on the timing offset. The starting playback position indicates when the first audio file and the second audio file start to have duplicates respectively. For example, if the timing offset is 20 and the intercept is (0, 20), it is determined that the first audio file starts from playback time 0, that is, from the first audio segment, and the second audio file starts from the 20th audio segment, and there are duplicates.
[0101] It should be noted that the server can perform audio recognition using one type of audio feature to obtain a recognition result. Alternatively, it can also perform audio recognition using multiple types respectively, and combine the multiple recognition results corresponding to the multiple audio features to comprehensively determine the final recognition result. Exemplarily, the server can use MFCC features to recognize two audio files and determine the recognition results of the two audio files. Alternatively, the server can also use MFCC features to determine a first recognition result, use audio fingerprint features to determine a second recognition result, and combine the two recognition results to determine whether there are duplicates in the first audio file and the second audio file; for example, when both the first recognition result and the second recognition result indicate that there are duplicates in the first audio file and the second audio file, the server determines that there are duplicates in the first audio file and the second audio file. Or, the server can also determine multiple MFCC feature audio segment pairs of the first audio file and the second audio file based on MFCC features, determine multiple audio fingerprint feature audio segment pairs of the first audio file and the second audio file based on audio fingerprint features, and perform regression analysis based on the timing pairs of the MFCC feature audio segment pairs and the timing pairs of the audio fingerprint feature audio segment pairs to obtain a regression coefficient, thereby obtaining a more reliable recognition result.
[0102] Figure 5 FIG. is a schematic diagram of the visualization result of performing regression analysis using audio fingerprint features and obtaining a regression coefficient. Figure 5 The left figure in shows the positions of multiple timing pairs in a coordinate system, and the right figure shows the result obtained by performing regression analysis on the multiple timing pairs using the RANSAC algorithm. As Figure 5 shown, using audio fingerprint features, multiple timing pairs with less noise can be obtained, that is, the multiple timing pairs are basically located around or on the straight line obtained by regression analysis; the regression coefficient of the timing change trend obtained using the RANSAC algorithm is close to 1, and there is an obvious timing match, indicating that there are duplicates between the two audio files.
[0103] Figure 6 Schematic diagram of the position of the time series pair determined using MFCC features in the coordinate system. As Figure 6 shown Figure 6 In the left figure, it is the time series pair corresponding to the audio segment pair of the feature matching of two repeated audio files. At this time, regression analysis has not been performed yet, but it can also be clearly observed that there is a straight line formed by multiple time series pairs in this coordinate system; the right figure is the time series pair corresponding to the audio segment pair of the audio file without repetition. As shown in the figure, when using MFCC features for audio feature matching, whether it is repeated audio or non-repeated audio, there will be a large number of time series pairs scattered in all directions such as up, down, left, and right in the coordinate system, almost covering all positions of the rectangular area corresponding to the coordinate system. That is to say, whether it is repeated audio or non-repeated audio, there will be a large number of mis-matching situations (that is, chaotic time series pairs). If methods such as counting and ratio in related technologies are used, it is very easy to cause the algorithm to fail, resulting in incorrect recognition. However, through the audio recognition method of the present application, regression analysis is used to determine the characteristic change trend in time series, and further obtain the recognition result, so that the entire recognition process can tolerate a large number of unmatched time series pairs, thereby improving the robustness of the audio recognition method.
[0104] Figure 7 For Figure 6 the visualization result schematic diagram of the regression analysis of the left figure in Figure 7 shown Figure 7 In it, the straight line with a slope of 1 is the true time series pair recognized by the RANSAC model, and the others are randomly matched time series pairs. In Figure 7 it, the proportion of mis-matching points using MFCC features is as high as 96.5%, but through the RANSAC algorithm, the matching time series features can still be accurately analyzed, so that the straight line representing the correlation between the time series change regions of the first segment and the second segment can still be accurately found, thus ensuring accuracy and improving the practicality of the audio recognition method.
[0105] Based on the above Figure 5 , Figure 6 and Figure 7From the visualization schematic diagram, it can be seen that from the perspective of the recognition of the time series pairs of audio fingerprint features and MFCC features, by performing regression analysis on the time series pairs of audio segments to analyze the correlation between the changing trends in time series, repeated audio can be accurately and efficiently recognized. This method can not only be used for the recognition of repeated audio in the above-mentioned feature matching results, but also be generalized to various audio feature-based methods, such as combining audio to perform repeated recognition on videos. By performing regression analysis on the time series pairs and analyzing the correlation between the changing trends of audio segments in time series, the interference of noise points on the recognition results can be effectively reduced. Especially when using MFCC features for recognition, even when the noise matching rate is as high as 96.5%, the audio recognition method provided by the embodiments of the present application can still recognize the correct repeated time series matching features without being interfered by noise, thus ensuring the accuracy of audio recognition.
[0106] In a possible example, as Figure 8 shown, the server may include an audio feature vector generation module, and use the audio feature vector generation module to determine the feature vector sets of two audio files; the server may include an audio time series repeated recognition module, and use the audio time series repeated recognition module to perform recognition on the audio based on the above steps 202-204.
[0107] In a possible implementation manner, the server may recognize the audio included in the video file, that is, the first audio file and the second audio file are respectively the background audio files of any two of at least two video files; the server may also process the video file based on the recognition result. For example, recommend videos to the user based on the background audio. Then, after step 204, the following process may further be included: the server divides the at least two video files into corresponding video push sets based on the recognition result, and the background audio files corresponding to the video files in each video push set are repeated; the server pushes the target video set in the at least two video push sets to the user, and the target video set includes video files whose corresponding background audio files have repetitions with the user-preferred audio. The server may obtain the user-preferred audio. For example, use the audio collected by the user as the user-preferred audio.
[0108] Exemplarily, the server can utilize the recognition result to aggregate the video files of the application program into different audio compilations. Different audio compilations can include multiple video files with the same or similar corresponding audio files, so as to achieve more accurate video recommendations. It should be noted that in video information stream recommendations, through research, some users like video compilations with the same background music, such as different video contents produced with the same background music. In the information stream scenario, when creating these videos, the audio in the music library may not be directly used, but can be freely edited. For example, operations such as adding special effects and weakening the sound can be performed, making it difficult to recognize even video compilations with the same background music without using the audio recognition method of the present application. By using the audio recognition method provided in the embodiments of the present application, for example, the MFCC feature and the audio fingerprint feature can be combined, that is, the two features are combined separately to recognize the audio file, so as to be able to find that the essential background music or human voice dubbing is the same, improve the aggregation accuracy of the audio compilation, and recommend it to users. The server can include an audio similarity system, which performs audio recognition by this audio similarity system and aggregates based on the audio to obtain multiple audio compilations. The product rules are as Figure 9 shown. The Xx audio compilation can include three videos with the same background music, and three videos with the same background music can be recommended to the user to achieve accurate recommendation. Exemplarily, the server can also identify duplicate videos by using a pre-trained model. For example, the server can use the regression analysis result of the time series pair as a strong feature and input it into a machine learning model to identify duplicate videos.
[0109] In a possible implementation manner, the first audio file and the second audio file are respectively the audio files of any two video files; the server can further identify the video based on the recognition result of the audio file. After step 204, the following process can also be included: the server determines, based on the recognition result, the overlapping video files to be confirmed with duplicate corresponding audio files among the at least two video files, detects the overlapping videos with duplicate images included in the overlapping video files to be confirmed, and deletes the overlapping video files. Exemplarily, the server can combine the audio recognition result and the image recognition result of the video file to identify whether there is a duplicate between two video files. For example, when the audio files of two video files are duplicate and the images included in the two video files are also duplicate, the server determines that there is a duplicate between the two video files.
[0110] For example, in video similarity detection, it is often used to determine whether a video is repeated or similar by the content frame dimension of the image. Since the content frames may be in lectures, music clips, animations, games or lectures, the pictures between two videos with different class hours and different explanations are often extremely similar, only with changes in subtitles and sounds. For such videos with the same content but different themes, if only the images in the video are detected, it may lead to misdetection. As Figure 10 shown, Figure 10 the video files with extremely similar images detected by the video similarity auxiliary detection system are as follows. As Figure 10 shown, the upper and lower pictures on the left are the video tutorials of the Rubik's Cube. The upper and lower images correspond to two non-repeating videos with different lecture contents; the upper and lower images in the middle are the tutorial videos of the same type of musical instrument, but the upper and lower images correspond to two non-repeating videos with different class hour contents of the same musical instrument; the upper and lower pictures on the right are the tutorial videos of the same doctor guiding and explaining medical knowledge in the same environmental background. The upper and lower images correspond to two non-repeating videos with different disease explanations. From the images on the left, middle and right, it can be seen that if only based on the images, it is easy to misidentify the videos corresponding to the upper and lower pictures as two repeated videos. To solve such problems, the embodiments of the present application provide an audio recognition method, which can be combined and applied to the process of video recognition to improve the recognition accuracy. Through audio-assisted detection, it can better distinguish two videos with highly similar pictures but different human voice audios, thereby improving the accuracy of video recognition.
[0111] The above is illustrated by taking two application scenarios of video repetition detection and recommended videos as examples. The audio recognition method of the present application can also be widely applied in the information flow industry. For example, recognizing repeated audio can be applied to various purposes such as low-quality transportation, video source exploration, and audio copyright protection.
[0112] In this step, through associated information, such as regression coefficients, time series offsets, etc., the degree of association of the characteristic change trends of the two audio files is accurately quantified with specific values, so as to more accurately measure the degree of association. Moreover, through the time series offset, the starting repetition position of the two repeated audios can be obtained, and a more accurate, reliable and comprehensive recognition result can be further obtained, thereby improving the accuracy and comprehensiveness of audio recognition. In addition, based on the audio recognition result, it can be applied to the video recognition scenario. For example, the repetition detection of videos, recommending videos based on audio compilations, etc., so that the audio recognition method of the embodiments of the present application can be applied to a wider range of scenarios, improving the applicability of the audio recognition method.
[0113] The audio recognition method provided by the embodiments of the present application performs temporal correlation analysis on the first audio file and the second audio file based on the playing timings of audio clip pairs that match in features, to obtain correlation information. This correlation information can indicate the degree of association between the feature change trends of the two audio files based on their playing timings respectively, thereby accurately defining the temporal relationship between the two audio files. Further, the feature of the continuity of the audio is clearly represented by the feature change trends in terms of time sequence. Based on the correlation information between the feature change trends, it can be accurately identified whether there is duplication between the two audio files, improving the accuracy of audio recognition.
[0114] Moreover, by performing temporal correlation analysis on the two audio files based on the playing timings of the audio clips, the degree of association between the feature change trends of the two audio files is transformed into the association relationship between the temporal change trends of the feature-matching clips, so as to map the feature change trends of the audio files with the temporal change trends of the audio clip pairs, improving the accuracy of the determined degree of association between the feature change trends, and further ensuring the high accuracy rate of audio recognition.
[0115] Moreover, through regression analysis, the association relationship presented in terms of time sequence by the clips with the same or similar features in the two audio files can be accurately measured by the numerical values of the regression coefficient and the temporal offset; and the correct matching rule can be accurately found in a large amount of noise data, ensuring the accuracy and precision of the determined correlation information.
[0116] Moreover, the regression analysis is performed on the time sequence pairs, reducing the over-reliance of audio recognition on features, reducing the requirements for the feature generalization ability, reducing the requirements for feature design, further expanding the range of available features for audio recognition, and improving the practicability and reliability of audio recognition.
[0117] The embodiments of the present application provide an audio recognition device, as Figure 11 shown. The audio recognition device may include:
[0118] A determination module 1101, configured to determine audio clip pairs that match in features in the first audio file and the second audio file based on the audio features of the first audio file and the audio features of the second audio file;
[0119] An analysis module 1102, configured to perform temporal correlation analysis on the first audio file and the second audio file based on the playing timings of the audio clip pairs, to obtain correlation information, where the correlation information is used to indicate the degree of association between the feature change trend of the first audio file based on the playing timing and the feature change trend of the second audio file based on the playing timing;
[0120] An identification module 1103, configured to determine an identification result of the first audio file and the second audio file based on the association information, where the identification result is at least used to indicate whether the first audio file and the second audio file are duplicates.
[0121] In a possible implementation, the audio segment pair includes a first segment of the first audio file and a second segment of the second audio file;
[0122] The analysis module 1102 is further configured to analyze the association information of the variation trend of the playing time sequence of the first segment and the playing time sequence of the second segment in each audio segment pair based on the playing time sequence of the audio segment pair.
[0123] In a possible implementation, the association information includes a regression coefficient;
[0124] The analysis module 1102 is further configured to perform a regression analysis on the playing time sequence of the first segment and the playing time sequence of the second segment in the at least two time sequence pairs based on the at least two time sequence pairs of at least two audio segment pairs, and determine the regression coefficient obtained from the regression analysis as the association information; wherein, the regression coefficient is used to indicate the variation relationship of the playing time sequence of the second segment with respect to the playing time sequence of the first segment.
[0125] In a possible implementation, the association information includes a regression coefficient and a time sequence offset;
[0126] The analysis module 1102 is further configured to perform a regression analysis on the playing time sequence of the first segment and the playing time sequence of the second segment in the at least two time sequence pairs based on the at least two time sequence pairs of at least two audio segment pairs, to obtain a regression coefficient and a time sequence offset; determine the regression coefficient and the time sequence offset as the association information; wherein, the time sequence offset is used to indicate the offset distance of the playing time sequence between the segments with duplicates in the first audio file and the second audio file.
[0127] In a possible implementation, the identification result is further used to indicate the starting playing position corresponding to the segments with duplicates in the first audio file and the second audio file;
[0128] The identification module 1103 is further configured to determine that the first audio file and the second audio file are duplicates when the regression coefficient is within a first threshold range; and obtain the starting playing position corresponding to the segments with duplicates in the first audio file and the second audio file based on the time sequence offset.
[0129] In a possible implementation, the audio segment pair includes a first segment of the first audio file and a second segment of the second audio file; the determination module 1101 includes:
[0130] A matching unit, configured to perform feature matching on audio segments in the first audio file and audio segments in the second audio file based on the audio features of the first audio file and the audio features of the second audio file, so as to obtain at least two first segments and the second segments respectively corresponding thereto;
[0131] An obtaining unit, configured to, for each pair of audio segments, obtain the playing time sequence of the first segment of the pair of audio segments and the playing time sequence of the second segment corresponding thereto.
[0132] In a possible implementation manner, the matching unit is further configured to determine the feature distance between the feature vectors of each audio segment in the first audio file and the feature vectors of each audio segment in the second audio file, and determine the first segment whose feature distance is within the second threshold range and the second segment corresponding thereto as a pair of audio segments.
[0133] In a possible implementation manner, the determining module 1101 is further configured to perform any one of the following:
[0134] Determine the pairs of audio segments with feature matching in the target segment of the first audio file and the target segment of the second audio file, where the target segment is an audio segment in the corresponding audio file that is before the first time sequence threshold or after the second time sequence threshold, and the first time sequence threshold is greater than the second time sequence threshold;
[0135] Crop out a first part from the first audio file respectively, crop out a second part with the same playing duration as the first part from the second audio file, and determine the pairs of audio segments with feature matching in the first part and the second part.
[0136] In a possible implementation manner, the first audio file and the second audio file are respectively the background audio files of any two of at least two video files; the apparatus further includes:
[0137] A partitioning module, configured to partition the at least two video files into corresponding video push sets based on the recognition result, and there are duplicates in the background audio files corresponding to the video files in each video push set;
[0138] A pushing module, configured to push a target video set in the at least two video push sets to a user, where the target video set includes video files whose corresponding background audio files have duplicates with the user-preferred audio.
[0139] In a possible implementation manner, the first audio file and the second audio file are respectively the audio files of any two video files; the apparatus further includes:
[0140] A deletion module, configured to determine, based on the recognition result, overlapping video files to be confirmed in the at least two video files, where the corresponding audio files are duplicate; detect overlapping video files in the overlapping video files to be confirmed, where the included images are duplicate, and delete the overlapping video files.
[0141] The audio recognition method provided by the embodiments of the present application performs timing correlation analysis on the first audio file and the second audio file based on the playing timings of the audio segment pairs that match features, and obtains correlation information. The correlation information can indicate the degree of correlation between the characteristic change trends of the two audio files based on the playing timings respectively, so as to accurately define the relationship in timing between the two audio files. Further, the characteristic of the continuity of the audio is clearly represented by the characteristic change trends in timing. Based on the correlation information between the characteristic change trends, it can be accurately recognized whether there is duplication between the two audio files, improving the accuracy of audio recognition.
[0142] Moreover, by performing timing correlation analysis on the two audio files based on the playing timings of the audio segments, the degree of correlation between the characteristic change trends of the two audio files is transformed into the correlation relationship between the timing change trends of the segments that match features, so as to map the characteristic change trends of the audio files with the timing change trends of the audio segment pairs, improving the accuracy of the determined degree of correlation between the characteristic change trends, and further ensuring the high accuracy rate of audio recognition.
[0143] Moreover, through regression analysis, the correlation relationship presented in timing by the segments with the same or similar features in the two audio files can be accurately measured by the numerical values of the regression coefficient and the timing offset; and the correct matching rule can be accurately found in a large amount of noise data, ensuring the accuracy and precision of the determined correlation information.
[0144] Moreover, regression analysis is performed on the timing pairs, reducing the over-reliance of audio recognition on features, reducing the requirements for the feature generalization ability, reducing the requirements for feature design, further expanding the range of available features for audio recognition, and improving the practicability and reliability of audio recognition.
[0145] The audio recognition device of this embodiment can execute the audio recognition method shown in the foregoing embodiments of the present application, and its implementation principle is similar and will not be elaborated here.
[0146] In an embodiment of the present application, a computer device is provided. The computer device includes: a memory and a processor; at least one program stored in the memory and configured to, when executed by the processor, compared with the prior art, achieve: by analyzing the playback timings of audio segment pairs based on feature matching to perform temporal correlation analysis on a first audio file and a second audio file, obtaining correlation information, and the correlation information can indicate the degree of correlation between the characteristic change trends of the two audio files based on the playback timings respectively, so as to accurately define the temporal relationship between the two audio files, and thus clearly represent the characteristic of audio continuity through the characteristic change trends in time series. Based on the correlation information between the characteristic change trends, it is possible to accurately identify whether there is duplication between the two audio files, improving the accuracy of audio recognition.
[0147] In an alternative embodiment, a computer device is provided, as Figure 12 shown Figure 12 The computer device 4000 shown includes: a processor 4001 and a memory 4003. Among them, the processor 4001 and the memory 4003 are connected, such as connected through a bus 4002. Optionally, the computer device 4000 may further include a transceiver 4004, and the transceiver 4004 can be used for data interaction between the computer device and other computer devices, such as data sending and / or data receiving, etc. It should be noted that in practical applications, the transceiver 4004 is not limited to one, and the structure of the computer device 4000 does not constitute a limitation to the embodiments of the present application.
[0148] The processor 4001 may be a CPU (Central Processing Unit, central processor), a general-purpose processor, a DSP (Digital Signal Processor, digital signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of the present application. The processor 4001 may also be a combination that implements computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0149] The bus 4002 may include a path for transmitting information between the above components. The bus 4002 can be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus 4002 can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 12 it is only represented by a thick line in the figure, but it does not mean that there is only one bus or one type of bus.
[0150] The memory 4003 can be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0151] The memory 4003 is used to store the application program code (computer program) for executing the solution of this application, and is controlled by the processor 4001 for execution. The processor 4001 is used to execute the application program code stored in the memory 4003 to implement the content shown in the foregoing method embodiments.
[0152] Among them, the computer device includes but is not limited to: servers, computer devices, service clusters composed of multiple servers, etc.
[0153] The embodiment of this application provides a computer-readable storage medium, on which a computer program is stored. When it runs on a computer, it enables the computer to execute the corresponding content in the foregoing audio recognition method embodiments.
[0154] An embodiment of the present application provides a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the corresponding content in the above-mentioned embodiment of the audio recognition method.
[0155] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, there is no strict order restriction for the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.
[0156] The above are only some embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. An audio recognition method, characterized in that, The method includes: Based on the audio features of the first audio file and the audio features of the second audio file, determining audio segment pairs with matching features in the first audio file and the second audio file; the audio segment pairs include a first segment of the first audio file and a second segment of the second audio file; Based on the playback timings of the audio segment pairs, performing a timing correlation analysis on the first audio file and the second audio file to obtain correlation information, including: performing a regression analysis on the playback timings of the first segments and the playback timings of the second segments in at least two timing pairs based on at least two audio segment pairs, and determining the regression coefficient obtained from the regression analysis as the correlation information; the correlation information is used to indicate the degree of association between the feature change trend of the first audio file based on the playback timing and the feature change trend of the second audio file based on the playback timing, and the regression coefficient is used to indicate the change relationship of the playback timing of the second segment with respect to the playback timing of the first segment; Based on the correlation information, determining the recognition results of the first audio file and the second audio file.
2. The audio recognition method according to claim 1, wherein The correlation information further includes a timing offset obtained when performing the regression analysis; wherein, the timing offset is used to indicate the offset distance of the playback timings between the segments with duplicates in the first audio file and the second audio file.
3. The audio recognition method according to claim 2, wherein The determining the recognition results of the first audio file and the second audio file based on the correlation information includes: When the regression coefficient is within a first threshold range, determining that the first audio file and the second audio file have duplicates; Based on the timing offset, obtaining the starting playback positions of the segments with duplicates in the first audio file and the second audio file.
4. The audio recognition method according to claim 1, wherein The audio segment pairs include a first segment of the first audio file and a second segment of the second audio file; The determining the audio segment pairs with matching features in the first audio file and the second audio file based on the audio features of the first audio file and the audio features of the second audio file includes: Based on the audio features of the first audio file and the audio features of the second audio file, performing feature matching on the audio segments in the first audio file and the audio segments in the second audio file to obtain at least two first segments and their respectively corresponding second segments; For each audio segment pair, obtaining the playback timing of the first segment of the audio segment pair and the playback timing of its corresponding second segment.
5. The audio recognition method according to claim 4, wherein The performing feature matching on the audio segments in the first audio file and the audio segments in the second audio file based on the audio features of the first audio file and the audio features of the second audio file to obtain at least two first segments and their respectively corresponding second segments includes: Determining the feature distances between the feature vectors of each audio segment in the first audio file and the feature vectors of each audio segment in the second audio file, and determining a first segment with a feature distance within a second threshold range and its corresponding second segment as an audio segment pair.
6. The audio recognition method according to claim 1, wherein Determining the audio segment pairs with feature matching in the first audio file and the second audio file includes any of the following: Determining the audio segment pairs with feature matching in the target segment of the first audio file and the target segment of the second audio file, where the target segment is the audio segment in the corresponding audio file before the first timing threshold or after the second timing threshold, and the first timing threshold is greater than the second timing threshold; Determining the audio segment pairs with feature matching in the first part of the first audio file and the second part of the second audio file, where the first part and the second part are obtained by respectively cropping the first audio file and the second audio file based on the target playback duration.
7. The audio recognition method according to claim 1, characterized in that, The first audio file and the second audio file are respectively the background audio files of any two of at least two video files; After determining the recognition results of the first audio file and the second audio file based on the association information, the method further includes: Based on the recognition results, dividing the at least two video files into corresponding video push sets, where the background audio files corresponding to the video files in each video push set have duplicates; Pushing the target video set in the video push set to the user, where the target video set includes the video files whose corresponding background audio files have duplicates with the user-preferred audio.
8. The audio recognition method according to claim 1, wherein The first audio file and the second audio file are respectively the audio files of any two video files; After determining the recognition results of the first audio file and the second audio file based on the association information, the method further includes: Based on the recognition results, determining the overlapping video files to be confirmed in the two video files where the corresponding audio files have duplicates; Detecting the overlapping video files in the overlapping video files to be confirmed where the included images have duplicates, and deleting the overlapping video files.
9. An audio recognition device, characterized in that, The device includes: A determination module, configured to determine the audio segment pairs with feature matching in the first audio file and the second audio file based on the audio features of the first audio file and the audio features of the second audio file; the audio segment pair includes the first segment of the first audio file and the second segment of the second audio file; An analysis module, configured to perform timing correlation analysis on the first audio file and the second audio file based on the playback timing of the audio segment pair to obtain association information, including: performing regression analysis on the playback timing of the first segment and the playback timing of the second segment in at least two timing pairs based on at least two audio segment pairs, and determining the regression coefficient obtained from the regression analysis as the association information; the association information is used to indicate the association degree between the feature change trend of the first audio file based on the playback timing and the feature change trend of the second audio file based on the playback timing, and the regression coefficient is used to indicate the change relationship of the playback timing of the second segment with the playback timing of the first segment; A recognition module, configured to determine the recognition results of the first audio file and the second audio file based on the association information.
10. The device according to claim 9, characterized in that, The associated information further includes a timing offset obtained during the regression analysis; wherein the timing offset is used to indicate the offset distance of the playback timing between the repeated segments in the first audio file and the second audio file.
11. The device according to claim 10, characterized in that, The recognition result is further used to indicate the starting playback positions corresponding to the repeated segments in the first audio file and the second audio file; The recognition module is further configured to: When the regression coefficient is within the first threshold range, determine that the first audio file and the second audio file have repetitions; Based on the timing offset, obtain the starting playback positions corresponding to the repeated segments in the first audio file and the second audio file.
12. The device according to claim 9, characterized in that, The audio segment pair includes a first segment of the first audio file and a second segment of the second audio file; The determination module includes: A matching unit, configured to perform feature matching on the audio segments in the first audio file and the audio segments in the second audio file based on the audio features of the first audio file and the audio features of the second audio file, to obtain at least two first segments and the second segments respectively corresponding thereto; An obtaining unit, configured to, for each audio segment pair, obtain the playback timing of the first segment of the audio segment pair and the playback timing of the second segment corresponding thereto.
13. The device according to claim 12, characterized in that, The matching unit is further configured to determine the feature distance between the feature vectors of each audio segment in the first audio file and the feature vectors of each audio segment in the second audio file, and determine a first segment with the feature distance within the second threshold range and the second segment corresponding thereto as an audio segment pair.
14. The device according to claim 9, characterized in that, The determination module is further configured to perform any one of the following: Determine the audio segment pairs with feature matching in the target segment of the first audio file and the target segment of the second audio file, where the target segment is an audio segment in the corresponding audio file that is before the first timing threshold or after the second timing threshold, and the first timing threshold is greater than the second timing threshold; Determine the audio segment pairs with feature matching in the first part of the first audio file and the second part of the second audio file, where the first part and the second part are obtained by respectively cropping the first audio file and the second audio file based on the target playback duration.
15. The device according to claim 9, characterized in that, The first audio file and the second audio file are respectively the background audio files of any two of at least two video files; The apparatus further includes: A partitioning module, configured to partition the at least two video files into corresponding video push sets based on the recognition result, where the background audio files corresponding to the video files in each video push set have repetitions; A pushing module, configured to push a target video set in the at least two video push sets to the user, where the target video set includes video files whose corresponding background audio files have repetitions with the user-preferred audio.
16. The device according to claim 9, characterized in that, The first audio file and the second audio file are respectively the audio files of any two video files; The apparatus further includes: A deletion module, configured to determine, based on the recognition result, overlapping video files to be confirmed in the two video files where the corresponding audio files are duplicate; detect overlapping video files in the overlapping video files to be confirmed where the included images are duplicate, and delete the overlapping video files.
17. A computer device, characterized in that, The computer device includes: One or more processors; A memory; One or more computer programs, wherein the one or more application programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to: execute the audio recognition method according to any one of claims 1 to 8.
18. A computer-readable storage medium, characterized in that, The computer storage medium is used to store computer instructions, and when the computer instructions are run on a computer, the computer is caused to execute the audio recognition method according to any one of claims 1 to 8 above.
19. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the audio recognition method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Voice similarity determination method and device, and program product
CN112951274A