Audio processing method and device, electronic equipment, computer readable storage medium and computer program product

By extracting audio fingerprints and calculating offset values ​​to automatically synthesize audio, the problem of difficult voice aligning in karaoke software is solved, accuracy is improved, the cost of manual adjustment is reduced, and the user experience is enhanced.

CN120656433APending Publication Date: 2025-09-16TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410307817.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-15
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

In karaoke software, the delay between the accompaniment and the vocals when users record songs makes it difficult to align the accompaniment. Existing technologies consume a lot of manpower and have low accuracy.

Method used

By obtaining dry audio and reference audio data, extracting audio fingerprints and calculating offset values, the audio is automatically synthesized to achieve voice partner alignment.

Benefits of technology

The accuracy of voice companion alignment is improved, the cost of manual adjustment is reduced, and the user experience is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656433A_ABST
    Figure CN120656433A_ABST
Patent Text Reader

Abstract

The invention provides an audio processing method and device, electronic equipment, a computer readable storage medium and a computer program product. The method comprises the steps that dry sound audio data and reference audio data of a target object are acquired, and the reference audio data comprise dry sound audio data of an original singing and accompaniment audio data; performing fingerprint extraction on the dry audio data of the target object to obtain a first audio fingerprint corresponding to the dry audio data of the target object; performing fingerprint extraction on the reference audio data to obtain a second audio fingerprint corresponding to the reference audio data; based on a comparison result of the first audio fingerprint and the second audio fingerprint, determining an offset value between dry audio data and accompaniment audio data of the target object; and synthesizing the dry audio data and the accompaniment audio data of the target object based on the offset value to obtain a synthesized audio. According to the method and the device, the accuracy of sound accompanying alignment can be improved, the labor cost of manual adjustment is reduced, and the use experience of a user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of Internet technology, and in particular to an audio processing method, device, electronic device, computer-readable storage medium, and computer program product. Background Art

[0002] With the popularity of karaoke software, more and more users are singing and recording their favorite songs through karaoke software, and using karaoke software to post-process the user's dry voice, such as background noise processing, pitch and rhythm calibration, etc. However, in the process of recording songs, due to hardware or personal reasons of the user, there is often a certain delay between the accompaniment and the vocals, which requires manual adjustment of the recorded vocals later. This method consumes a lot of manpower costs, and when the delay difference is small, the manually adjusted vocal accompaniment alignment method has low accuracy. Summary of the Invention

[0003] The embodiments of the present application provide an audio processing method, device, electronic device, computer-readable storage medium and computer program product, which can improve the accuracy of sound companion alignment while reducing the labor cost of manual adjustment, thereby improving the user experience.

[0004] The technical solution of the embodiment of the present application is implemented as follows:

[0005] The present invention provides an audio processing method, including:

[0006] Acquire dry audio data and reference audio data of a target object, wherein the reference audio data includes the dry audio data of the original singer and the accompaniment audio data;

[0007] Performing fingerprint extraction on the dry audio data of the target object to obtain a first audio fingerprint corresponding to the dry audio data of the target object;

[0008] Extracting a fingerprint from the reference audio data to obtain a second audio fingerprint corresponding to the reference audio data;

[0009] determining an offset value between the dry audio data of the target object and the accompaniment audio data based on a comparison result of the first audio fingerprint and the second audio fingerprint;

[0010] The dry audio data of the target object and the accompaniment audio data are synthesized based on the offset value to obtain synthesized audio.

[0011] The present invention provides an audio processing device, including:

[0012] An acquisition module, configured to acquire dry audio data and reference audio data of a target object, wherein the reference audio data includes the dry audio data of the original singer and the accompaniment audio data;

[0013] a fingerprint extraction module, configured to extract a fingerprint from the dry audio data of the target object to obtain a first audio fingerprint corresponding to the dry audio data of the target object;

[0014] The fingerprint extraction module is further configured to extract a fingerprint from the reference audio data to obtain a second audio fingerprint corresponding to the reference audio data;

[0015] a determination module, configured to determine an offset value between the dry audio data of the target object and the accompaniment audio data based on a comparison result of the first audio fingerprint and the second audio fingerprint;

[0016] A synthesis module is used to synthesize the dry sound audio data of the target object and the accompaniment audio data based on the offset value to obtain synthesized audio.

[0017] An embodiment of the present application provides an electronic device, including:

[0018] a memory for storing executable instructions;

[0019] The processor is configured to implement the audio processing method provided in the embodiment of the present application when executing the executable instructions stored in the memory.

[0020] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions for implementing the audio processing method provided in the embodiment of the present application when executed by a processor.

[0021] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions, which are used to implement the audio processing method provided in the embodiment of the present application when executed by a processor.

[0022] The embodiments of the present application have the following beneficial effects:

[0023] By obtaining the dry audio data and reference audio data of the target object, and performing fingerprint extraction on the dry audio data and reference audio data of the target object respectively, the corresponding first audio fingerprint and second audio fingerprint are obtained, and according to the comparison result of the first audio fingerprint and the second audio fingerprint, the offset value between the dry audio data and the accompaniment audio data of the target object is determined, and finally the dry audio data and the accompaniment audio data of the target object are synthesized based on the offset value to obtain synthesized audio. That is to say, the embodiment of the present application can improve the accuracy of sound accompaniment alignment while reducing the manpower cost of manual adjustment by converting audio data into waveforms and performing visual comparison through image form, thereby improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 1 is a schematic diagram of the architecture of the audio processing system 100 provided in an embodiment of the present application;

[0025] Figure 2 is a structural diagram of an electronic device 500 provided in an embodiment of the present application;

[0026] Figure 3 This is a flowchart of the audio processing method provided in an embodiment of the present application;

[0027] Figure 4A This is a flowchart of the audio processing method provided in an embodiment of the present application;

[0028] Figure 4B This is a flowchart of the audio processing method provided in an embodiment of the present application;

[0029] Figure 5A This is a flowchart of the audio processing method provided in an embodiment of the present application;

[0030] Figure 5B This is a flowchart of the audio processing method provided in an embodiment of the present application;

[0031] Figure 6 This is a schematic diagram of the process of extracting audio fingerprints based on the perceptual hash algorithm provided in an embodiment of the present application;

[0032] Figure 7 This is a schematic diagram of an enlarged local area of ​​a spectrogram provided in an embodiment of the present application;

[0033] Figure 8 This is a graph showing the similarity change trend of songs after alignment provided by an embodiment of the present application;

[0034] Figure 9 This is a schematic diagram of manually calculating offsets and comparing audio fingerprints provided in an embodiment of the present application. DETAILED DESCRIPTION

[0035] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0036] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0037] It is understandable that in the embodiments of the present application, when user information and other related data (such as audio data recorded by the user) are involved, when the embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.

[0038] In the following description, the terms "first\second\..." are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understandable that "first\second\..." can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0039] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0040] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.

[0041] 1) Dry sound: Dry sound refers to sound that has not been modified or processed later. It is also called raw sound or unpolished sound. It is recorded using professional recording equipment (such as microphones, sound cards, mixing consoles, etc.), maintaining the original sound quality and sound effects.

[0042] 2) Audio fingerprint: Audio fingerprint is a set of biometric recognition technologies used for audio content retrieval. It is mainly used to match similar song content within a few seconds of audio, that is, to extract the unique "fingerprint" of the audio to achieve rapid comparison and retrieval of audio content.

[0043] 3) Offset: The offset value refers to the time difference between the target dry sound and the accompaniment. This difference can be caused by a variety of factors, such as recording equipment latency, post-processing inaccuracies, and personal vocal delays. By adjusting the target dry sound and accompaniment based on the offset value, the target dry sound and accompaniment can be perfectly synchronized in time, improving the listening experience of the musical work. Specifically, the offset value can be positive or negative, indicating whether the target dry sound is ahead or behind the accompaniment.

[0044] 4) Fourier Transform (FT): A method for analyzing signals. It can be used to analyze signal components or synthesize signals from these components. In signal processing, the Fourier transform is typically used to decompose a signal into amplitude and frequency components.

[0045] 5) Spectrogram: A spectrogram is a graph of the speech spectrum, obtained by processing the received time-domain signal. A spectrogram can be generated with any time-domain signal of sufficient length. In a spectrogram, the horizontal axis represents time, the vertical axis represents frequency, and the value at each coordinate represents the energy of the speech data. A spectrogram uses a two-dimensional plane to represent three-dimensional information, with color representing energy levels. Darker colors indicate higher speech energy at that point.

[0046] 6) Spectrogram: A spectrogram is a graphical representation of a signal's frequency components. It plots signal frequency on the horizontal axis and signal strength on the vertical axis, using different colors or grayscales to display signal strength at different frequencies, thus clearly illustrating the signal's spectral characteristics. A spectrogram visually illustrates the signal strength, or "loudness," of a signal at various frequencies within a specific waveform over time. It helps people more intuitively understand and analyze the signal's spectral characteristics. Common spectrograms include amplitude spectrograms and phase spectrograms.

[0047] 7) Perceptual Hash Algorithm (PHA): This is a general term for a class of hash algorithms that generates an image "fingerprint" and compares its similarity. This algorithm determines image similarity by comparing the fingerprint information of different images. There are various specific implementations of perceptual hashing algorithms, including average hashing (aHash), perceptual hashing (pHash), and differential value hashing (dHash).

[0048] 8) Amplitude information: Amplitude information refers to the range or size of the change in sound intensity in the audio signal, and is a key parameter for describing the change in sound intensity.

[0049] 9) Phase information: Phase information refers to the parameter of waveform offset in the audio signal, which describes the position of the sound waveform at any time and is an important parameter for describing the offset of the sound waveform.

[0050] Artificial intelligence technology plays a vital role in audio processing, with widespread applications in speech recognition, audio synthesis, audio classification and annotation, audio enhancement and noise reduction, and intelligent audio editing. It can improve the accuracy, efficiency, and user experience of audio processing. By leveraging machine learning, deep learning, and other technologies, audio processing tools can better process audio data and provide enhanced audio processing solutions.

[0051] Key technologies in speech technology include automatic speech recognition (ASR), text-to-speech (TTS), and voiceprint recognition. Enabling computers to hear, see, speak, and feel is the future direction of human-computer interaction, with speech becoming one of the most promising methods of human-computer interaction. Large model technology is revolutionizing the development of speech technology. Pre-trained models such as WavLM and UniSpeech, which leverage the Transformer architecture, possess strong generalization and versatility, enabling them to effectively handle a wide range of speech processing tasks.

[0052] Among them, automatic voice-accompaniment alignment refers to calculating a recommended offset value for the synthesis of the user's dry voice and accompaniment after the recording is completed, so that the synthesized audio achieves a more ideal listening effect. During the audio recording process, users often sing along with the rhythm, but the recorded result will be different from the accompaniment, which will affect the final effect of the work. A common problem in mobile phone recording is that the user feels that he has kept up with the rhythm during recording, but the user's dry voice and accompaniment are not aligned when previewing the page; for ordinary users, under normal circumstances, a delay within 25 milliseconds (ms, millisecond) is almost imperceptible to the ear, and a delay of about 20ms-50ms, if the user is familiar with the song, can be perceived to a certain extent but it is not obvious; when the delay exceeds 50ms, the user's listening experience will be more obviously dull. The main links that cause delays include the following:

[0053] (1) Delay caused by system processing. The recording delay of the operating system is relatively ideal. When the user uses wired headphones for recording, the delay caused by the system is almost imperceptible. When on the Android system, due to the complex hardware environment and split system versions, the recording delays caused by different devices are not the same. In most cases, audio playback and recording will execute a separate thread, which may also cause the user's dry voice and accompaniment to be out of sync. It should be noted here that when using Bluetooth headphones for recording, the delay caused by wireless transmission will show a more obvious delay than wired transmission, especially in scenarios that are relatively sensitive to time delay. For example, there will be more obvious differences between different recording software. For example, the delay of the AirPods series of wireless headphones when recording on an iPhone is about 150ms-180ms. If the sound accompaniment is not aligned, it will seriously affect the user's listening experience.

[0054] (2) Offset value of the reference object. The recorded user dry voice will eventually be mixed with the accompaniment, which means that the alignment is based on the accompaniment. However, there is no guarantee that the user will sing along with the accompaniment. Many users like to record songs in the original singing mode, or sing unfamiliar songs following the rendering progress of the lyrics text. Due to copyright, version / live and other reasons, there is no guarantee that the time axis of the accompaniment / original singing / lyrics file is completely consistent. If the original singing is about 40ms slower than the accompaniment, then if the user sings along with the original singing, the dry voice recorded at the end will be aligned with the original singing, and there is a high possibility that it will deviate from the accompaniment.

[0055] (3) Personal subjective influence. Singing is a way for users to express their emotions. When users are singing, they may be immersed in it and lose control of themselves, causing them to forget to control the rhythm of the singing. The recorded user's dry voice may have a subjective advance or delay.

[0056] Furthermore, during the implementation of the embodiments of the present application, the applicant discovered that the solutions provided by the related art had the following problems:

[0057] In some implementations, it is difficult for users to grasp the timing of starting to sing. Therefore, during the recording process, a countdown reminder will be used to assist users in judging the time to start singing. However, the prelude rhythm of some songs is not very obvious, and it is difficult for users to grasp it accurately. Therefore, there is a high probability that users will have a little deviation when they start singing the first sentence of the song. As the user gets into the state, they will slowly find the rhythm. The relevant voice alignment technology will align according to the time point of starting to sing, and there is a high probability that problems will occur. This voice alignment implementation process needs to rely on the accuracy of fundamental frequency extraction and lyrics text information. This solution relies more on external conditions; in other implementations, voice recognition is performed on the user's dry voice and alignment is performed according to the boundary of each word. This voice alignment implementation solution has high requirements on voice quality and consumes a lot of computing resources. Due to the diversity of languages, the accuracy of voice alignment will be limited.

[0058] In view of this, embodiments of the present application provide an audio processing method, apparatus, electronic device, computer-readable storage medium, and computer program product that can improve the accuracy of voice partner alignment while reducing the labor cost of manual adjustment, thereby enhancing the user experience. The following describes an exemplary application of an electronic device for performing audio processing provided by embodiments of the present application. The electronic device in the embodiments of the present application can be a server or a terminal.

[0059] See also Figure 1 , Figure 1 This is an architectural diagram of the audio processing system 100 provided in an embodiment of the present application. In order to support an application for improving the accuracy of sound companion alignment, the terminal 400 is connected to the server 200 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.

[0060] The terminal 400 is used to display an audio upload interface to the user on the graphical interface 410, and obtain the dry audio data and reference audio data of the target object (for example, user A) uploaded by the user from the interface, wherein the reference audio data includes the dry audio data and accompaniment audio data of the original singer. The terminal 400 uploads the dry audio data and the reference audio data of the target object to the server 200 through the network 300. The server 200 extracts the fingerprint of the dry audio data of the target object to obtain a first audio fingerprint corresponding to the dry audio data of the target object; then the server 200 extracts the fingerprint of the reference audio data to obtain a second audio fingerprint corresponding to the reference audio data; then the server 200 determines the offset value between the dry audio data and the accompaniment audio data of the target object based on the comparison result of the first audio fingerprint and the second audio fingerprint; then the server 200 synthesizes the dry audio data and the accompaniment audio data of the target object based on the offset value to obtain synthesized audio; finally, the server 200 transmits the synthesized audio to the terminal 400 through the network 300, and plays the synthesized audio at the terminal 400.

[0061] In other embodiments, the embodiments of the present application can also be implemented with the help of cloud technology (Cloud Technology) and artificial intelligence technology (AI, Artificial Intelligence technology). Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and network within a wide area network or local area network to realize data calculation, storage, processing and sharing; artificial intelligence technology refers to a discipline and technology that uses computer science and technology to simulate, extend and expand human intelligence, and simulates various aspects of human intelligence through learning, reasoning, perception, understanding, communication, etc.

[0062] Cloud technology is a general term for network, information, integration, management platform, and application technologies used in the cloud computing business model. It can form a resource pool that can be used flexibly and conveniently on demand. Cloud computing technology will become a key support. The backend services of technical network systems require a large amount of computing and storage resources.

[0063] For example, Figure 1The server 200 shown in the figure can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks (CDN, Content Delivery Network), and basic cloud computing services such as big data and artificial intelligence platforms. The terminal can be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, car terminal, etc., but is not limited to this. The terminal 400 and the server 200 can be directly or indirectly connected by wired or wireless communication, which is not limited in the embodiments of the present application.

[0064] It should be noted that the audio processing method provided in the embodiments of the present application can be applied to various audio processing scenarios such as speech recognition and synthesis, audio editing and processing, and music creation.

[0065] In some embodiments, the terminal or server can also implement the audio processing method provided in the embodiment of the present application by running various computer executable instructions or computer programs. For example, computer executable instructions can be microprogram-level commands, machine instructions or software instructions. The computer program can be a native program or software module in the operating system; it can be a local (Native) application (APP, APPlication), that is, a program that needs to be installed in the operating system to run, such as an audio playback APP (such as a karaoke software); it can also be a small program that can be embedded in any APP, that is, a program that can be run only by downloading it to a browser environment. In short, the above-mentioned computer executable instructions can be instructions in any form, and the above-mentioned computer program can be an application, module or plug-in in any form.

[0066] The following continues to describe the structure of the electronic device provided by the embodiment of the present application. Take the electronic device as an example, see Figure 2 , Figure 2 is a structural diagram of an electronic device 500 provided in an embodiment of the present application, Figure 2 The electronic device 500 shown includes: at least one processor 510, a memory 550, at least one network interface 520 and a user interface 530. The various components in the electronic device 500 are coupled together via a bus system 540. It is understood that the bus system 540 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 540 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 540 is not shown in FIG. Figure 2 Various buses are labeled as bus system 540 .

[0067] The processor 510 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0068] The user interface 530 includes one or more output devices 531 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 530 also includes one or more input devices 532, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0069] The memory 550 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 550 may optionally include one or more storage devices that are physically remote from the processor 510.

[0070] The memory 550 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 550 described in the embodiments of the present application is intended to include any suitable type of memory.

[0071] In some embodiments, the memory 550 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.

[0072] Operating system 551, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic services and processing hardware-based tasks;

[0073] A network communication module 552 for reaching other computing devices via one or more (wired or wireless) network interfaces 520 , exemplary network interfaces 520 including Bluetooth, WiFi, and USB;

[0074] a presentation module 553 for enabling presentation of information via one or more output devices 531 (e.g., a display screen, a speaker, etc.) associated with the user interface 530 (e.g., a user interface for operating peripheral devices and displaying content and information);

[0075] The input processing module 554 is configured to detect one or more user inputs or interactions from one of the one or more input devices 532 and to translate the detected inputs or interactions.

[0076] In some embodiments, the apparatus provided in the embodiments of the present application may be implemented in software. Figure 2 An audio processing device 555 stored in the memory 550 is shown, which can be software in the form of a program and plug-in, etc., including the following software modules: an acquisition module 5551, a fingerprint extraction module 5552, a determination module 5553 and a synthesis module 5554. These modules are logical and can therefore be arbitrarily combined or further split according to the functions implemented.

[0077] In other embodiments, the apparatus provided in the embodiments of the present application may be implemented in hardware. As an example, the apparatus provided in the embodiments of the present application may be a processor in the form of a hardware decoding processor, which is programmed to execute the audio processing method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor may be one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0078] The audio processing method provided in the embodiment of the present application will be described in detail below in conjunction with the exemplary application and implementation of the terminal provided in the embodiment of the present application.

[0079] See also Figure 3 , Figure 3 This is a flowchart of the audio processing method provided by the embodiment of the present application, which will be combined with Figure 3 The steps shown are explained.

[0080] It should be noted that Figure 3The audio processing method shown can be executed by various computer programs running on a terminal, not limited to a client. For example, it can also be an operating system, software module, script, and applet as described above. Therefore, the example of a client below should not be considered a limitation of the embodiments of the present application. In addition, for the sake of convenience, the following description does not specifically distinguish between a terminal and a client running on the terminal.

[0081] In step 101 , dry audio data and reference audio data of a target object are obtained, wherein the reference audio data includes the dry audio data and accompaniment audio data of the original singer.

[0082] It should be noted here that the dry audio data of the target object (such as user 1) can be obtained by user 1 uploading audio data, or by recording through relevant audio recording equipment, which is not specifically limited here; and the reference audio data can include not only the dry audio data of the original singer and the accompaniment audio data, but also the lyrics file (QRC, Qt Resources File) corresponding to the audio. The reference audio data is the audio data corresponding to the dry audio data of the target object. For example, when the reference audio data is a song, the dry audio data of the target object is the dry voice of the user singing and recording into the terminal according to the reference audio, or the dry voice of the user singing according to the reference audio is uploaded to the terminal.

[0083] In step 102 , fingerprint extraction is performed on the dry audio data of the target object to obtain a first audio fingerprint corresponding to the dry audio data of the target object.

[0084] In some embodiments, see Figure 4A , Figure 4A is a flowchart of the audio processing method provided in an embodiment of the present application, Figure 3 Step 102 shown may be performed by Figure 4A Steps 1021 to 1023 shown are implemented by combining Figure 4A The steps shown are explained.

[0085] In step 1021, Fourier transform is performed on the dry audio data of the target object to obtain a first spectrogram corresponding to the audio data of the target object.

[0086] In some embodiments, the dry sound audio data of the target object is first loaded. For example, an audio processing library (for example, Librosa, PyDub, etc.) can be used to load the dry sound audio data of the target object, and the dry sound audio data of the target object is subjected to signal conversion to convert the dry sound audio data of the target object into a digital signal. Before Fourier transforming the audio data, it is usually necessary to preprocess the dry sound audio data of the target object to obtain the preprocessed dry sound audio data of the target object. For example, the dry sound audio data can be subjected to denoising, normalization, segmentation and other processing operations to ensure the quality and stability of the audio data. Then, Fourier transform is performed on the preprocessed dry sound audio data of the target object to convert the dry sound audio data of the target object from the time domain to the frequency domain, and then the frequency characteristics of the audio data are analyzed to obtain the Fourier transform result. The power spectral density (PSD) is calculated from the Fourier transform result, and according to the calculated power spectral density data, the corresponding first spectrogram is drawn through a graphics library (for example, matplotlib).

[0087] In step 1022, the first spectrogram is filtered to obtain a first frequency spectrum corresponding to the dry audio data of the target object.

[0088] In some embodiments, the first spectrogram may include first amplitude information and first phase information. The above-mentioned filtering of the first spectrogram to obtain the first spectrum graph corresponding to the dry sound audio data of the target object can be achieved in the following way: calling a filter to filter the first amplitude information and the first phase information to obtain the first spectrum graph corresponding to the dry sound audio data of the target object.

[0089] In some embodiments, it is first necessary to load the acquired first spectrogram, and an image processing library (such as OpenCV, matplotlib, etc.) can be used to load the image file corresponding to the spectrogram; before filtering the spectrogram, it is necessary to select a suitable filter type according to the filtering operation to be performed, and set the corresponding filter parameters according to specific needs, and then apply the defined filter to the first spectrogram to obtain the filtered first spectrum graph.

[0090] For example, OpenCV can be used to load the image file of the first spectrogram, define a Bark filter (BarkFilter) and set the corresponding filter parameters, use the Bark filter to filter the first spectrogram, and finally draw the filtered first spectrum to achieve filtering processing of the spectrogram.

[0091] In step 1023, hash sensing is performed on the first spectrogram to obtain a first audio fingerprint corresponding to the dry audio data of the target object.

[0092] In some embodiments, see Figure 4B , Figure 4B is a flowchart of the audio processing method provided in an embodiment of the present application, Figure 4A Step 1023 shown may be performed by Figure 4B Steps 10231 to 10235 shown are implemented by combining Figure 4B The steps shown are explained.

[0093] In step 10231, the first spectrogram is reduced to a fixed size, and the reduced first spectrogram is converted into a first grayscale image.

[0094] For example, the first spectrum graph can be reduced to 32×32 pixels, and the pixel values ​​of the three RGB channels corresponding to the reduced first spectrum graph are averaged. The obtained average is used as the grayscale value of the pixel point, and the first grayscale image is obtained by drawing according to the grayscale value. This can effectively remove the differences between various image sizes and image ratios, retaining only basic information such as structure and brightness, so as to ensure image consistency and effectively reduce the complexity of calculation.

[0095] In step 10232, a discrete cosine transform is performed on the first grayscale image to obtain a first discrete cosine transform coefficient matrix.

[0096] For example, the fft.dct2d function in the Numpy library may be called to convert the first grayscale image into a first discrete cosine transform coefficient matrix of size 32×32.

[0097] It should be noted here that the discrete cosine transform (DCT) converts an image in one-dimensional or multi-dimensional space from the spatial domain to a discrete cosine transform coefficient matrix in the frequency domain to capture the low-frequency information in the image and ignore the high-frequency information to achieve image compression. This method not only retains the main features of the image, but also greatly reduces the data storage and transmission requirements.

[0098] In step 10233, the first discrete cosine transform coefficient matrix is ​​reduced, and the mean value of the reduced first discrete cosine transform coefficient matrix is ​​determined.

[0099] In some embodiments, after DCT transformation, if the frequency characteristics of the image are concentrated in the upper left corner of the coefficient matrix, only the 8×8 coefficient sub-matrix in the upper left corner of the coefficient matrix can be retained to further compress the data and reduce redundancy, while further expanding the low-frequency information of the image; then the overall mean of the image block after DCT transformation is calculated to subsequently determine the brightness and darkness of each image block.

[0100] In step 10234, the multiple elements included in the reduced first discrete cosine transform coefficient matrix are compared with the mean value in sequence, and the values ​​of the multiple elements are updated according to the comparison results.

[0101] In some embodiments, the DCT coefficient values ​​of each element in the reduced first discrete cosine transform coefficient matrix can be compared with the mean of the entire discrete cosine transform coefficient matrix, and elements with DCT coefficients above the mean are represented as 1, and otherwise represented as 0. The first discrete cosine transform coefficient matrix is ​​updated based on the comparison results to obtain the same number of binary values ​​as the discrete cosine transform coefficient matrix. Continuing with the above example, the 64 elements included in the 8×8 coefficient submatrix are sequentially compared with the mean of the entire submatrix, and the values ​​of the 64 elements are updated based on the comparison results, ultimately obtaining a 64-bit binary value.

[0102] In step 10235 , the updated values ​​of the multiple elements are combined into a first character string to serve as a first audio fingerprint corresponding to the dry audio data of the target object.

[0103] In some embodiments, since the values ​​of multiple elements are updated according to the comparison results to obtain binary values, and the representation of binary values ​​is lengthy, the numerical values ​​of the binary values ​​are grouped according to characters to generate a first character string as the first audio fingerprint corresponding to the dry audio of the target object.

[0104] For example, continuing with the above example, since a 64-bit binary value is too long, it can be converted from binary to hexadecimal in groups of 4 characters. In this way, a string of 16 characters can be obtained. This string is the first audio fingerprint corresponding to the first spectrogram obtained through hash perception, that is, the audio features contained in the first spectrogram.

[0105] In other embodiments, the above-mentioned fingerprint extraction of the dry sound audio data of the target object to obtain the first audio fingerprint corresponding to the dry sound audio data of the target object can also be achieved by inputting the dry sound audio data of the target object into a trained audio fingerprint extraction model to obtain the first audio fingerprint corresponding to the dry sound audio data of the target object.

[0106] For example, the dry audio data of the target object is used as input data and input into the trained audio fingerprint extraction model. The audio fingerprint extraction model generates a first audio fingerprint corresponding to the dry audio data of the target object based on the input audio data.

[0107] Before extracting fingerprints from dry audio data of a target object using an audio fingerprint extraction model, the audio fingerprint extraction model needs to be trained in advance.

[0108] In some embodiments, the audio fingerprint extraction model can be trained by: obtaining a training sample, wherein the training sample includes sample audio data and an audio fingerprint corresponding to the sample audio data; inputting the sample audio data into the audio fingerprint extraction model, performing audio fingerprint prediction on the sample audio data, and obtaining a predicted audio fingerprint; obtaining the difference between the predicted audio fingerprint and the audio fingerprint corresponding to the sample audio data, and updating the model parameters of the audio fingerprint extraction model based on the difference. Here, the audio fingerprint extraction model can be trained using machine learning methods (such as deep learning, generative adversarial networks, etc.), or other commonly used model training methods (such as pre-training methods, etc.), and is not specifically limited here.

[0109] In practical applications, the value of the loss function of the audio fingerprint extraction model is determined based on the difference between the predicted audio fingerprint and the audio fingerprint corresponding to the sample audio data. When the value of the loss function of the audio fingerprint extraction model reaches a parameter threshold, the corresponding target audio fingerprint is determined based on the loss function of the audio fingerprint extraction model, and the target audio fingerprint is back-propagated in the audio fingerprint extraction model, and the model parameters of the audio fingerprint extraction model are updated during the propagation process.

[0110] Here we explain back propagation. The training sample data is input into the input layer of the neural network model, passes through the hidden layer, and finally reaches the output layer and outputs the result. This is the forward propagation process of the neural network model. Since there is an error between the output result of the neural network model and the actual result, the error between the result and the actual value is calculated, and the error is backpropagated from the output layer to the hidden layer until it propagates to the input layer. During the back propagation process, the value of the model parameter is adjusted according to the error; the above process is continuously iterated until convergence.

[0111] In step 103, fingerprint extraction is performed on the reference audio data to obtain a second audio fingerprint corresponding to the reference audio data.

[0112] Here, fingerprint extraction can be performed on the dry audio data of the original singer included in the reference audio data to obtain a second audio fingerprint corresponding to the dry audio data of the original singer, or fingerprint extraction can be performed on the synthesized audio data (i.e., reference audio data) obtained by synthesizing the dry audio data of the original singer and the accompaniment audio data to obtain a second audio fingerprint corresponding to the synthesized audio data. No specific limitation is made here.

[0113] It should be noted that the implementation method of extracting the fingerprint of the reference audio data to obtain the second audio fingerprint corresponding to the reference audio data in step 103 is similar to the implementation method of step 102. For details, please refer to the implementation method of the above step 102 and will not be repeated here.

[0114] In some embodiments, see Figure 5A , Figure 5A is a flowchart of the audio processing method provided in an embodiment of the present application, Figure 3 Step 103 shown can be performed by Figure 5A Steps 1031 to 1033 shown are implemented by combining Figure 5A The steps shown are explained.

[0115] In step 1031, Fourier transform is performed on the reference audio data to obtain a second spectrogram corresponding to the reference audio data.

[0116] It should be noted that the implementation method of step 1031 of performing Fourier transform on the reference audio data to obtain the second spectrogram corresponding to the reference audio data is similar to the implementation method of step 1021. For details, please refer to the implementation method of the above-mentioned step 1021, which will not be repeated here.

[0117] In step 1032, the second spectrogram is filtered to obtain a second spectrogram corresponding to the reference audio data.

[0118] It should be noted that the implementation method of step 1032 of filtering the second spectrogram to obtain the second spectrogram corresponding to the reference audio data is similar to the implementation method of step 1022. For details, please refer to the implementation method of the above step 1022, which will not be repeated here.

[0119] In step 1033, hash sensing is performed on the second spectrogram to obtain a second audio fingerprint corresponding to the reference audio data.

[0120] It should be noted that the implementation method of performing hash sensing on the second spectrogram to obtain the second audio fingerprint corresponding to the reference audio data in step 1033 is similar to the implementation method of step 1023. For details, please refer to the implementation method of the above-mentioned step 1023 and will not be repeated here.

[0121] In some embodiments, the second spectrogram may include second amplitude information and second phase information. The above-mentioned filtering of the second spectrogram to obtain the second spectrum graph corresponding to the reference audio data can be achieved in the following way: calling a filter to filter the second amplitude information and the second phase information to obtain the second spectrum graph corresponding to the reference audio data.

[0122] In some embodiments, see Figure 5B , Figure 5B is a flowchart of the audio processing method provided in an embodiment of the present application, Figure 5A Step 1033 shown may be performed by Figure 5B Steps 10331 to 10335 shown are implemented by combining Figure 5B The steps shown are explained.

[0123] In step 10331, the second spectrum map is reduced to a fixed size, and the reduced second spectrum map is converted into a second grayscale image.

[0124] It should be noted that the implementation method of step 10331 of reducing the second spectrum map to a fixed size and converting the reduced second spectrum map into a second grayscale image is similar to the implementation method of step 10231. For details, please refer to the implementation method of the above step 10231, which will not be repeated here.

[0125] In step 10332, a discrete cosine transform is performed on the second grayscale image to obtain a second discrete cosine transform coefficient matrix.

[0126] It should be noted that the implementation method of step 10332 of performing discrete cosine transform on the second grayscale image to obtain the second discrete cosine transform coefficient matrix is ​​similar to the implementation method of step 10232. For details, please refer to the implementation method of the above step 10232, which will not be repeated here.

[0127] In step 10333, the second discrete cosine transform coefficient matrix is ​​reduced, and the mean value of the reduced second discrete cosine transform coefficient matrix is ​​determined.

[0128] It should be noted that the implementation method of step 10333 of reducing the second discrete cosine transform coefficient matrix and determining the mean of the reduced second discrete cosine transform coefficient matrix is ​​similar to the implementation method of step 10233. For details, please refer to the implementation method of the above-mentioned step 10233, which will not be repeated here.

[0129] In step 10334, the multiple elements included in the reduced second discrete cosine transform coefficient matrix are compared with the mean value in sequence, and the values ​​of the multiple elements are updated according to the comparison results.

[0130] It should be noted that the implementation method of step 10334, which compares the multiple elements included in the reduced second discrete cosine transform coefficient matrix with the mean in turn, and updates the values ​​of the multiple elements according to the comparison results, is similar to the implementation method of step 10234. For details, please refer to the implementation method of the above-mentioned step 10234, which will not be repeated here.

[0131] In step 10335, the updated values ​​of the multiple elements are combined into a second string to serve as the second audio fingerprint corresponding to the reference audio data.

[0132] It should be noted that the implementation method of combining the updated values ​​of the multiple elements into a second string as the second audio fingerprint corresponding to the reference audio data in step 10335 is similar to the implementation method of step 10235. For details, please refer to the implementation method of the above-mentioned step 10235 and will not be repeated here.

[0133] In step 104 , based on the comparison result of the first audio fingerprint and the second audio fingerprint, an offset value between the dry audio data and the accompaniment audio data of the target object is determined.

[0134] In some embodiments, determining the offset value between the dry audio data and the accompaniment audio data of the target object based on the comparison result of the first audio fingerprint and the second audio fingerprint can be achieved by: determining the similarity between the first audio fingerprint and the second audio fingerprint; and in response to the similarity satisfying the audio synthesis condition, determining the offset value between the dry audio data and the accompaniment audio data of the target object based on the similarity.

[0135] It should be noted here that different algorithms can be used to compare two audio fingerprints to determine the similarity between the two audio fingerprints. Commonly used audio fingerprint similarity comparison algorithms include Euclidean distance, Pearson correlation coefficient, cosine similarity, etc. These algorithms can effectively measure the degree of difference and consistency between two audio fingerprints.

[0136] In some embodiments, determining the similarity between the first audio fingerprint and the second audio fingerprint can be achieved by dividing the first audio fingerprint into multiple first sub-audio fingerprints, and dividing the second audio fingerprint into multiple second sub-audio fingerprints; and determining, for each first sub-audio fingerprint, the similarity between the first sub-audio fingerprint and the corresponding second sub-audio fingerprint.

[0137] For example, the durations corresponding to the first audio fingerprint and the second audio fingerprint are both 10 seconds. The first audio fingerprint and the second audio fingerprint are evenly divided according to the time dimension to ensure that each sub-audio fingerprint has a duration of 1 second. That is, 10 first sub-audio fingerprints and 10 second sub-audio fingerprints are obtained. The correspondence between the first sub-audio fingerprints and the second sub-audio fingerprints is determined. Assuming that the first first sub-audio fingerprint corresponds to the first second sub-audio fingerprint, the similarity between the two sub-audio fingerprints is determined.

[0138] In some embodiments, determining the similarity between the first audio fingerprint and the second audio fingerprint can also be achieved by: visualizing the first audio fingerprint to obtain a first image; visualizing the second audio fingerprint to obtain a second image; and overlapping and comparing the first image and the second image to determine the similarity between the first audio fingerprint and the corresponding second audio fingerprint.

[0139] It should be noted that when performing overlapping comparison on images, manual comparison can be adopted, or automatic comparison can be achieved through an algorithm, which is not specifically limited here.

[0140] In some embodiments, in response to the similarity satisfying the audio synthesis condition, the offset value between the dry audio data and the accompaniment audio data of the target object based on the similarity is determined, which can be achieved in the following way: based on the maximum similarity and the minimum similarity among multiple similarities, the similarity range is determined; in response to the maximum similarity being greater than a first threshold, the similarity range being greater than a second threshold and the offset value corresponding to the maximum similarity being greater than a third threshold, it is determined that the audio synthesis condition is satisfied, and the offset value corresponding to the maximum similarity is used as the offset value between the dry audio data and the accompaniment audio data of the target object.

[0141] For example, continuing with the above example, the similarity of 10 first sub-audio fingerprints and 10 second sub-audio fingerprints is calculated to obtain 10 audio similarities, and the similarity range and the offset values ​​corresponding to the 10 audio similarities are calculated. A maximum audio similarity is selected for verification. It can be determined whether the maximum similarity, the similarity range and the offset value corresponding to the maximum similarity meet the audio synthesis conditions, and then whether the target synthesized audio can be obtained according to the offset value corresponding to the maximum audio similarity. If the audio synthesis conditions are met, the offset value corresponding to the maximum similarity is the offset value between the dry sound audio data and the accompaniment audio data of the target object.

[0142] In some embodiments, in response to the similarity satisfying the audio synthesis condition, the offset value between the dry audio data and the accompaniment audio data of the target object based on the similarity is determined, which can also be achieved in the following way: the dry audio data of the target object is divided according to the time dimension corresponding to the first sub-audio fingerprint to obtain multiple first sub-audio data; based on the similarity between each first sub-audio fingerprint and the corresponding second sub-audio fingerprint, the offset value corresponding to each first sub-audio fingerprint is determined.

[0143] In step 105 , the dry audio data and the accompaniment audio data of the target object are synthesized based on the offset value to obtain synthesized audio.

[0144] In some embodiments, the above-mentioned synthesis of the dry sound audio data and the accompaniment audio data of the target object based on the offset value to obtain the synthesized audio can be achieved in the following ways: offsetting the dry sound audio data of the target object based on the offset value to obtain the dry sound audio data of the target object after offset; synthesizing the dry sound audio data and the accompaniment audio data of the target object after offset to obtain the synthesized audio.

[0145] In some embodiments, the above-mentioned synthesis of the dry sound audio data and accompaniment audio data of the target object based on the offset value to obtain the synthesized audio can also be achieved in the following way: according to the offset value corresponding to the first sub-audio fingerprint, the first sub-audio corresponding to the first sub-audio fingerprint is offset to obtain the offset first sub-audio data; and multiple offset first sub-audio data and accompaniment data are synthesized to obtain the synthesized audio.

[0146] The following describes an exemplary application of the embodiment of the present application in a practical application scenario, which describes the specific implementation process of extracting audio fingerprints based on a hash-based algorithm.

[0147] In some embodiments, see Figure 6 , Figure 6 This is a schematic diagram of the process of extracting audio fingerprints based on the perceptual hash algorithm provided by the embodiment of the present application. Figure 6 As shown in the figure, first obtain the audio data to be processed, and perform signal conversion on the audio data to obtain the input audio signal. Then, the amplitude information and phase information corresponding to the audio data are obtained through Fourier transform. Then, the amplitude information and phase information are filtered using a Bark filter, and the amplitude information and phase information corresponding to the audio are displayed in the Bark domain to obtain the corresponding spectrum diagram. Finally, the pHash algorithm is used to process the audio data to obtain the audio fingerprint corresponding to the audio data.

[0148] In practical applications, conventional hash algorithms may cause drastic changes in output due to a small change in the input. However, pHash is different from conventional hash algorithms. If the input is similar, the output results will also be similar. The implementation process of pHash is similar to the process of amplifying the sensitive area of ​​the human ear and then compressing and storing it.

[0149] In some embodiments, see Figure 7 , Figure 7 is a schematic diagram of an enlarged local area of ​​a spectrogram provided in an embodiment of the present application, Figure 7 The spectrogram corresponding to the user's dry voice of a specific song is shown. Spectrogram 702 is an amplified representation of the low-frequency band of spectrogram 701 in the Bark domain and a visualization of the corresponding audio fingerprint. Spectrogram 703 is a further amplified representation of the low-frequency band of spectrogram 702 in the Bark domain and a visualization of the corresponding audio fingerprint. The vertical axis of the spectrogram is a linear representation of frequency. At a sampling rate of 44.1k, Figure 7 As shown in 701, it can be seen that the information content of the person who needs to be paid attention to is very unclear; after that, after processing and mapping the amplitude-frequency information to the Bark domain, the following is obtained: Figure 7As shown in 702, it can be seen that the low frequency (that is, the human voice information that needs attention) is amplified, and the high frequency information is relatively compressed. The human ear's perception of light sound is logarithmic, so the Bark domain can more realistically reflect the real feeling of the human ear; a differential operation is performed on the basis of the Bark domain to further compress the information, retaining only the rising and falling information of the image, which is described by a 32-bit unsigned integer uint32_t. That is to say, the final audio fingerprint is as follows Figure 7 As shown in 703.

[0150] For example, the operation of voice accompaniment alignment is actually to match the audio fingerprint of the user's dry voice with the audio fingerprint of the original singer. Since the users are singing the same song, the alignment result is proportional to the degree of similarity. Figure 8 , Figure 8 This is a trend diagram of the similarity change after song alignment provided by the embodiment of the present application. That is to say, the higher the similarity, the better the alignment effect when users sing the same song. Figure 9 , Figure 9 This is a schematic diagram of the manual calculation of offset and audio fingerprint comparison provided by an embodiment of the present application. The method of manually calculating the offset value and confirming the delay size is to put the comparison audio into the left and right channels, and then find some relatively obvious feature points on the spectrogram (for example: audio turning point, audio highest point, audio lowest point, etc.) to compare the delay size of the two-channel audio; the process of comparing through audio fingerprints is actually simulating the two-dimensional images of two audio fingerprints, dividing the audio fingerprints into multiple sub-audio fingerprints in the time dimension, and then performing image overlapping comparison in units of sub-audio to determine the delay size of each sub-audio, and finally adjusting the sub-audio according to the delay size of each sub-audio, and then synthesizing the adjusted multiple sub-audio with the accompaniment audio to achieve automatic sound accompaniment alignment.

[0151] Compared with other sound accompaniment alignment methods, the advantage of the audio processing method provided in the embodiment of the present application is that it is more robust, grasps the alignment effect from a holistic perspective, and describes the alignment effect by the overall similarity, which is more reliable in the face of complex online scenarios; and the reference object of the audio comparison is the original singer information, which is absolutely credible, and there is no need to worry about the inaccuracy of external dependencies affecting the results. In addition, the audio processing method provided in the embodiment of the present application achieves a 2s offset effect on low-end Android machines. The vast majority of songs can automatically correct the offset problem without the user noticing through the audio processing method provided in the embodiment of the present application.

[0152] The embodiment of the present application provides the use of audio fingerprints to achieve automatic sound companion alignment, uses hash perception to perform visual comparison of audio fingerprints, and adjusts the fingerprints in local areas to achieve accurate matching alignment, thereby improving the accuracy of sound companion alignment while reducing the labor cost of manual adjustment and improving the user experience.

[0153] The following continues to describe the exemplary structure of the audio processing device 555 provided in the embodiment of the present application as a software module. In some embodiments, such as Figure 2 As shown, the software modules stored in the audio processing device 555 of the memory 550 may include: an acquisition module 5551, used to acquire dry audio data and reference audio data of the target object, wherein the reference audio data includes the dry audio data and accompaniment audio data of the original singer; a fingerprint extraction module 5552, used to perform fingerprint extraction on the dry audio data of the target object to obtain a first audio fingerprint corresponding to the dry audio data of the target object; the fingerprint extraction module 5552 is also used to perform fingerprint extraction on the reference audio data to obtain a second audio fingerprint corresponding to the reference audio data; a determination module 5553, used to determine the offset value between the dry audio data and the accompaniment audio data of the target object based on the comparison result of the first audio fingerprint and the second audio fingerprint; a synthesis module 5554, used to synthesize the dry audio data and the accompaniment audio data of the target object based on the offset value to obtain synthesized audio.

[0154] In some embodiments, the fingerprint extraction module 5552 is further used to perform Fourier transform on the dry audio data of the target object to obtain a first spectrogram corresponding to the dry audio data of the target object; filter the first spectrogram to obtain a first frequency spectrum corresponding to the dry audio data of the target object; and perform perceptual hashing on the first frequency spectrum to obtain a first audio fingerprint corresponding to the dry audio data of the target object.

[0155] In some embodiments, the first spectrogram includes first amplitude information and first phase information. The fingerprint extraction module 5552 is further used to call a filter to filter the first amplitude information and the first phase information to obtain a first spectrogram corresponding to the dry sound audio data of the target object.

[0156] In some embodiments, the fingerprint extraction module 5552 is further used to reduce the first spectrogram to a fixed size and convert the reduced first spectrogram into a first grayscale image; perform discrete cosine transform on the first grayscale image to obtain a first discrete cosine transform coefficient matrix; reduce the first discrete cosine transform coefficient matrix and determine the mean of the reduced first discrete cosine transform coefficient matrix; compare the multiple elements included in the reduced first discrete cosine transform coefficient matrix with the mean in turn, and update the values ​​of the multiple elements according to the comparison result; combine the updated values ​​of the multiple elements into a first character string as the first audio fingerprint corresponding to the dry sound audio data of the target object.

[0157] In some embodiments, the fingerprint extraction module 5552 is further used to perform Fourier transform on the reference audio data to obtain a second spectrogram corresponding to the reference audio data; filter the second spectrogram to obtain a second frequency spectrum corresponding to the reference audio data; and perform perceptual hashing on the second frequency spectrum to obtain a second audio fingerprint corresponding to the reference audio data.

[0158] In some embodiments, the second spectrogram includes second amplitude information and second phase information. The fingerprint extraction module 5552 is further used to call a filter to filter the second amplitude information and the second phase information to obtain a second spectrogram corresponding to the reference audio data.

[0159] In some embodiments, the fingerprint extraction module 5552 is further used to reduce the second spectrogram to a fixed size and convert the reduced second spectrogram into a second grayscale image; perform a discrete cosine transform on the second grayscale image to obtain a second discrete cosine transform coefficient matrix; reduce the second discrete cosine transform coefficient matrix and determine the mean of the reduced second discrete cosine transform coefficient matrix; compare multiple elements included in the reduced second discrete cosine transform coefficient matrix with the mean in turn, and update the values ​​of the multiple elements according to the comparison result; combine the updated values ​​of the multiple elements into a second character string to serve as the second audio fingerprint corresponding to the reference audio data.

[0160] In some embodiments, the determination module 5553 is further configured to determine a similarity between the first audio fingerprint and the second audio fingerprint; and in response to the similarity satisfying an audio synthesis condition, determine an offset value between the dry audio data and the accompaniment audio data of the target object based on the similarity.

[0161] In some embodiments, the determination module 5553 is also used to determine the similarity range based on the maximum similarity and the minimum similarity among the multiple similarities; in response to the maximum similarity being greater than the first threshold, the similarity range being greater than the second threshold, and the offset value corresponding to the maximum similarity being greater than the third threshold, it is determined that the audio synthesis condition is met, and the offset value corresponding to the maximum similarity is used as the offset value between the dry audio data and the accompaniment audio data of the target object.

[0162] In some embodiments, the synthesis module 5554 is further used to offset the dry audio data of the target object based on the offset value to obtain the dry audio data of the target object after offset; and to synthesize the dry audio data of the target object after offset and the accompaniment audio data to obtain synthesized audio.

[0163] It should be noted that the description of the device of the embodiment of the present application is similar to the description of the method embodiment above, and has similar beneficial effects as the method embodiment, so it will not be repeated here. Figure 3 、 Figure 4A 、 Figure 4B 、 Figure 5A ,or Figure 5B The present invention should be understood by referring to the description of any one of the accompanying drawings.

[0164] The present invention provides a computer program product comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the computer device to perform the audio processing method described in the present invention.

[0165] The embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, the processor will execute the audio processing method provided by the embodiment of the present application, for example, Figure 3 、 Figure 4A 、 Figure 4B 、 Figure 5A ,or Figure 5B The audio processing method shown.

[0166] In some embodiments, the computer-readable storage medium may be a ferroelectric random access memory (FRAM), ROM, programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic surface memory, optical disk, or compact disc read-only memory (CD-ROM); it may also be various devices including one or any combination of the above memories.

[0167] In some embodiments, executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0168] As an example, executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.

[0169] To summarize, the embodiment of the present application obtains the dry audio data and reference audio data of the target object, and performs fingerprint extraction on the dry audio data and accompaniment audio data of the target object respectively to obtain the corresponding first audio fingerprint and second audio fingerprint. According to the comparison result of the first audio fingerprint and the second audio fingerprint, the offset value between the dry audio data and the accompaniment audio data of the target object is determined, and finally the dry audio data and the accompaniment audio data of the target object are synthesized based on the offset value to obtain synthesized audio. In this way, automatic sound accompaniment alignment is realized, which not only improves the accuracy of sound accompaniment alignment, but also saves the cost of manual adjustment and improves the user experience.

[0170] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.

Claims

1. An audio processing method, characterized in that: The method comprises: Acquire dry audio data and reference audio data of a target object, wherein the reference audio data includes the dry audio data of the original singer and the accompaniment audio data; Performing fingerprint extraction on the dry audio data of the target object to obtain a first audio fingerprint corresponding to the dry audio data of the target object; Extracting a fingerprint from the reference audio data to obtain a second audio fingerprint corresponding to the reference audio data; determining an offset value between the dry audio data of the target object and the accompaniment audio data based on a comparison result of the first audio fingerprint and the second audio fingerprint; The dry audio data of the target object and the accompaniment audio data are synthesized based on the offset value to obtain synthesized audio.

2. The method according to claim 1, characterized in that The extracting a fingerprint from the dry audio data of the target object to obtain a first audio fingerprint corresponding to the dry audio data of the target object includes: Performing Fourier transform on the dry audio data of the target object to obtain a first spectrogram corresponding to the dry audio data of the target object; Filtering the first spectrogram to obtain a first spectrogram corresponding to the dry voice audio data of the target object; Perform perceptual hashing on the first spectrogram to obtain a first audio fingerprint corresponding to the dry audio data of the target object.

3. The method according to claim 2, characterized in that The first spectrogram includes first amplitude information and first phase information; The filtering of the first spectrogram to obtain a first spectrogram corresponding to the dry voice audio data of the target object includes: A filter is called to filter the first amplitude information and the first phase information to obtain a first frequency spectrum corresponding to the dry sound audio data of the target object.

4. The method according to claim 2, characterized in that The performing perceptual hashing on the first spectrogram to obtain a first audio fingerprint corresponding to the dry audio data of the target object includes: reducing the first spectrogram to a fixed size, and converting the reduced first spectrogram into a first grayscale image; Performing a discrete cosine transform on the first grayscale image to obtain a first discrete cosine transform coefficient matrix; reducing the first discrete cosine transform coefficient matrix and determining a mean value of the reduced first discrete cosine transform coefficient matrix; Comparing a plurality of elements included in the reduced first discrete cosine transform coefficient matrix with the mean value in sequence, and updating the values ​​of the plurality of elements according to the comparison results; The updated values ​​of the multiple elements are combined into a first character string to serve as a first audio fingerprint corresponding to the dry audio data of the target object.

5. The method according to claim 1, wherein Extracting the fingerprint of the reference audio data to obtain a second audio fingerprint corresponding to the reference audio data includes: Performing a Fourier transform on the reference audio data to obtain a second spectrogram corresponding to the reference audio data; filtering the second spectrogram to obtain a second spectrogram corresponding to the reference audio data; Perform perceptual hashing on the second spectrogram to obtain a second audio fingerprint corresponding to the reference audio data.

6. The method according to claim 5, characterized in that The second spectrogram includes second amplitude information and second phase information; The filtering the second spectrogram to obtain a second spectrogram corresponding to the reference audio data includes: A filter is called to filter the second amplitude information and the second phase information to obtain a second frequency spectrum corresponding to the reference audio data.

7. The method according to claim 5, characterized in that The performing perceptual hashing on the second spectrogram to obtain a second audio fingerprint corresponding to the reference audio data includes: reducing the second spectrogram to a fixed size, and converting the reduced second spectrogram into a second grayscale image; performing a discrete cosine transform on the second grayscale image to obtain a second discrete cosine transform coefficient matrix; reducing the second discrete cosine transform coefficient matrix and determining a mean value of the reduced second discrete cosine transform coefficient matrix; Comparing a plurality of elements included in the reduced second discrete cosine transform coefficient matrix with the mean value in sequence, and updating the values ​​of the plurality of elements according to the comparison results; The updated values ​​of the multiple elements are combined into a second character string to serve as a second audio fingerprint corresponding to the reference audio data.

8. The method according to claim 1, characterized in that The determining, based on the comparison result of the first audio fingerprint and the second audio fingerprint, an offset value between the dry audio data of the target object and the accompaniment audio data includes: determining a similarity between the first audio fingerprint and the second audio fingerprint; In response to the similarity satisfying an audio synthesis condition, an offset value between the dry audio data of the target object and the accompaniment audio data is determined based on the similarity.

9. The method according to claim 8, characterized in that The determining the similarity between the first audio fingerprint and the second audio fingerprint includes: Dividing the first audio fingerprint into a plurality of first sub-audio fingerprints, and dividing the second audio fingerprint into a plurality of second sub-audio fingerprints; For each of the first sub-audio fingerprints, a similarity between the first sub-audio fingerprint and the corresponding second sub-audio fingerprint is determined.

10. The method according to claim 9, characterized in that In response to the similarity satisfying an audio synthesis condition, determining an offset value between the dry audio data of the target object and the accompaniment audio data based on the similarity includes: Determining a similarity range based on a maximum similarity and a minimum similarity among the plurality of similarities; In response to the maximum similarity being greater than a first threshold, the similarity range being greater than a second threshold, and the offset value corresponding to the maximum similarity being greater than a third threshold, it is determined that the audio synthesis condition is met, and the offset value corresponding to the maximum similarity is used as the offset value between the dry audio data of the target object and the accompaniment audio data.

11. The method according to claim 1, characterized in that The synthesizing the dry audio data of the target object and the accompaniment audio data based on the offset value to obtain synthesized audio includes: offsetting the dry audio data of the target object based on the offset value to obtain the offset dry audio data of the target object; The dry audio data after the target object is offset and the accompaniment audio data are synthesized to obtain synthesized audio.

12. An audio processing device, characterized in that: The device comprises: An acquisition module, configured to acquire dry audio data and reference audio data of a target object, wherein the reference audio data includes the dry audio data of the original singer and the accompaniment audio data; a fingerprint extraction module, configured to extract a fingerprint from the dry audio data of the target object to obtain a first audio fingerprint corresponding to the dry audio data of the target object; The fingerprint extraction module is further configured to extract a fingerprint from the reference audio data to obtain a second audio fingerprint corresponding to the reference audio data; a determination module, configured to determine an offset value between the dry audio data of the target object and the accompaniment audio data based on a comparison result of the first audio fingerprint and the second audio fingerprint; A synthesis module is used to synthesize the dry sound audio data of the target object and the accompaniment audio data based on the offset value to obtain synthesized audio.

13. An electronic device, characterized in that: include: a memory for storing executable instructions; The processor is configured to implement the audio processing method according to any one of claims 1 to 11 when executing the executable instructions stored in the memory.

14. A computer-readable storage medium storing computer-executable instructions, characterized in that: When the computer-executable instructions are executed by a processor, the audio processing method according to any one of claims 1 to 11 is implemented.

15. A computer program product comprising a computer program or computer executable instructions, characterized in that When the computer program or computer executable instructions are executed by a processor, the audio processing method according to any one of claims 1 to 11 is implemented.