An audio data mixing method, a terminal, and a computer-readable storage medium

By obtaining the standard synchronous audio source transmitted by the audio channel, determining the audio channel delay time, and adjusting the queue cache length, the problem of out-of-synchronization of the audio channel transmission time is solved and the mixing effect is improved.

CN114333864BActive Publication Date: 2025-06-20ZHEJIANG DAHUA TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111553710.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-17
Publication Date
2025-06-20
Estimated Expiration
2041-12-17

AI Technical Summary

Technical Problem

In the prior art, the time of transmission of data by various audio channels is not synchronized, resulting in poor mixing effects.

Method used

By obtaining the same audio information transmitted by at least two audio channels, including a standard synchronous audio source, the delay time of the audio channel is determined, and the queue cache length corresponding to each audio channel is determined according to the delay time, so as to synchronize the transmission of audio data.

Benefits of technology

By detecting the delay time of different audio channels, the queue cache length of other audio channels is determined based on the delay time of all audio channels, so that the mixing module can receive audio data transmitted synchronously from different audio channels, thereby improving the mixing effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114333864B_ABST
    Figure CN114333864B_ABST
Patent Text Reader

Abstract

The present invention provides an audio data mixing method, a terminal, and a computer-readable storage medium. The audio data mixing method includes: obtaining the same audio information transmitted by at least two audio channels respectively, where the audio information includes a standard synchronization sound source; determining the delay duration of the audio channels according to the time of receiving the standard synchronization sound source transmitted by the at least two audio channels respectively; and determining the buffer lengths of the queues respectively corresponding to the at least two audio channels according to the delay durations corresponding to the at least two audio channels. This application detects the delay times of different audio channels, and determines the buffer lengths of the queues of other audio channels based on the delay durations of all audio channels, so that the mixing module can receive audio data synchronously transmitted by different audio channels, thereby improving the mixing effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio data processing, and particularly to an audio data mixing method, a terminal and a computer-readable storage medium. Background Art

[0002] Voice mixing is an important part of multimedia conferences. Since the audio sources for mixing come from different devices and pass through different transmission paths, there is a delay between the moment when the audio data collected from each channel actually reaches the mixing module and the moment when the sound in the real world is generated. There may be significant differences in the delays between each channel. Especially for audio data transmitted over a network, due to additional processes such as encoding, network transmission, and decoding compared to analog audio acquisition, its delay will be significantly higher than that of analog audio sources. There may also be different delay performances between different audio acquisition devices due to internal processing procedures and network fluctuations. If the delay differences of each channel of audio are not processed and directly sent to the mixing module, the mixed audio data may have an overlapping sound problem, seriously affecting the mixing effect. Summary of the Invention

[0003] The main technical problem to be solved by the present invention is to provide an audio data mixing method, a terminal and a computer-readable storage medium, so as to solve the problem in the prior art that the data transmission times of each audio channel are out of sync, resulting in poor mixing effects.

[0004] To solve the above technical problem, the first technical solution adopted by the present invention is: to provide an audio data mixing method, the audio data mixing method comprising: obtaining the same audio information respectively transmitted by at least two audio channels, the audio information including a standard synchronization sound source; determining the delay duration of the audio channels according to the times of receiving the standard synchronization sound sources respectively transmitted by at least two audio channels; and determining the buffer lengths of the queues respectively corresponding to at least two audio channels according to the delay durations corresponding to at least two audio channels.

[0005] Among them, determining the delay duration of the audio channels according to the times of receiving the standard synchronization sound sources respectively transmitted by at least two audio channels comprises: identifying the audio information transmitted by the audio channels to determine the time of receiving the standard synchronization sound source; and determining the delay duration of the audio channels according to the time of receiving the standard synchronization sound source.

[0006] Among them, identifying the audio information transmitted by the audio channel and determining the time to receive the standard synchronous sound source includes: caching audio samples for a preset duration; converting the audio samples into an audio spectrum; determining whether the similarity between the audio spectrum and the preset spectrum of the standard synchronous sound source exceeds a preset similarity; if the similarity between the audio spectrum and the preset spectrum of the standard synchronous sound source exceeds the preset similarity, determining the audio samples as the standard synchronous sound source; and determining the time of receiving the audio samples as the time of receiving the standard synchronous sound source.

[0007] Among them, before caching the audio samples for a preset duration, it further includes: obtaining the audio sample points in the audio information; determining whether the amplitude of the audio sample points exceeds a threshold; if the amplitude of the audio sample points exceeds the threshold, starting to cache the audio samples.

[0008] Among them, converting the audio samples into an audio spectrum includes: converting the audio samples into an audio spectrum through fast Fourier transform.

[0009] Among them, determining the delay duration of the audio channel according to the time of receiving the standard synchronous sound source includes: the moment of playing the standard synchronous sound source is the first time; the moment of receiving the standard synchronous sound source transmitted by the audio channel is the second time; the difference between the second time and the first time is the delay duration of the audio channel.

[0010] Among them, after the difference between the second time and the first time is the delay duration of the audio channel, it includes: determining whether the number of times of playing the standard synchronous sound source exceeds a preset number of times; if the number of times of playing the standard synchronous sound source exceeds the preset number of times, determining the average delay duration of the audio channel according to the multiple delay durations obtained from the audio channel.

[0011] Among them, determining the cache lengths of the queues respectively corresponding to at least two audio channels according to the delay durations corresponding to the at least two audio channels includes: selecting the maximum delay duration among the delay durations corresponding to the at least two audio channels; respectively calculating the differences between the other delay durations except the maximum delay duration and the maximum delay duration among the delay durations corresponding to the at least two audio channels to determine the cache lengths of the queues respectively corresponding to the audio channels.

[0012] Among them, respectively calculating the differences between the other delay durations except the maximum delay duration and the maximum delay duration among the delay durations corresponding to the at least two audio channels to determine the cache lengths of the queues respectively corresponding to the audio channels includes: calculating the difference between the delay duration corresponding to the audio channel and the maximum delay duration; determining the cache lengths of the queues respectively corresponding to the audio channels according to the difference and the sampling rate of obtaining the audio information.

[0013] To solve the above technical problems, the second technical solution adopted by the present invention is: to provide a terminal, which includes a memory, a processor, and a computer program stored in the memory and running on the processor, and the processor is used to execute program data to implement the steps in the above audio data mixing method.

[0014] To solve the above technical problems, the third technical solution adopted by the present invention is: to provide a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the above audio data mixing method are implemented.

[0015] The beneficial effect of the present invention is: different from the prior art, an audio data mixing method, a terminal, and a computer-readable storage medium are provided. The audio data mixing method includes: obtaining the same audio information transmitted by at least two audio channels respectively, where the audio information includes a standard synchronization sound source; determining the delay duration of the audio channels according to the time of receiving the standard synchronization sound sources transmitted by at least two audio channels respectively; and determining the buffer lengths of the queues respectively corresponding to at least two audio channels according to the delay durations corresponding to at least two audio channels respectively. By detecting the delay times of different audio channels and determining the buffer lengths of the queues respectively corresponding to other audio channels based on the delay durations of all audio channels, the mixing module can receive the audio data synchronously transmitted by different audio channels, thereby improving the mixing effect. Description of the Drawings

[0016] Figure 1 is a schematic flowchart of the audio data mixing method provided by the present invention;

[0017] Figure 2 is a schematic flowchart of a specific embodiment of the audio data mixing method provided by the present invention;

[0018] Figure 3 is a schematic block diagram of an embodiment of the terminal provided by the present invention;

[0019] Figure 4 is a schematic block diagram of an embodiment of the computer-readable storage medium provided by the present invention. Detailed Embodiments

[0020] The following combines the accompanying drawings of the specification to detail the solutions of the embodiments of the present application.

[0021] In the following description, specific details such as specific system structures, interfaces, and technologies are proposed for the purpose of illustration rather than limitation, so as to thoroughly understand the present application.

[0022] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0023] The terms "first", "second", and "third" in the present application are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", and "third" may explicitly or implicitly include at least one of such features. In the description of the present application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined. All directional indications (such as up, down, left, right, front, back...) in the embodiments of the present application are only used to explain the relative positional relationship and movement conditions between components in a specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indications will also change accordingly. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes unlisted steps or units, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.

[0024] Referring to "embodiment" in this context means that the specific features, structures, or characteristics described in connection with the embodiment may be included in at least one embodiment of the present application. The phrase appears in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments.

[0025] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of the audio data mixing method provided by the present invention. In this embodiment, an audio data mixing method is provided. This method can be applied to the recording and broadcasting products in the education industry, and can also be applied to industries such as conference recording and broadcasting products. The audio data mixing method includes the following steps.

[0026] S11: Obtain the same audio information transmitted by at least two audio channels, where the audio information includes a standard synchronization sound source.

[0027] Specifically, a standard synchronous sound source played by a recording and playing host is collected through at least two audio channels, and the audio channels that collect the standard synchronous sound source transmit audio data to a mixing module in the recording and playing host. Among them, the transmission methods for the audio channels to transmit audio data to the mixing module in the recording and playing host include wired transmission, Bluetooth transmission, WIFI transmission, Zigbee transmission, etc. Since different audio data can transmit the collected standard synchronous sound source to the mixing module in the recording and playing host through different transmission methods, the transmission speeds of different transmission methods are different, and during the transmission process of the standard synchronous sound source, the encoding and decoding methods for the standard synchronous sound source are also different, resulting in a delay in the time when different audio channels transmit to the mixing module, and making it impossible for different audio channels to synchronously transmit the collected standard synchronous sound source to the mixing module.

[0028] S12: Determine the delay duration of the audio channel according to the time when the standard synchronous sound sources respectively transmitted by at least two audio channels are received.

[0029] Specifically, identify the audio information transmitted by the audio channel to determine the time when the standard synchronous sound source is received; according to the time when the standard synchronous sound source is received, determine the delay duration of the audio channel. In a specific embodiment, cache audio samples for a preset duration; convert the audio samples into an audio spectrum; determine whether the similarity between the audio spectrum and a preset spectrum of the standard synchronous sound source exceeds a preset similarity; if the similarity between the audio spectrum and the preset spectrum exceeds the preset similarity, then determine that the audio sample is the standard synchronous sound source; determine the time when the audio sample is received as the time when the standard synchronous sound source is received. Among them, the audio samples are converted into an audio spectrum through fast Fourier transform. In an alternative embodiment, obtain the audio sample points in the audio information; determine whether the amplitude of the audio sample points exceeds a threshold; if the amplitude of the audio sample points exceeds the threshold, then start caching the audio samples.

[0030] In a specific embodiment, the moment when the standard synchronous sound source is sent is the first time; the moment when the standard synchronous sound source transmitted by the audio channel is received is the second time; the difference between the second time and the first time is the delay duration of the audio channel.

[0031] In a specific embodiment, determine whether the number of times the standard synchronous sound source is played exceeds a preset number of times; if the number of times the standard synchronous sound source is played exceeds the preset number of times, then determine the average delay duration of the audio channel according to the multiple delay durations obtained for the audio channel.

[0032] S13: According to the delay durations corresponding to at least two audio channels, respectively determine the cache lengths of the queues respectively corresponding to at least two audio channels.

[0033] Specifically, select the maximum delay duration among the delay durations corresponding to at least two audio channels; calculate the differences between the other delay durations and the maximum delay duration among the delay durations corresponding to at least two audio channels respectively, so as to determine the buffer lengths of the queues corresponding to each audio channel respectively.

[0034] This embodiment provides an audio data mixing method. By obtaining the same audio information transmitted by at least two audio channels respectively, the audio information includes a standard synchronization sound source; according to the time of receiving the standard synchronization sound sources transmitted by at least two audio channels respectively, determine the delay duration of the audio channels; according to the delay durations corresponding to at least two audio channels, determine the buffer lengths of the queues corresponding to at least two audio channels respectively. This application detects the delay times of different audio channels, and determines the buffer lengths of the queues corresponding to other audio channels respectively based on the delay durations of all audio channels, so that the mixing module can receive the audio data transmitted synchronously by different audio channels, thereby improving the mixing effect.

[0035] Please refer to Figure 2 , Figure 2 FIG. is a schematic flowchart of a specific embodiment of the audio data mixing method provided by the present invention. In this embodiment, an audio data mixing method is provided. This method can be applied to the recording and broadcasting products in the education industry, and can also be applied to industries such as conference recording and broadcasting products. The audio data mixing method includes the following steps.

[0036] S201: At least two audio channels simultaneously collect the same standard synchronization sound source.

[0037] Specifically, the recording and broadcasting host emits a standard synchronization sound source, and records the time when the recording and broadcasting host emits the standard synchronization sound source as T0. The standard synchronization sound source emitted by the recording and broadcasting host is simultaneously collected through at least two audio channels. Among them, the number of audio channels can be two or more. Among them, the audio channels can be pickups, analog high-definition cameras, network cameras (IP Cameras, IPCs) and / or network dome cameras, or other audio acquisition devices, which are not limited here. The standard synchronization sound source is a beep with a fixed frequency and a fixed length, or other audio data. When the recording and broadcasting host emits the standard synchronization sound source, it is necessary to keep the environment quiet and play the standard synchronization sound source at a relatively high volume, so as to facilitate the subsequent recognition of the collected audio information.

[0038] S202: Obtain the audio sample points in the audio information.

[0039] Specifically, when the recording and broadcasting host receives the audio information sent by the audio channel, it obtains the audio sample points in the audio information in real time.

[0040] S203: Determine whether the amplitude of the audio sample point exceeds the threshold.

[0041] Specifically, first screen the audio information to filter out the audio before the standard synchronization sound source, thereby reducing the workload of converting audio samples into audio spectra subsequently. By comparing the audio sample points in the acquired audio information with a threshold value, it is then determined whether the audio sample points exceed the threshold value. Among them, the threshold value is determined according to the amplitude of the standard synchronization sound source to distinguish the standard synchronization sound source from noise.

[0042] If the audio sample points in the audio information exceed the threshold value, directly jump to step S204; if the audio sample points in the audio information do not exceed the threshold value, directly jump to step S202.

[0043] S204: Start caching audio samples.

[0044] If the amplitude of the audio sample points exceeds the threshold value, start caching audio samples. Specifically, the recording and playing host receives the audio information transmitted through the audio channel, and the recording and playing host caches the audio samples of a preset duration in real time. Among them, the audio samples can at least include part of the standard synchronization sound source.

[0045] S205: Convert the audio samples into audio spectra.

[0046] Specifically, in order to facilitate the comparison of the acquired audio samples with the standard synchronization sound source and more conveniently identify the audio samples, the audio samples are converted into audio spectra through fast Fourier transform.

[0047] S206: Determine whether the similarity between the audio spectrum and the preset spectrum of the standard synchronization sound source exceeds the preset similarity.

[0048] Specifically, compare the audio spectrum converted from the audio samples with the preset spectrum of the standard synchronization sound source, and then determine whether the acquired audio samples are the standard synchronization sound source or whether the acquired audio samples include the standard synchronization sound source. In a specific embodiment, by comparing whether the laws presented by the amplitudes at the frequency points of the audio spectrum converted from the audio samples and the amplitudes at the frequency points of the preset spectrum of the standard synchronization sound source are the same. It is also possible to perform the comparison through other mathematical methods, which are not limited here as long as the comparison between the audio spectrum and the preset spectrum can be achieved.

[0049] If the similarity between the audio spectrum and the preset spectrum of the standard synchronization sound source exceeds the preset similarity, directly jump to step S207; if the similarity between the audio spectrum and the preset spectrum of the standard synchronization sound source does not exceed the preset similarity, directly jump to step S202.

[0050] S207: Determine that the audio sample is the standard synchronization sound source; and determine the time when the audio sample is received as the time when the standard synchronization sound source is received.

[0051] Specifically, if the similarity between the audio spectrum and the preset spectrum of the standard synchronous sound source exceeds the preset similarity, it indicates that the audio sample is the standard synchronous sound source; and the reception time of the received audio sample is determined as the time T1 when the standard synchronous sound source is received, and then the time when the standard synchronous sound source collected by the audio channel is transmitted to the mixing module is obtained.

[0052] S208: Determine the delay duration of the audio channel according to the time when the standard synchronous sound source is sent by the received audio channel.

[0053] Specifically, the moment when the recording and playback host sends the standard synchronous sound source is the first time, recorded as T0; the moment when the standard synchronous sound source is received is the second time, recorded as T1; the delay duration DT of the audio channel is determined according to the difference between the second time and the first time, that is, DT = T1 - T0.

[0054] S209: Determine whether the number of times of playing the standard synchronous sound source exceeds the preset number of times.

[0055] Specifically, in order to improve the detection accuracy of the delay duration when the audio channel transmits audio information, the standard synchronous sound source needs to be sent by the recording and playback host multiple times, that is, repeat the above steps S201 to S208, and multiple delay durations need to be obtained for the same audio channel. And accumulate the number of times of repeating the above steps, and determine whether the number of times of repeating the above steps exceeds the preset number of times.

[0056] If the number of times of repeating the above steps exceeds the preset number of times, directly jump to step S210; if the number of times of repeating the above steps does not exceed the preset number of times, directly jump to step S201.

[0057] S210: Determine the average delay duration of the audio channel according to the multiple delay durations obtained by the audio channel.

[0058] Specifically, in order to improve the stability of the delay duration of the audio channel transmitting audio information, if the number of times of sending the standard synchronous sound source exceeds the preset number of times, the average delay duration of the audio channel is determined according to the multiple delay durations obtained by the audio channel. That is, calculate the average delay duration of the multiple delay durations corresponding to the same audio channel.

[0059] At this time, the average delay durations corresponding to different audio channels are obtained.

[0060] S211: Select the maximum delay duration among the delay durations corresponding to at least two audio channels.

[0061] Specifically, the average delay duration corresponding to different audio channels is obtained through the above steps. The average delay durations corresponding to different audio channels are sorted from largest to smallest, and the largest average delay duration is selected as the maximum delay duration corresponding to multiple audio channels. For example, the sorting of the maximum delay durations of each audio channel is as follows: DavgT0 (channel 0), DavgT1 (channel 1), …, DavgT n (channel n), and the maximum channel delay time among them is taken: DT max = Max{DavgT0, DavgT1, …, DavgT n} and the corresponding audio channel number is recorded as m.

[0062] S212: Calculate the differences between the other delay durations except the maximum delay duration and the maximum delay duration among the delay durations corresponding to at least two audio channels respectively, so as to determine the buffer lengths of the queues corresponding to each audio channel.

[0063] Specifically, adjust the buffer lengths of the queues corresponding to the other audio channels according to the maximum delay duration. That is, the delay durations of the other audio channels need to be kept consistent with the delay duration of audio channel m. Calculate the differences T i = DT max - DavgT i , where i is the audio channel number. Among them, the buffer length of the queue represents the time length of the audio samples that need to be buffered in the audio channel. In this embodiment, the audio samples can be blank data, that is, empty audio data.

[0064] According to the difference between the delay duration corresponding to the audio channel and the maximum delay duration and the sampling rate of obtaining the audio information, determine the buffer length L i of the queue of the audio channel, that is, L i = T i * F. Among them, L i is the number of audio sample points that need to be buffered, L i takes a positive integer, i is the audio channel number, and F is the sampling rate of the audio channel for obtaining audio information.

[0065] This embodiment provides an audio data mixing method, which acquires the same audio information transmitted by at least two audio channels respectively, and the audio information includes a standard synchronization sound source; determines the delay duration of the audio channels according to the time of receiving the standard synchronization sound sources transmitted by at least two audio channels respectively; and determines the buffer lengths of the queues corresponding to at least two audio channels respectively according to the delay durations corresponding to at least two audio channels. By detecting the delay times of different audio channels and determining the buffer lengths of the queues corresponding to other audio channels based on the delay durations of all audio channels, this application enables the mixing module to receive audio data synchronously transmitted by different audio channels, thereby improving the mixing effect.

[0066] Refer to Figure 3 , Figure 3 FIG. is a schematic block diagram of an embodiment of a terminal provided by the present invention. The terminal 70 in this embodiment includes: a processor 71, a memory 72, and a computer program stored in the memory 72 and executable on the processor 71. When the computer program is executed by the processor 71, it implements the above audio data mixing method. To avoid repetition, details are not described herein one by one.

[0067] Refer to Figure 4 , Figure 4 FIG. is a schematic block diagram of an embodiment of a computer-readable storage medium provided by the present invention.

[0068] In an embodiment of the present application, a computer-readable storage medium 90 is further provided. The computer-readable storage medium 90 stores a computer program 901, and the computer program 901 includes program instructions. When the processor executes the program instructions, it implements the audio data mixing method provided by the embodiment of the present application.

[0069] Among them, the computer-readable storage medium 90 may be an internal storage unit of the computer device in the foregoing embodiment, such as the hard disk or memory of the computer device. The computer-readable storage medium 90 may also be an external storage device of the computer device, such as a plug-in hard disk equipped on the computer device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc.

[0070] The above are only embodiments of the present invention, and do not limit the patent protection scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be included in the patent protection scope of the present invention by the same token.

Claims

1. An audio data mixing method, characterized in that, The audio data mixing method includes: Obtaining the same audio information transmitted by at least two audio channels respectively, where the audio information includes a standard synchronization sound source; Determining the delay duration of the audio channels according to the time of receiving the standard synchronization sound source transmitted through the at least two audio channels respectively; Determining the buffer lengths of the queues respectively corresponding to the at least two audio channels according to the delay durations corresponding to the at least two audio channels; The determining the buffer lengths of the queues respectively corresponding to the at least two audio channels according to the delay durations corresponding to the at least two audio channels includes: Selecting the maximum delay duration among the delay durations corresponding to the at least two audio channels; Calculating the differences between the other delay durations except the maximum delay duration among the delay durations corresponding to the at least two audio channels and the maximum delay duration respectively, to determine the buffer lengths of the queues respectively corresponding to each audio channel.

2. The audio data mixing method according to claim 1, characterized in that, The determining the delay duration of the audio channels according to the time of receiving the standard synchronization sound source transmitted through the at least two audio channels respectively includes: Identifying the audio information transmitted by the audio channel to determine the time of receiving the standard synchronization sound source; Determining the delay duration of the audio channel according to the time of receiving the standard synchronization sound source.

3. The audio data mixing method according to claim 2, characterized in that, The identifying the audio information transmitted by the audio channel to determine the time of receiving the standard synchronization sound source includes: Caching audio samples for a preset duration; Converting the audio samples into an audio frequency spectrum; Judging whether the similarity between the audio frequency spectrum and the preset frequency spectrum of the standard synchronization sound source exceeds a preset similarity; If the similarity between the audio frequency spectrum and the preset frequency spectrum of the standard synchronization sound source exceeds the preset similarity, determining that the audio sample is the standard synchronization sound source; Determining the time of receiving the audio sample as the time of receiving the standard synchronization sound source.

4. The audio data mixing method according to claim 3, characterized in that, Before the caching audio samples for a preset duration, it further includes: Obtaining the audio sample points in the audio information; Judging whether the amplitude of the audio sample points exceeds a threshold; If the amplitude of the audio sample points exceeds the threshold, starting to cache the audio samples.

5. The audio data mixing method according to claim 3, characterized in that, The converting the audio samples into an audio frequency spectrum includes: Converting the audio samples into an audio frequency spectrum through fast Fourier transform.

6. The audio data mixing method according to claim 2, characterized in that, The determining the delay duration of the audio channel according to the time of receiving the standard synchronization sound source includes: The moment of playing the standard synchronization sound source is the first time; The moment of receiving the standard synchronization sound source transmitted by the audio channel is the second time; The difference between the second time and the first time is the delay duration of the audio channel.

7. The audio data mixing method according to claim 6, characterized in that, After the difference between the second time and the first time is the delay duration of the audio channel, it includes: Judging whether the number of times of playing the standard synchronization sound source exceeds a preset number of times; If the number of times of playing the standard synchronization sound source exceeds the preset number of times, determining the average delay duration of the audio channel according to the multiple delay durations obtained by the audio channel.

8. The audio data mixing method according to claim 1, characterized in that Calculating the differences between the other delay durations except the maximum delay duration among the delay durations corresponding to the at least two audio channels respectively to determine the buffer lengths of the queues respectively corresponding to the audio channels, includes: Calculating the difference between the delay duration corresponding to the audio channel and the maximum delay duration; Determining the buffer lengths of the queues respectively corresponding to the audio channels according to the difference and the sampling rate for obtaining the audio information.

9. A terminal, characterized in that The terminal includes a memory, a processor, and a computer program stored in the memory and running on the processor, and the processor is configured to execute program data to implement the steps in the audio data mixing method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps in the audio data mixing method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Human voice accompaniment alignment method and device

    CN112216259A

  • Multi-channel audio mixing processing method and system, audio mixing processor and storage medium

    CN114512139A

  • Audio data delay transmission method and device, terminal and storage medium

    CN114974251A