Audio transmission synchronization method and device
By cache audio data in the slave device of the USB microphone and processing it within the synchronization time, the problem of mismatch between the audio transmission rate between the slave device and the host is solved, and stable audio data synchronization is achieved.
Patent Information
- Application Number
- CN202110769958.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-07
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-07-07
AI Technical Summary
In the audio transmission of USB microphones, due to the difference in clocks between the slave device and the host, the audio transmission rate is mismatched, resulting in synchronization problems. The synchronization mode in the prior art requires high hardware support, the asynchronous mode support is inconsistent, and the adaptive mode is prone to noise.
By cache audio data in the slave device and processing the cached data when synchronization timing is monitored, non-voice data is added or discarded to match the audio transmission rate between the slave device and the host.
The audio data synchronization between the slave device and the host is realized, avoiding sound lag and noise generation, and improving the stability and consistency of audio transmission.
Smart Images

Figure CN115604514B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of media processing technology, and in particular to an audio transmission synchronization method and device. Background Art
[0002] A universal serial bus (USB) microphone acts as a USB slave device (also known as a USB sound card) and transmits audio to a USB host through the USB audio class (UAC) 1 or UAC2 protocol. Due to the clock difference between the two, the rate at which the USB slave device generates audio and the rate at which the USB host retrieves audio data are inconsistent, so synchronization is required.
[0003] USB sound cards using UAC1 or UAC2 have three audio synchronization modes: synchronous mode, asynchronous mode, and adaptive mode. In synchronous mode, the clocks of the USB host and the USB slave are synchronized with the USB start frame (SOF); in asynchronous mode, the USB slave provides direct or indirect feedback to the USB host, and the USB host adapts to the clock of the USB slave; and in adaptive mode, the USB slave adapts to the clock of the USB host.
[0004] The synchronous mode solution requires the USB host and USB slave hardware to support clock synchronization with USB SOF, which has high hardware requirements, and the accuracy of SOF as a clock source is not high. The asynchronous mode solution has inconsistent support on different systems. The adaptive mode solution is common in which the USB slave adjusts the clock through a phase locked loop (PLL), or randomly drops or supplements data, which is prone to noise. Summary of the invention
[0005] The embodiments of the present disclosure provide an audio transmission synchronization method and apparatus to synchronize the transmission rate between a slave device and a host.
[0006] In a first aspect, a method for synchronizing audio transmission is provided, the method comprising:
[0007] The slave device caches the audio data collected by the slave device at a first rate, wherein the first rate does not match a second rate at which the host obtains the audio data from the slave device;
[0008] The slave device monitors synchronization timing;
[0009] The slave device obtains the size of the cached data;
[0010] The slave device determines that the size of the cached data is within a first set range, and performs a first processing operation on the cached data within the synchronization opportunity, wherein the first rate is less than the second rate; or
[0011] The slave device determines that the size of the cached data is within a second set range, and performs a second processing operation on the cached data within the synchronization opportunity, wherein the first rate is greater than the second rate;
[0012] The first processing operation or the second processing operation is used to match the first rate with the second rate, and the first setting range is smaller than the second setting range.
[0013] In a possible implementation, before the slave device caches the audio data collected by the slave device at the first rate, the method further includes:
[0014] The slave device performs denoising on the collected audio data, wherein the sound amplitude of the non-speech data in the denoised audio data is equal to 0, or the sound amplitude of the non-speech data in the denoised audio data is less than or equal to the set threshold;
[0015] The slave device buffers the audio data collected by the slave device at a first rate, including:
[0016] The slave device caches the audio data after noise elimination.
[0017] In yet another possible implementation, the slave device monitoring the synchronization opportunity includes:
[0018] The slave device detects that the buffered data contains non-voice data, wherein the sound amplitude of the non-voice data is equal to 0, or the sound amplitude of the non-voice data is less than or equal to a set threshold.
[0019] In yet another possible implementation, the slave device performs a first processing operation on the cached data within the synchronization opportunity, including:
[0020] The slave device increases a set value of a first set amount after the non-voice data within the synchronization opportunity.
[0021] In another possible implementation, the first setting range is that the buffered data is less than a second threshold, and the slave device increases a setting value of a first setting amount after the non-voice data, including:
[0022] If the cached data is less than or equal to the first threshold, the slave device increases the set value of the first set number after the cached data, so that the cached data is greater than the first threshold and less than a third threshold; or
[0023] If the cached data is larger than the first threshold value and smaller than the second threshold value, the slave device increases the set value of the first set number after the non-voice data so that the cached data is larger than the second threshold value and smaller than a third threshold value;
[0024] Among them, the first threshold is the minimum retained data length of the cache data, the first threshold is associated with the system operation jitter state, the second threshold is the minimum length of the cache data after increasing the set value, the third threshold is the maximum length of the cache data after increasing the set value, and the third threshold is associated with the maximum allowed delay of the system.
[0025] In yet another possible implementation, the slave device performs a second processing operation on the cached data within the synchronization opportunity, including:
[0026] The slave device discards a second set amount of the non-voice data within the synchronization opportunity.
[0027] In another possible implementation, the second set range is greater than a third threshold, and the slave device discards the second set amount of non-voice data, including:
[0028] The buffered data is greater than the third threshold value and less than the fourth threshold value, and the slave device discards the second set amount of non-voice data in the buffered data; or
[0029] The buffered data is larger than the fourth threshold, and the slave device discards the second set amount of non-voice data in the buffered data, so that the buffered data is larger than the first threshold and smaller than the third threshold;
[0030] The fourth threshold is the maximum retained data length of the cache data.
[0031] In a second aspect, an audio transmission synchronization device is provided, the device comprising:
[0032] a cache module, configured to cache the audio data collected by the slave device at a first rate, wherein the first rate does not match a second rate at which the host obtains the audio data from the slave device;
[0033] A processing module for monitoring synchronization timing;
[0034] The processing module is further used to obtain the size of the cache data;
[0035] The processing module is further configured to determine that the size of the cached data is within a first set range, and perform a first processing operation on the cached data within the synchronization opportunity, wherein the first rate is less than the second rate; or
[0036] The processing module is further configured to determine that the size of the cached data is within a second set range, and perform a second processing operation on the cached data within the synchronization opportunity, wherein the first rate is greater than the second rate;
[0037] The first processing operation or the second processing operation is used to match the first rate with the second rate, and the first setting range is smaller than the second setting range.
[0038] In a possible implementation, the device further includes:
[0039] A noise reduction module, configured to reduce noise on the collected audio data, wherein the sound amplitude of the non-speech data in the audio data after the noise reduction is equal to 0, or the sound amplitude of the non-speech data in the audio data after the noise reduction is less than or equal to the set threshold;
[0040] The cache module is used to cache the audio data after noise elimination.
[0041] In another possible implementation, the processing module is used to detect that the buffered data contains non-voice data, wherein the sound amplitude of the non-voice data is equal to 0, or the sound amplitude of the non-voice data is less than or equal to a set threshold.
[0042] In yet another possible implementation, the processing module is configured to increase a set value of a first set amount after the non-voice data within the synchronization opportunity.
[0043] In another possible implementation, the first set range is that the cached data is smaller than a second threshold;
[0044] The processing module is used for, when the cached data is less than or equal to a first threshold, increasing the set value of the first set number after the cached data, so that the cached data is greater than the first threshold and less than a third threshold; or
[0045] The processing module is further configured to, when the cached data is larger than the first threshold and smaller than the second threshold, increase the set value of the first set number after the non-voice data, so that the cached data is larger than the second threshold and smaller than a third threshold;
[0046] Among them, the first threshold is the minimum retained data length of the cache data, the first threshold is associated with the system operation jitter state, the second threshold is the minimum length of the cache data after increasing the set value, the third threshold is the maximum length of the cache data after increasing the set value, and the third threshold is associated with the maximum allowed delay of the system.
[0047] In yet another possible implementation, the processing module is further configured to discard a second set amount of the non-voice data within the synchronization opportunity.
[0048] In another possible implementation, the second setting range is greater than a third threshold;
[0049] The processing module is further configured to discard the second set amount of non-voice data in the cached data when the cached data is greater than the third threshold and less than a fourth threshold; or
[0050] The processing module is further configured to discard the second set amount of non-voice data in the cached data when the cached data is larger than the fourth threshold, so that the cached data is larger than the first threshold and smaller than the third threshold;
[0051] The fourth threshold is the maximum retained data length of the cache data.
[0052] In a third aspect, an audio transmission synchronization device is provided, including an input device and an output device, and further comprising:
[0053] a processor adapted to implement one or more instructions; and,
[0054] A computer storage medium storing one or more instructions, wherein the one or more instructions are suitable for being loaded by the processor and executed by the method described in the first aspect or any implementation of the first aspect.
[0055] According to a fourth aspect, a computer storage medium is provided, wherein the computer storage medium stores one or more instructions, wherein the one or more instructions are suitable for being loaded by a processor and executed by the method described in the first aspect or any implementation of the first aspect.
[0056] The solution provided by the embodiment of the present disclosure has the following beneficial effects:
[0057] By monitoring the synchronization timing, the cached data is processed without causing sound freeze to the cached audio data during the synchronization timing, thereby achieving audio data synchronization between the slave device and the host. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0059] Figure 1 is a structural schematic diagram of a USB device provided by an embodiment of the present disclosure;
[0060] Figure 2 It is a structural schematic diagram of an audio transmission synchronization device provided by an embodiment of the present disclosure;
[0061] Figure 3 It is a flowchart of an audio transmission synchronization method provided by an embodiment of the present disclosure;
[0062] Figure 4 It is a flowchart of another audio transmission synchronization method provided by an embodiment of the present disclosure;
[0063] Figure 5 is the audio data after noise reduction of the example of the present disclosure;
[0064] Figure 6 It is a flowchart of an audio transmission synchronization method provided by an embodiment of the present disclosure;
[0065] Figure 7 It is a structural diagram of another audio transmission synchronization device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0066] The following will be combined with the drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.
[0067] The present disclosure can be applied to the above-mentioned USB slave device. The audio transmission synchronization device in the present disclosure can be a USB slave device or a part of the USB slave device.
[0068] In one embodiment, Figure 1The USB slave device shown in FIG. 1 includes a processor for logic operations (e.g., a system on chip (SoC)) and one or more microphones. The processor communicates with the integrated circuit through an internal audio bus (inter-IC sound, I 2 The processor includes a USB slave device peripheral, and the audio data is sent from the USB peripheral to the USB host after passing through the processing module.
[0069] When recording, the USB host and the USB slave record at an agreed sampling rate fs. Due to the difference in clocks between the two, the rate at which the USB host obtains audio from the USB peripheral is slightly different from the acquisition rate of the microphone. If the USB host obtains data faster, the USB slave needs to make up the data, and inserting data into the collected audio can easily cause sound stuttering; if the USB host obtains data more slowly, the USB slave accumulates more and more data, which will cause the audio delay to continue to increase. When it increases to a certain extent, some data needs to be discarded, and some key audio may be lost.
[0070] like Figure 2 1 is a schematic diagram of the structure of an audio transmission synchronization device provided by an embodiment of the present disclosure. The device 100 includes a processing module 11 and a buffer module 12, and may also include a noise reduction module 13. The audio transmission synchronization device may be a part of a USB slave device, or may also be a USB slave device. Assuming that the audio transmission synchronization device in this embodiment is a USB slave device, the processing module 11, the buffer module 12 and the noise reduction module 13 may be a part of a processor 101, and the device 100 also includes a microphone 102. The processor 101 may also include an I 2 S or PDM or analog interface 14, USB peripheral 16. Microphone 106 is used to collect audio data from the surrounding environment or external devices and transmit it to the I 2S or PDM or analog interface 14 is transmitted to processor 101. Cache module 12 is used to cache audio data collected by the audio transmission synchronization device at a first rate, wherein the first rate does not match the second rate at which the host obtains audio data from the audio transmission synchronization device. Processing module 11 is used to monitor the synchronization timing, obtain the size of the cached data, determine within the synchronization timing that the size of the cached data is within a first set range, and perform a first processing operation on the cached data, wherein the first rate is less than the second rate; or determine that the size of the cached data is within the second set range, and perform a second processing operation on the cached data within the synchronization timing, wherein the first rate is greater than the second rate, wherein the first processing operation or the second processing operation is used to match the first rate with the second rate.
[0071] Optionally, the denoising module 13 is used to denoise the collected audio data and transmit it to the cache module 12; wherein the sound amplitude of the non-speech data in the denoised audio data is equal to 0, or the sound amplitude of the non-speech data in the denoised audio data is less than or equal to a set threshold;
[0072] The buffer module 12 is used to buffer the audio data after noise elimination.
[0073] Optionally, the processing module 11 is used to detect that the buffered data contains non-voice data, wherein the sound amplitude of the non-voice data is equal to 0, or the sound amplitude of the non-voice data is less than or equal to a set threshold.
[0074] Optionally, the processing module 11 is configured to increase a set value of a first set amount after the non-voice data within the synchronization opportunity.
[0075] Optionally, in another possible implementation, the first set range is that the cached data is smaller than a second threshold;
[0076] The processing module 11 is used for, when the cached data is less than or equal to a first threshold, increasing a set value of a first set number after the cached data, so that the cached data is greater than the first threshold and less than a third threshold; or
[0077] The processing module 11 is further configured to, when the cached data is greater than a first threshold value and less than a second threshold value, increase a first set number of set values after the non-voice data so that the cached data is greater than the second threshold value and less than a third threshold value;
[0078] Among them, the first threshold is the minimum retained data length of the cached data, the first threshold is associated with the system operation jitter state, the second threshold is the minimum length of the cached data after increasing the set value, the third threshold is the maximum length of the cached data after increasing the set value, and the third threshold is associated with the maximum allowed delay of the system.
[0079] In yet another possible implementation, the processing module 11 is further configured to discard a second set amount of non-voice data within the synchronization opportunity.
[0080] In yet another possible implementation, the second setting range is greater than a third threshold;
[0081] The processing module 11 is further configured to discard a second set amount of non-voice data in the cached data when the cached data is greater than a third threshold and less than a fourth threshold; or
[0082] The processing module 11 is further configured to discard a second set amount of non-voice data in the cached data when the cached data is larger than a fourth threshold, so that the cached data is larger than the first threshold and smaller than the third threshold;
[0083] The fourth threshold is the maximum retained data length of the cached data.
[0084] Based on the above audio transmission synchronization device, Figure 3 As shown, the embodiment of the present disclosure provides an audio transmission synchronization method, which may include the following steps:
[0085] S101. A slave device caches audio data collected by the slave device at a first rate.
[0086] The microphone 102 collects audio data in the surrounding environment or from the user. Specifically, the host and the slave device may pre-agreed on a first rate, which is the rate at which the microphone 102 collects audio data. The microphone 102 collects audio data at the first rate.
[0087] In this embodiment, a buffer module 12 is set in the slave device. The buffer module 12 can be any storage device or memory. The size of the buffer module 12 can be set according to actual experience, theoretical values, or the purpose of the slave device. 2 The S or PDM or analog interface 14 transmits the audio data collected by the microphone 102 to the processing module 11, and then the processing module 11 transmits the collected audio data to the buffer module 12, and the buffer module buffers the audio data. In the voice call scenario, the voice emitted by the user is sparse, so the audio data includes voice data and may also include non-voice data.
[0088] Exemplarily, the cache module 12 can cache the audio data collected by the microphone 102 in real time, or the processing module 11 can acquire an integer number of data blocks of audio data and then transmit them to the cache module 12 for storage. The present disclosure does not limit the way in which the cache module 12 stores audio data.
[0089] S102: The slave device monitors synchronization timing.
[0090] The host and the slave device agree that the microphone collects audio data at a first rate. However, due to the difference in clocks between the slave device and the host, the second rate at which the host obtains data from the USB peripheral of the slave device does not match the first rate. There is a slight difference between the first rate and the second rate, which causes the audio transmission between the host and the slave device to be out of sync.
[0091] For example, in the above-mentioned voice call scenario, due to the sparseness of voice, it is difficult for the slave device to find the right time to synchronize the audio transmission.
[0092] In this embodiment, the processing module 11 may monitor the synchronization opportunity and synchronize the audio transmission within the synchronization opportunity.
[0093] Specifically, the processing module 11 may monitor the synchronization opportunity periodically or irregularly. A time window may be pre-set, and the processing module 11 scans the arrival of the synchronization opportunity within the time window.
[0094] S103: Obtain the size of cached data from the device.
[0095] It is understandable that the cache module 12 receives the audio data collected by the microphone 102, and at the same time, the host obtains data from the cache module 12, and the size of the audio data cached in the cache module 12 is constantly changing. When the processing module 11 detects the synchronization opportunity, it obtains the size of the data currently cached in the cache module 12.
[0096] Specifically, the cache module 12 can store audio data in order of storage addresses from large to small or from small to large. The size of the data block corresponding to each storage address is fixed. Writing the collected data into the cache module 12 and the host obtaining the data in the cache module 12 need to pass through the processing module 11. The processing module 11 can obtain the size of the cached data in the current cache module 12 based on the number of writes or reads, the size of the data block each time written or read, etc.
[0097] S104a: The slave device determines that the size of the cached data is within a first set range, and performs a first processing operation on the cached data within a synchronization opportunity, wherein the first rate is less than the second rate.
[0098] When the first rate is less than the second rate, that is, the host obtains data from the peripheral device faster, while the microphone collects audio data relatively slowly, the processing module 11 determines that the size of the cached data is within the first set range. The processing module 11 performs a first processing operation on the cached data without causing sound jamming to the cached audio data, and the first processing operation is used to match the first rate with the second rate. It can be understood that matching the first rate with the second rate here does not mean adjusting the first rate or the second rate, but performing a first processing operation on the cached data, and increasing the cached data without causing sound jamming to the cached audio data, so that the cached data obtained by the host and the data cached by the cache module 12 are balanced.
[0099] S104b: The slave device determines that the size of the cached data is within a second set range, and performs a second processing operation on the cached data within a synchronization opportunity, wherein the first rate is greater than the second rate.
[0100] When the first rate is greater than the second rate, that is, the microphone collects audio data faster, while the host obtains data from the peripheral device relatively slowly, the processing module 11 determines that the size of the cached data is within the second set range. The processing module 11 performs a second processing operation on the cached data without causing sound jamming to the cached audio data, and the second processing operation is used to match the first rate with the second rate. It can be understood that matching the first rate with the second rate here does not mean adjusting the first rate or the second rate, but performing a second processing operation on the cached data, reducing the cached data without causing sound jamming to the cached audio data, so that the cached data obtained by the host and the data cached by the cache module 12 are balanced.
[0101] It can be seen that compared with the existing adaptive mode, the random loss or supplement of data from the slave device is prone to noise. The present disclosure monitors the synchronization timing and processes the cached data without causing sound jamming to the cached audio data during the synchronization timing, which can well solve the audio data synchronization problem between the slave device and the host.
[0102] According to an audio transmission synchronization method provided by an embodiment of the present disclosure, by monitoring the synchronization timing, the cached data is processed within the synchronization timing without causing sound freeze to the cached audio data, thereby achieving audio data synchronization between the slave device and the host.
[0103] like Figure 4 As shown, the embodiment of the present disclosure also provides an audio transmission synchronization method, which may include the following steps:
[0104] S201. The slave device denoises audio data collected by the slave device at a first rate, wherein the first rate does not match a second rate at which the host obtains audio data from the slave device, and the sound amplitude of non-voice data in the denoised audio data is equal to 0, or the sound amplitude of non-voice data in the denoised audio data is less than or equal to a set threshold.
[0105] As mentioned above, in a voice call scenario, the audio data includes voice data and non-voice data. The sound amplitude of the non-voice data before noise reduction is not uniform, and subsequent processing of the non-voice data is prone to mutations, and data splicing of the non-voice data is prone to generate noise. Therefore, the noise reduction module 13 can perform noise reduction on the audio data collected by the microphone. Optionally, the data can be denoised by making the sound amplitude of the non-voice data in the denoised audio data equal to 0, or the sound amplitude of the non-voice data in the denoised audio data is less than or equal to a set threshold. For example, a noise suppressor (NS) can be used to denoise the data. Figure 5 As shown, it is a schematic diagram of the audio data after noise elimination. It can be seen that the sound amplitude of the non-speech data in the audio data after noise elimination is equal to 0, while the speech data remains unchanged. In another implementation, the sound amplitude of the non-speech data in the audio data after noise elimination can also be less than or equal to a set threshold, and other modules can distinguish whether the audio data is non-speech data or speech data based on the set threshold. The set threshold can be set according to experience or experiment.
[0106] S202: Buffer the de-noised audio data from the device.
[0107] After the noise elimination module 13 eliminates the noise of the non-speech data, the noise elimination module 13 can output the noise elimination audio data to the buffer module 12 through the processing module 11. The buffer module 12 buffers the noise elimination audio data.
[0108] S203: The slave device detects that the buffered data contains non-voice data, wherein the sound amplitude of the non-voice data is equal to 0, or the sound amplitude of the non-voice data is less than or equal to a set threshold.
[0109] The host and the slave device agree that the microphone collects audio data at a first rate. However, due to the difference in clocks between the slave device and the host, the second rate at which the host obtains data from the USB peripheral of the slave device does not match the first rate. There is a slight difference between the first rate and the second rate, which causes the audio transmission between the host and the slave device to be out of sync.
[0110] For example, in the above-mentioned voice call scenario, due to the sparseness of voice, it is difficult for the slave device to find the right time to synchronize the audio transmission.
[0111] In this embodiment, the processing module 11 may monitor the synchronization opportunity and synchronize the audio transmission within the synchronization opportunity.
[0112] The audio transmission is synchronized within the synchronization opportunity, and it is necessary not to cause sound jamming to the cached audio data. However, if data is directly inserted into the cached audio data, it is easy to cause sound jamming. Therefore, in this embodiment, the synchronization opportunity can be when the user is not speaking or the surrounding environment is in a silent state. At this time, the microphone collects a section of non-voice data, and the cache module 12 caches the non-voice data. The cache data stored in the cache module 12 contains a section of non-voice data. That is, when the processing module 11 detects that the cache data contains non-voice data, the cache data is processed without causing sound jamming.
[0113] Specifically, the processing module 11 detects audio data with a sound amplitude equal to 0 or a sound amplitude less than or equal to a set threshold, and then determines that the buffered data contains non-voice data.
[0114] S204: Obtain the size of cached data from the device.
[0115] It is understandable that the cache module 12 receives the audio data collected by the microphone 102, and at the same time, the host obtains data from the cache module 12, and the size of the audio data cached in the cache module 12 is constantly changing. When the processing module 11 detects that the cache data contains non-voice data, it obtains the size of the data currently cached in the cache module 12.
[0116] S205a: The slave device determines that the size of the cached data is within a first set range, and increases a set value of a first set amount after the non-voice data within a synchronization opportunity, wherein the first rate is less than the second rate.
[0117] When the first rate is less than the second rate, that is, the host obtains data from the peripheral device faster, while the microphone collects audio data relatively slowly, the processing module 11 determines that the size of the cached data is within the first set range. The processing module 11 increases the set value of the first set number after the non-voice data within the synchronization opportunity to match the first rate with the second rate. The first set number can be determined according to the first rate and the second rate. The set value can be 0, for example, or other values. For example, the set value can be the same as the non-voice data whose sound amplitude after noise elimination by the noise elimination module 13 is less than or equal to the set threshold. It can be understood that matching the first rate with the second rate here does not mean adjusting the first rate or the second rate, but increasing the set value of the first set number after the non-voice data, so that the cached data obtained by the host and the data cached by the cache module 12 are balanced without causing sound jamming to the cached audio data.
[0118] The first setting range is that the cache data is less than the second threshold value L2. The second threshold value L2 is the minimum length of the cache data after the set value is increased. Among them, the cache data less than the second threshold value L2 can be divided into two intervals: the first interval is that the cache data is less than or equal to the first threshold value L1, and the second interval is that the cache data is greater than the first threshold value L1 and less than the second threshold value L2. Among them, the first threshold value L1 is the minimum retained data length of the cache data.
[0119] When the cached data is in the first interval, specifically, Figure 6 As shown, the processing module 11 determines whether the cached data is less than or equal to the first threshold value L1. If the cached data is less than or equal to the first threshold value L1, it indicates that the host retrieves data too quickly, and the data accumulated from the device is too little, and the set value of the first set number can be increased after the cached data. Specifically, the processing module 11 can fill in the non-voice data with 0. Optionally, 0 can be filled after the non-voice data. Filling the non-voice data with 0, since the sound amplitude of the non-voice data itself is equal to 0, or less than or equal to the set threshold, will not cause a sudden change in the sound, and splicing a set number of 0s into the non-voice data will not produce noise.
[0120] When the cached data is in the second interval, that is, if the cached data is greater than the first threshold value L1, the processing module 11 further determines whether the cached data is less than the second threshold value L2. If the cached data is greater than the first threshold value L1 and less than the second threshold value L2, it indicates that the data accumulated from the device is still too little. It is further determined whether the cached data is voice data. If it is voice data, no processing is performed; if the cached data contains non-voice data, the non-voice data is further increased by a set value of the first set quantity. Specifically, 0 can be added after the non-voice data so that the cached data is greater than the second threshold value L1 and less than the third threshold value L3, that is, the cached data is maintained at an appropriate size. Among them, the third threshold value L3 is the maximum length of the cached data after adding 0.
[0121] Optionally, the first threshold L1 can be determined according to the jitter state of the system operation. For example, if the system operation jitter is relatively large and the microphone collection cycle cannot be guaranteed, the first threshold L1 can be set larger; if the system operation jitter is relatively small, the first threshold L1 can be set smaller.
[0122] S205b: The slave device determines that the size of the buffered data is within a second set range, and discards a second set amount of non-voice data within a synchronization opportunity, wherein the first rate is greater than the second rate.
[0123] When the first rate is greater than the second rate, that is, the microphone collects audio data faster, while the host obtains data from the peripheral device relatively slowly, the processing module 11 determines that the size of the cached data is within the second set range. The processing module 11 discards the second set amount of non-voice data without causing sound jamming to the cached audio data. The second processing operation is used to match the first rate with the second rate. The second set amount can be determined based on the first rate and the second rate. It can be understood that matching the first rate with the second rate here does not mean adjusting the first rate or the second rate, but discarding the second set amount of non-voice data, reducing the cached data without causing sound jamming to the cached audio data, so that the cached data obtained by the host and the data cached by the cache module 12 are balanced.
[0124] The second setting range is that the cache data is greater than the third threshold value L3. The cache data greater than the third threshold value L3 can be divided into: a third interval and a fourth interval. The third interval is that the cache data is greater than the third threshold value L3 and less than the fourth threshold value L4; the fourth interval is greater than the fourth threshold value L4. The fourth threshold value is the maximum retained data length of the cache data.
[0125] Specifically, Figure 6 As shown, when the processing module 11 determines that the cached data is in the third interval, that is, the processing module 11 determines that the cached data is greater than the third threshold value L3 and less than the fourth threshold value L4, it indicates that the cache module 12 of the slave device has accumulated too much data, and then further determines whether the cached data contains non-voice data. If non-voice data is contained, the second set amount of non-voice data in the detected cached data is discarded.
[0126] like Figure 6 As shown, when the processing module 11 determines that the cached data is in the fourth interval, that is, the processing module 11 determines that the cached data is greater than the fourth threshold value L4, it indicates that the cached data in the current cache module 12 is too large, and the data accumulated from the device is too much, then the second set amount of data in the cached data is discarded, so that the cached data is greater than the first threshold value L1 and less than the third threshold value L3. Optionally, the second set amount of data can be a segment of audio data that has been cached recently. Optionally, the discarded data can be non-voice data or voice data.
[0127] Optionally, the cached data in the cache module 12 cannot be increased blindly, which will cause data transmission delay. Therefore, the fourth threshold L4 can be determined according to the maximum allowed delay of the system, and the fourth threshold L4 can be determined according to the longest time the system allows the audio data to be cached in the cache module 12.
[0128] It can be seen that compared with the existing adaptive mode, the random loss or supplement of data from the slave device is prone to noise. However, when the present disclosure detects that the user is not speaking or the surrounding environment is in a silent state, that is, when the cached data contains non-voice data, the cached data is processed without causing sound jamming to the cached audio data, which can well solve the problem of audio data synchronization between the slave device and the host.
[0129] According to an audio transmission synchronization method provided by an embodiment of the present disclosure, by detecting that non-voice data is contained in cached data, the size of the cached data is obtained, and the cached data is processed according to the size of the cached data so that the size of the cached data is within a set range, thereby achieving audio data synchronization between a slave device and a host.
[0130] This method can be applied to slave devices to avoid noise or increasing delays due to audio asynchrony, especially to solve the audio synchronization problem of slave devices in voice call scenarios, because the voice in voice call scenarios is sparse, and the processing module 11 can find the right time for synchronization.
[0131] According to another embodiment of the present disclosure, Figure 2 The various units or modules in the audio transmission synchronization device shown can be separately or completely combined into one or several other units to form, or one (some) of the units can be further divided into multiple functionally smaller units to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present disclosure. The above-mentioned units are divided based on logical functions. In actual applications, the functions of one unit can also be implemented by multiple units, or the functions of multiple units can be implemented by one unit. In other embodiments of the present disclosure, the media resource dynamic display device can also include other units. In actual applications, these functions can also be implemented with the assistance of other units, and can be implemented by the collaboration of multiple units.
[0132] According to another embodiment of the present disclosure, a computer program (including program code) capable of executing each step involved in the corresponding method shown in the above method embodiment can be constructed by running on a general computing device such as a computer including a central processing unit (CPU), a random access memory (RAM), a read-only memory (ROM) and other processing elements and storage elements. Figure 2The audio transmission synchronization device shown in and the audio transmission synchronization method of the embodiment of the present disclosure are implemented. The computer program can be recorded on, for example, a computer readable recording medium, and loaded into the above-mentioned computing device through the computer readable recording medium and run therein.
[0133] Based on the description of the above method embodiment and device embodiment, the present disclosure also provides an audio transmission synchronization device. Figure 7 The device at least includes a processor 301, an input device 302, an output device 303, and a computer storage medium 304. The processor 301, the input device 302, the output device 303, and the computer storage medium 304 in the device can be connected via a bus or other means.
[0134] The computer storage medium 304 may be stored in the memory of the device, the computer storage medium 304 is used to store a computer program, the computer program includes program instructions, and the processor 301 is used to execute the program instructions stored in the computer storage medium 304. The processor 301 (or CPU) is the computing core and control core of the device, which is suitable for implementing one or more instructions, and is specifically suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function.
[0135] In one embodiment, the processor 301 described in the embodiment of the present disclosure can be used to load and execute Figure 3 or Figure 4 Method steps in the illustrated embodiment.
[0136] It should be noted that the above units or one or more of the units can be implemented by software, hardware or a combination of the two. When any of the above units or units is implemented by software, the software exists in the form of computer program instructions and is stored in a memory, and the processor can be used to execute the program instructions and implement the above method flow. The processor can be built into a system on chip (SoC) or ASIC, or it can be an independent semiconductor chip. In addition to the core used to execute software instructions for calculation or processing in the processor, it can also further include necessary hardware accelerators, such as field programmable gate arrays (FPGA), programmable logic devices (PLD), or logic circuits that implement dedicated logic operations.
[0137] When the above units or units are implemented in hardware, the hardware can be any one or any combination of a CPU, a microprocessor, a digital signal processing (DSP) chip, a microcontroller unit (MCU), an artificial intelligence processor, an ASIC, a SoC, an FPGA, a PLD, a dedicated digital circuit, a hardware accelerator or a non-integrated discrete device, which can run the necessary software or not rely on the software to execute the above method flow.
[0138] Optionally, the embodiment of the present application further provides a chip system, including: at least one processor and an interface, the at least one processor is coupled to a memory via the interface, and when the at least one processor runs a computer program or instruction in the memory, the chip system executes a method in any of the above method embodiments. Optionally, the chip system may be composed of a chip, or may include a chip and other discrete devices, which is not specifically limited in the embodiment of the present application.
[0139] It should be understood that in the description of the present application, unless otherwise specified, " / " indicates that the objects associated before and after are in an "or" relationship, for example, A / B can represent A or B; wherein A and B can be singular or plural. Also, in the description of the present application, unless otherwise specified, "multiple" refers to two or more than two. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, wherein a, b, c can be single or multiple. In addition, in order to facilitate the clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, the words "first", "second", etc. are used to distinguish the same items or similar items with substantially the same functions and effects. Those skilled in the art can understand that the words "first", "second", etc. do not limit the quantity and execution order, and the words "first", "second", etc. do not limit them to be necessarily different. Meanwhile, in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete manner for ease of understanding.
[0140] The disclosed embodiment also provides a computer storage medium (memory), which is a memory device in the device for storing programs and data. It is understandable that the computer storage medium here can include both the built-in storage medium in the device and the extended storage medium supported by the device. The computer storage medium provides a storage space, which stores the operating system of the device. In addition, one or more instructions suitable for being loaded and executed by the processor 301 are also stored in the storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that the computer storage medium here can be a high-speed RAM memory, or a non-volatile memory (non-volatile memory), such as at least one disk storage; optionally, it can also be at least one computer storage medium located away from the aforementioned processor.
[0141] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0142] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the division of the unit is only a logical function division, and there may be other division methods in actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. The mutual coupling, direct coupling, or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0143] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0144] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from a website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (digital subscriber line, DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that contains one or more available media integrations. The available medium may be a read-only memory (ROM), or a random access memory (RAM), or a magnetic medium, such as a floppy disk, a hard disk, a tape, a magnetic disk, or an optical medium, such as a digital versatile disc (DVD), or a semiconductor medium, such as a solid state disk (SSD), etc.
Claims
1. An audio transmission synchronization method, characterized in that: The method comprises: The slave device caches the audio data collected by the slave device at a first rate, wherein the first rate does not match a second rate at which the host obtains the audio data from the slave device; The slave device monitors synchronization timing; The size of cached data obtained from the device; The slave device determines that the size of the cached data is within a first set range, and performs a first processing operation on the cached data within the synchronization opportunity, wherein the first rate is less than the second rate; The slave device determines that the size of the cached data is within a second set range, and performs a second processing operation on the cached data within the synchronization opportunity, wherein the first rate is greater than the second rate; The first processing operation and the second processing operation are used to match the first rate with the second rate, the first setting range is less than a second threshold, and the second setting range is greater than a third threshold; The slave device determines that the size of the cached data is within a first set range, and performs a first processing operation on the cached data within the synchronization opportunity, including: The slave device determines that the size of the cached data is within a first set range, and the cached data is less than or equal to a first threshold, then the slave device adds a set value of a first set number after the cached data, so that the cached data is greater than the first threshold and less than a third threshold, wherein the first threshold is the minimum retained data length of the cached data, the first threshold is associated with the system operation jitter state, the second threshold is the minimum length of the cached data after the set value is increased, the third threshold is the maximum length of the cached data after the set value is increased, the third threshold is associated with the maximum allowable delay of the system, the first threshold is less than the second threshold, and the second threshold is less than the third threshold; The slave device determines that the size of the cached data is within a first set range, and the cached data is greater than the first threshold and less than the second threshold, then the slave device increases the set value of the first set number after the non-voice data included in the cached data, so that the cached data is greater than the second threshold and less than a third threshold.
2. The method according to claim 1, characterized in that Before the slave device caches the audio data collected by the slave device at the first rate, the method further includes: The slave device performs denoising on the collected audio data, wherein the sound amplitude of the non-speech data in the denoised audio data is equal to 0, or the sound amplitude of the non-speech data in the denoised audio data is less than or equal to a set threshold; The slave device buffers the audio data collected by the slave device at a first rate, including: The slave device caches the de-noised audio data.
3. The method according to claim 1, characterized in that The slave device monitors synchronization timing, including: The slave device detects that the buffered data contains non-voice data, wherein the sound amplitude of the non-voice data is equal to 0, or the sound amplitude of the non-voice data is less than or equal to a set threshold.
4. The method according to claim 1, characterized in that: The slave device performs a second processing operation on the cached data within the synchronization opportunity, including: The slave device discards a second set amount of non-voice data in the buffered data within the synchronization opportunity.
5. The method according to claim 4, characterized in that The slave device discards a second set amount of non-voice data in the buffered data within the synchronization opportunity, including: The buffered data is greater than the third threshold value and less than the fourth threshold value, and the slave device discards the second set amount of non-voice data in the buffered data; The cached data is larger than the fourth threshold, the slave device discards the second set amount of non-voice data in the cached data, so that the cached data is larger than the first threshold and smaller than the third threshold; the third threshold is smaller than the fourth threshold; The fourth threshold is the maximum retained data length of the cache data.
6. An audio transmission synchronization device, characterized in that: The device comprises: a cache module, configured to cache audio data collected from a slave device at a first rate, wherein the first rate does not match a second rate at which a host obtains the audio data from the slave device; A processing module for monitoring synchronization timing; The processing module is further used to obtain the size of cached data; The processing module is further configured to determine that the size of the cached data is within a first set range, and perform a first processing operation on the cached data within the synchronization opportunity, wherein the first rate is less than the second rate; The processing module is further configured to determine that the size of the cached data is within a second set range, and perform a second processing operation on the cached data within the synchronization opportunity, wherein the first rate is greater than the second rate; The first processing operation and the second processing operation are used to match the first rate with the second rate, the first setting range is less than a second threshold, and the second setting range is greater than a third threshold; Wherein, when the processing module is used to determine that the size of the cache data is within the first set range and performs the first processing operation on the cache data within the synchronization opportunity, it is specifically used to: Determining that the size of the cache data is within a first set range and the cache data is less than or equal to a first threshold, then adding a set value of a first set number after the cache data so that the cache data is greater than the first threshold and less than a third threshold, wherein the first threshold is the minimum retained data length of the cache data, the first threshold is associated with the system operation jitter state, the second threshold is the minimum length of the cache data after the set value is increased, the third threshold is the maximum length of the cache data after the set value is increased, the third threshold is associated with the maximum allowable delay of the system, the first threshold is less than the second threshold, and the second threshold is less than the third threshold; Determine that the size of the cached data is within a first set range, and the cached data is greater than the first threshold and less than the second threshold, then increase the set value of the first set number after the non-voice data included in the cached data, so that the cached data is greater than the second threshold and less than a third threshold.
7. An audio transmission synchronization device, characterized in that: Includes input devices and output devices, and also includes: a processor adapted to implement one or more instructions; and, A computer storage medium storing one or more instructions, wherein the one or more instructions are suitable for being loaded by the processor and executing the method according to any one of claims 1 to 5.
8. A computer storage medium, characterized in that: The computer storage medium stores one or more instructions, and the one or more instructions are suitable for being loaded by a processor and executing the method according to any one of claims 1-5.
Citation Information
Patent Citations
Method and device for solving network jitter
CN101119323A
Method for processing acoustic frequency flow playback in network terminal buffer
CN1464685A