Audio synchronization method and apparatus for distributed microphones, and storage medium
By calculating the time difference between multiple microphones and the host, performing time correction and feature information adjustment, the sound delay interference problem in the multi-microphone system is solved, and the synchronization and high-quality output of the audio signal are achieved.
Patent Information
- Application Number
- CN202211296909.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-21
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2042-10-21
AI Technical Summary
In a multi-microphone system, the sound delay interference caused by the different microphone positions affects the auditory effect of the sound.
By calculating the time difference between multiple microphones and the host, time correction and feature information adjustment are performed to synchronize multiple audio signals and normalize them to the audio signal containing the most volume features for filtering and fusion.
The audio signal of the multi-microphone system is output synchronously, which reduces interference and improves the sound reception effect.
Smart Images

Figure CN115631764B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of microphones, and in particular to an audio synchronization method, device and storage medium for distributed microphones. Background Art
[0002] For slightly larger conference rooms, placing multiple audio equipment can better collect the voices of the participants. When there are multiple microphones in the conference room, due to the different positions of the multiple microphones, there will be a delay between the sounds collected by different microphones. If these sounds are played directly, there will be great interference and the auditory effect of the sound will be relatively poor. Summary of the Invention
[0003] The present invention provides an audio synchronization method, device and storage medium for distributed microphones, aiming to improve the sound reception effect of the distributed microphones and enable the sounds of multiple microphones to be output synchronously.
[0004] In a first aspect, the present invention provides an audio synchronization method for distributed microphones, comprising:
[0005] Get audio signals from multiple microphones;
[0006] calculating a time difference between the audio signals of the plurality of microphones and a host receiving the audio signals;
[0007] performing time correction on the audio signals of the plurality of microphones according to the time difference;
[0008] Obtaining feature information of multiple audio signals after time correction;
[0009] The multiple audio signals are adjusted according to the feature information so that the multiple audio signals are normalized to an audio signal containing the most volume features, thereby obtaining a final audio signal.
[0010] In one embodiment, adjusting the multiple audio signals according to the feature information so that the modified multiple audio signals are normalized to an audio signal containing the most volume features to obtain a final audio signal includes:
[0011] Aligning the multiple audio signals according to the characteristic information so that peaks and troughs of the multiple audio signals are at the same time point;
[0012] The gain of other audio signals is adjusted according to the audio signal containing the most loud volume feature in the feature information, so that the multiple audio signals are normalized to the audio signal containing the most loud volume feature to obtain a final audio signal.
[0013] In one embodiment, the feature information includes feature points, lateral distances between feature points, and change trends between feature points;
[0014] The characteristic points include peaks and troughs.
[0015] In one embodiment, after adjusting the modified multiple audio signals according to the feature information so that the modified multiple audio signals are normalized to an audio signal containing the most volume features, the method further includes:
[0016] filtering the plurality of audio signals after normalizing the audio signal to include the most loud volume features;
[0017] Fusion of multiple audio signals after filtering.
[0018] In one embodiment, after adjusting the multiple audio signals according to the feature information so that the multiple audio signals are normalized to the audio signal containing the most volume features, and before obtaining the final audio signal, the method further includes:
[0019] Normalizing the multiple audio signals to an audio signal containing the most volume features to obtain a normalized audio signal, wherein the normalized audio signal includes an overlapping region and a non-overlapping region, wherein the overlapping region is a region where the audio signals have a consistent change trend, and the non-overlapping region is a region where the audio signals have a divergent change trend;
[0020] Determining a starting region signal of the normalized audio signal according to the non-overlapping region;
[0021] If the amplitudes of the end signal of the audio signal obtained in the previous sampling time period and the start signal of the overlapping area of the current sampling time period are continuous, the start area signal is discarded; if the amplitudes of the end signal of the audio signal obtained in the previous sampling time period and the start signal of the overlapping area of the current sampling time period are discontinuous, the start signal of the overlapping area of the current sampling time period is corrected based on the start area signal so that the corrected start signal of the overlapping area of the current sampling time period is continuous with the end signal of the audio signal obtained in the previous sampling time period.
[0022] In one embodiment, the process of correcting the start signal of the overlapping area of the current sampling time period according to the start area signal includes:
[0023] If the amplitude of the start region signal is closer to the end signal of the audio signal obtained in the previous sampling time period than the amplitude of the start signal of the overlapping region of the current sampling time period, then adding the start region signal to the audio signal of the overlapping region of the current sampling time period so that the corrected start signal of the overlapping region of the current sampling time period is continuous with the end signal of the audio signal obtained in the previous sampling time period;
[0024] If the amplitude of the start area signal is further away from the end signal of the audio signal obtained in the previous sampling time period than the amplitude of the start signal of the overlapping area of the current sampling time period, part of the start signal of the overlapping area of the current sampling time period is deleted so that the corrected start signal of the overlapping area of the current sampling time period is continuous with the end signal of the audio signal obtained in the previous sampling time period.
[0025] In one embodiment, before the start region signal obtained in the current sampling time period is corrected based on the end signal of the audio signal obtained in the previous sampling time period, the method further comprises:
[0026] A plurality of change trends of the start region signal are fused to obtain a fused audio signal.
[0027] In one embodiment, fusing the multiple change trends of the start region signal includes:
[0028] One of the intermediate trend, average trend and main trend of the multiple changing trends is taken as the final trend of the multiple changing trends.
[0029] In a second aspect, the present invention further provides an audio synchronization device for distributed microphones, comprising:
[0030] An audio signal acquisition unit, configured to acquire audio signals from multiple microphones;
[0031] a delay calculation unit, configured to calculate a time difference between the audio signals of the plurality of microphones and a host receiving the audio signals;
[0032] a time correction unit, configured to perform time correction on the audio signals of the plurality of microphones according to the time difference;
[0033] A feature acquisition unit, configured to acquire feature information of a plurality of audio signals after time correction;
[0034] The signal adjustment unit is configured to adjust the multiple audio signals according to the feature information so that the multiple audio signals are normalized to an audio signal containing the most volume features, thereby obtaining a final audio signal.
[0035] In a third aspect, the present invention further proposes a computer storage medium, wherein a computer program is stored in the computer storage medium. When the computer program is executed, the audio synchronization method of the distributed microphone described in any one of the above items is implemented.
[0036] The present invention provides an audio synchronization method, device and storage medium for distributed microphones. Based on the time difference between the time when the microphone sends the audio signal and the time when the host receives the audio signal, time correction is performed on the audio signals of multiple microphones at different positions, so that the multiple audio signals tend to be synchronized in time. Then, based on the characteristic information of the multiple audio signals after time correction, the multiple audio signals are normalized to the audio signal containing the most volume characteristics to obtain the final audio signal. The audio signals of the distributed microphones can be output synchronously, and the sound reception effect is good. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. The drawings described below are only drawings corresponding to some embodiments of the present invention. For ordinary technicians in this field, without paying any creative work, they can also obtain drawings of other embodiments based on these drawings.
[0038] Figure 1 This is a flow chart of a method for audio synchronization of distributed microphones in one embodiment of the present invention;
[0039] Figure 2 Schematic diagram of characteristic information of multiple audio signals after time correction in one embodiment of the present invention;
[0040] Figure 3 A schematic diagram of aligning multiple audio signals in one embodiment of the present invention;
[0041] Figure 4 A schematic diagram of gain adjustment for multiple audio signals in one embodiment of the present invention;
[0042] Figure 5 A flowchart of a method for audio synchronization of distributed microphones in another embodiment of the present invention;
[0043] Figure 6 This is a schematic diagram of fusing multiple audio signals in one embodiment of the present invention;
[0044] Figure 7 This is a structural diagram of an audio synchronization device for distributed microphones in one embodiment of the present invention. DETAILED DESCRIPTION
[0045] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0046] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0047] The distributed microphone audio synchronization method and apparatus of the present invention can be used on electronic devices having multiple distributed microphones. Such electronic devices include, but are not limited to, wearable devices, head-mounted devices, healthcare platforms, personal computers, server computers, handheld or laptop devices, mobile devices (such as mobile phones, personal digital assistants (PDAs), media players, etc.), multi-processor systems, consumer electronic devices, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. The electronic device is preferably an audio and video conferencing system having multiple distributed microphones to improve the audio output effect of the audio and video conferencing system.
[0048] See also Figure 1 The present invention provides an audio synchronization method for distributed microphones. In one embodiment, the audio synchronization method includes:
[0049] Step S101: Acquire audio signals from multiple microphones.
[0050] Multiple microphones are distributed microphones or sound receiving devices containing multiple single microphones, which are placed in different locations in the conference room to facilitate the collection of voice information of participants. During the meeting, audio signals are collected once every certain sampling time period.
[0051] Step S102: Calculate the time difference between the audio signals of the plurality of microphones and the host receiving the audio signals.
[0052] Step S103: performing time correction on the audio signals of the multiple microphones according to the time difference.
[0053] Since multiple microphones are collecting sound, when a participant speaks, there will be a certain delay between the voice information collected by microphones at different locations. Based on the time difference between the audio signals of multiple microphones and the host receiving the audio signals, the time of the voice information of multiple microphones can be adjusted to be consistent.
[0054] Specifically, if at a certain sampling moment, one microphone sends an audio signal to the host at time T1 and the host receives the audio signal at time T2, then the time difference between the microphone sending the audio signal to the host and the host receiving the audio signal is T = T2 - T1. Based on this time difference, the audio signal sent by each microphone is delayed by time T to make the audio signals of each microphone consistent in timing.
[0055] Step S104: Acquire feature information of the multiple audio signals after time correction.
[0056] Step S105: Adjust the multiple audio signals according to the feature information so that the multiple audio signals are normalized to the audio signal containing the most volume features, thereby obtaining a final audio signal.
[0057] After time correction is performed on the audio signals sent by multiple microphones, there are still differences in phase and amplitude between the multiple audio signals obtained. If the audio signals are played at this time, the noise will be relatively large.
[0058] Based on the characteristic information of the audio signal, multiple audio signals can be normalized to the audio signal containing the most volume features to obtain the final audio signal. The audio signal containing the most volume features can be inferred to be the microphone closest to the speaker, and the signal collected by it is the clearest and loudest. Using it as a benchmark to normalize other signals can maximize the elimination of the delay interference caused by distributed microphones and optimize the audio effect of distributed microphones.
[0059] Here, normalizing multiple audio signals to the audio signal containing the most loud volume features can be understood as processing other audio signals so that the other audio signals approach the audio signal containing the most loud volume features, thereby ultimately obtaining an audio signal with less interference.
[0060] In one embodiment, step S105 specifically includes:
[0061] Align multiple audio signals based on feature information so that the peaks and troughs of the multiple audio signals are at the same time point;
[0062] The gain of other audio signals is adjusted according to the audio signal containing the most loud volume feature in the feature information, so that the multiple audio signals are normalized to the audio signal containing the most loud volume feature to obtain a final audio signal.
[0063] For details, see Figures 2-4 The feature information includes feature points, lateral distances between feature points, and changing trends between feature points; wherein the feature points include peaks and troughs.
[0064] Based on the peaks, troughs, the horizontal distances between the peaks and troughs, and the changing trends between the peaks and troughs, the multiple audio signals obtained after time correction are aligned so that the peaks and troughs of the multiple audio signals are at the same time point; the characteristic point with the largest peak or trough amplitude at the same time point is used as the loud volume feature, and the audio signal containing the most loud volume features is used as the reference signal to adjust the gain of other audio signals, so that the multiple audio signals are normalized to the audio signal containing the most loud volume features, and the final audio signal is obtained.
[0065] The audio synchronization method of distributed microphones in this embodiment performs time correction on the audio signals of multiple microphones at different positions based on the time difference between the time when the microphone sends the audio signal and the time when the host receives the audio signal, so that the multiple audio signals tend to be synchronized in time, and then normalizes the multiple audio signals to the audio signal containing the most volume features according to the feature information of the multiple audio signals after time correction to obtain the final audio signal. Specifically, the multiple audio signals are aligned according to the feature information, and then the amplitude is adjusted to normalize the multiple audio signals to the audio signal containing the most volume features; the audio signals of the distributed microphones can be output synchronously, the interference caused by the distributed microphones can be reduced, and the sound reception effect is good.
[0066] See also Figure 5 In one embodiment, the audio synchronization method of the distributed microphone of the present application includes:
[0067] Step S201: Acquire audio signals from multiple microphones.
[0068] Step S202: Calculate the time difference between the audio signals of the plurality of microphones and the host receiving the audio signals.
[0069] Step S203: Time-correct the audio signals from the multiple microphones according to the time difference.
[0070] Step S204: Acquire feature information of the multiple audio signals after time correction.
[0071] Step S205: align the multiple audio signals according to the feature information so that the peaks and troughs of the multiple audio signals are at the same time point.
[0072] Step S206: Gain adjustment is performed on other audio signals according to the audio signal containing the largest volume feature in the feature information, so that the multiple audio signals are normalized to the audio signal containing the largest volume feature.
[0073] Step S207: Filter the multiple audio signals normalized to the audio signal containing the most loud volume features.
[0074] Step S208: Fusing the multiple audio signals after filtering to obtain a final audio signal.
[0075] See also Figure 6 , after normalizing the multiple audio signals to the audio signal containing the most loud volume features, filter the multiple audio signals to eliminate noise, and further fuse the multiple audio signals after filtering, so that the multiple audios are fused into one audio signal, completely eliminating the interference caused by the distributed microphones.
[0076] After aligning and gain-adjusting multiple audio signals based on feature information, there may still be multiple audio signals that do not overlap. Such audio signals still contain noise. Therefore, the multiple audio signals can be further filtered and fused to further eliminate interference.
[0077] In one embodiment, see Figure 4 , adjusting the multiple audio signals according to the feature information so that the multiple audio signals are normalized to the audio signal containing the most volume features and before obtaining the final audio signal, further comprising:
[0078] Normalizing the multiple audio signals to an audio signal containing the most volume features to obtain a normalized audio signal, wherein the normalized audio signal includes an overlapping region and a non-overlapping region, wherein the overlapping region is a region where the audio signals have the same change trend, and the non-overlapping region is a region where the audio signals have a divergent change trend;
[0079] determining a start region signal of the normalized audio signal according to the non-overlapping region;
[0080] If the amplitudes of the end signal of the audio signal obtained in the previous sampling time period and the start signal of the overlapping area of the current sampling time period are continuous, the start area signal is discarded; if the amplitudes of the end signal of the audio signal obtained in the previous sampling time period and the start signal of the overlapping area of the current sampling time period are discontinuous, the start signal of the overlapping area of the current sampling time period is corrected according to the start area signal so that the corrected start signal of the overlapping area of the current sampling time period is continuous with the end signal of the audio signal obtained in the previous sampling time period.
[0081] It can be seen that due to the different sampling times of the distributed microphones, the normalized audio signals obtained after alignment and gain adjustment are highly overlapping in region DE. Taking this region as the overlapping region, the region before point D as the starting region signal, and the signal after point E as the ending region signal, the changing trends of the starting region signal and the ending region signal diverge. In order to ensure the continuity between the audio signal obtained in the current sampling time period and the audio signals to be obtained in the previous sampling time period and the next sampling time period, it is necessary to correct the starting signal of the overlapping region of the current sampling time period so that the corrected starting signal of the overlapping region of the current sampling time period is continuous with the ending signal of the audio signal obtained in the previous sampling time period.
[0082] Furthermore, the process of correcting the start signal of the overlapping area of the current sampling time period according to the start area signal includes:
[0083] If the amplitude of the start region signal is closer to the end signal of the audio signal obtained in the previous sampling time period than the amplitude of the start signal of the overlapping region of the current sampling time period, the start region signal is added to the audio signal of the overlapping region of the current sampling time period so that the corrected start signal of the overlapping region of the current sampling time period is continuous with the end signal of the audio signal obtained in the previous sampling time period;
[0084] If the amplitude of the start area signal is further away from the end signal of the audio signal obtained in the previous sampling time period than the amplitude of the start signal of the overlapping area of the current sampling time period, part of the start signal of the overlapping area of the current sampling time period is deleted so that the corrected start signal of the overlapping area of the current sampling time period is continuous with the end signal of the audio signal obtained in the previous sampling time period.
[0085] Specifically, if the amplitude of the end signal of the audio signal obtained in the previous sampling time period is 0, the amplitude interval of the start area signal of the normalized audio signal obtained in the current sampling time period is [0, 1], and the amplitude of the start signal of the overlapping area is 1, then the start area signal is added to the audio signal of the overlapping area of the current sampling time period, so that the corrected start signal of the overlapping area of the current sampling time period is 0; if the amplitude interval of the start area signal of the normalized audio signal obtained in the current sampling time period is [-1, -0.2], and the amplitude of the start signal of the overlapping area is -0.2, then the start signal with an amplitude of the [-0.2, 0] interval in the overlapping area is deleted, so that the start signal of the overlapping area is continuous with the end signal of the audio signal obtained in the previous sampling time period.
[0086] In one embodiment, the method further comprises discarding an end region signal of the normalized audio signal.
[0087] In one embodiment, before the start region signal obtained in the current sampling time period is corrected based on the end signal of the audio signal obtained in the previous sampling time period, the method further comprises:
[0088] Fusing the multiple change trends of the start region signal to obtain a fused audio signal.
[0089] Specifically, fusing the multiple change trends of the start region signal comprises:
[0090] Taking one of the intermediate trend, the average trend, and the main trend of the multiple change trends as the final trend of the multiple change trends. Or other fusion manners, which are not limited in the embodiment.
[0091] When calculating the amplitude of the start region signal, the trend fusion can be performed first to determine the change trend of the start region signal, and then the amplitude interval of the start region signal is calculated and compared with the end signal of the previous sampling time period. The specific trend fusion manner comprises taking one of the intermediate trend, the average trend, and the main trend of the multiple change trends as the final trend of the multiple change trends, thereby improving the accuracy of the calculation of the amplitude of the start region signal.
[0092] In a specific implementation, after the multiple audio signals are aligned, gain-adjusted, and filtered, different fusion processing can be performed on the highly coincident DE region signal, the start region signal, and the end region signal whose change trends are different, so as to maximize the optimization of the signal of the effective region corresponding to the DE region, judge and crop the start region signal, discard the end region signal, and guarantee the integrity and continuity between the audio signals obtained in different time periods, thereby optimizing the quality of the audio signal.
[0093] Referring to Figure 7 The embodiment of the application further provides an audio synchronization device of a distributed microphone, which is characterized by comprising:
[0094] An audio signal acquisition unit 10 is configured to acquire audio signals of multiple microphones.
[0095] A time difference calculation unit 20 is configured to calculate time differences between the audio signals of the multiple microphones and a host receiving the audio signals.
[0096] A time correction unit 30 is configured to correct the audio signals of the multiple microphones according to the time differences.
[0097] A feature acquisition unit 40 is configured to acquire feature information of the multiple audio signals after the time correction.
[0098] The signal adjustment unit 50 is configured to adjust the multiple audio signals according to the feature information so as to normalize the multiple audio signals to an audio signal having the most volume features, thereby obtaining a final audio signal.
[0099] In one embodiment, the signal adjustment unit 50 is specifically configured to:
[0100] Align multiple audio signals based on feature information so that the peaks and troughs of the multiple audio signals are at the same time point;
[0101] The gain of other audio signals is adjusted according to the audio signal containing the most loud volume feature in the feature information, so that the multiple audio signals are normalized to the audio signal containing the most loud volume feature to obtain a final audio signal.
[0102] In one embodiment, the feature information includes feature points, lateral distances between feature points, and change trends between feature points;
[0103] The characteristic points include peaks and troughs.
[0104] In one embodiment, the audio synchronization device for distributed microphones further includes:
[0105] a fusion unit for filtering the plurality of audio signals after normalizing the audio signal containing the most loudness features;
[0106] Fusion of multiple audio signals after filtering.
[0107] In one embodiment, the audio synchronization device for distributed microphones further includes:
[0108] A signal processing unit is started, configured to normalize the multiple audio signals into an audio signal containing the most volume features to obtain a normalized audio signal, wherein the normalized audio signal includes an overlapping region and a non-overlapping region, wherein the overlapping region is a region where the audio signals have a consistent change trend, and the non-overlapping region is a region where the audio signals have a divergent change trend;
[0109] determining a start region signal of the normalized audio signal according to the non-overlapping region;
[0110] If the amplitudes of the end signal of the audio signal obtained in the previous sampling time period and the start signal of the overlapping area of the current sampling time period are continuous, the start area signal is discarded; if the amplitudes of the end signal of the audio signal obtained in the previous sampling time period and the start signal of the overlapping area of the current sampling time period are discontinuous, the start signal of the overlapping area of the current sampling time period is corrected according to the start area signal so that the corrected start signal of the overlapping area of the current sampling time period is continuous with the end signal of the audio signal obtained in the previous sampling time period.
[0111] In one embodiment, the process of correcting the start signal of the overlapping area of the current sampling time period according to the start area signal includes:
[0112] If the amplitude of the start region signal is closer to the end signal of the audio signal obtained in the previous sampling time period than the amplitude of the start signal of the overlapping region of the current sampling time period, the start region signal is added to the audio signal of the overlapping region of the current sampling time period so that the corrected start signal of the overlapping region of the current sampling time period is continuous with the end signal of the audio signal obtained in the previous sampling time period;
[0113] If the amplitude of the start area signal is further away from the end signal of the audio signal obtained in the previous sampling time period than the amplitude of the start signal of the overlapping area of the current sampling time period, part of the start signal of the overlapping area of the current sampling time period is deleted so that the corrected start signal of the overlapping area of the current sampling time period is continuous with the end signal of the audio signal obtained in the previous sampling time period.
[0114] In one embodiment, before correcting the start region signal obtained in the current sampling time period based on the end signal of the audio signal obtained in the previous sampling time period, the method further includes:
[0115] Multiple change trends of the starting area signal are fused to obtain a fused audio signal.
[0116] In one embodiment, fusing multiple change trends of the audio signal in the starting area includes:
[0117] Take one of the intermediate trend, average trend, and main trend of multiple changing trends as the final trend of multiple changing trends.
[0118] The specific process of each unit executing the above corresponding steps has been described in detail in the above method embodiment, and will not be repeated here for the sake of brevity.
[0119] An embodiment of the present invention further provides a computer device, comprising a processor and a memory, wherein the memory stores a computer program, and when the computer program is loaded and executed by the controller, the method steps described in any of the above method embodiments are implemented.
[0120] An embodiment of the present invention further provides a computer storage medium, wherein the computer storage medium stores a computer program. When the computer program is executed, the method steps described in any one of the above method embodiments are implemented.
[0121] In the above embodiments provided by the present invention, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.
[0122] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0123] In addition, the functional units in the various embodiments of the present application may be integrated into a processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or in the form of software functional units. If the integrated units are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium.
[0124] Based on this understanding, the technical solution of this application, or the contributing part, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a mobile terminal, personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of this application. The aforementioned storage medium includes: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc., various media that can store program code.
[0125] In summary, although the present invention has been disclosed as above in terms of preferred embodiments, the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the concept of the technical solution of the present invention, should be covered by the scope of protection of the present invention.
[0126] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
Claims
1. A distributed microphone audio synchronization method, characterized in that: include: Get audio signals from multiple microphones; calculating a time difference between the audio signals of the plurality of microphones and a host receiving the audio signals; performing time correction on the audio signals of the plurality of microphones according to the time difference; Obtaining feature information of multiple audio signals after time correction; Adjusting the multiple audio signals according to the feature information so that the multiple audio signals are normalized to an audio signal containing the most volume features, thereby obtaining a final audio signal; After adjusting the multiple audio signals according to the feature information so that the multiple audio signals are normalized to an audio signal containing the most volume features, and before obtaining the final audio signal, the method further includes: Normalizing the multiple audio signals to an audio signal containing the most volume features to obtain a normalized audio signal, wherein the normalized audio signal includes an overlapping region and a non-overlapping region, wherein the overlapping region is a region where the audio signals have a consistent change trend, and the non-overlapping region is a region where the audio signals have a divergent change trend; Determining a starting region signal of the normalized audio signal according to the non-overlapping region; If the amplitudes of the end signal of the audio signal obtained in the previous sampling time period and the start signal of the overlapping area of the current sampling time period are continuous, the start area signal is discarded; if the amplitudes of the end signal of the audio signal obtained in the previous sampling time period and the start signal of the overlapping area of the current sampling time period are discontinuous, the start signal of the overlapping area of the current sampling time period is corrected based on the start area signal so that the corrected start signal of the overlapping area of the current sampling time period is continuous with the end signal of the audio signal obtained in the previous sampling time period.
2. The audio synchronization method of distributed microphones according to claim 1, characterized in that: The adjusting the multiple audio signals according to the feature information so that the modified multiple audio signals are normalized to an audio signal containing the most volume features to obtain a final audio signal includes: Aligning the multiple audio signals according to the characteristic information so that peaks and troughs of the multiple audio signals are at the same time point; The gain of other audio signals is adjusted according to the audio signal containing the most loud volume feature in the feature information, so that the multiple audio signals are normalized to the audio signal containing the most loud volume feature to obtain a final audio signal.
3. The audio synchronization method of distributed microphones according to claim 1, characterized in that: The feature information includes feature points, lateral distances between feature points, and change trends between feature points; The characteristic points include peaks and troughs.
4. The audio synchronization method of distributed microphones according to claim 1, characterized in that: The method further comprises: adjusting the modified multiple audio signals according to the feature information so that the modified multiple audio signals are normalized to the audio signal containing the most volume features; filtering the plurality of audio signals after normalizing the audio signal to include the most loud volume features; Fusion of multiple audio signals after filtering.
5. The audio synchronization method of distributed microphones according to claim 1, characterized in that: The process of correcting the start signal of the overlapping area of the current sampling time period according to the start area signal includes: If the amplitude of the start region signal is closer to the end signal of the audio signal obtained in the previous sampling time period than the amplitude of the start signal of the overlapping region of the current sampling time period, then adding the start region signal to the audio signal of the overlapping region of the current sampling time period so that the corrected start signal of the overlapping region of the current sampling time period is continuous with the end signal of the audio signal obtained in the previous sampling time period; If the amplitude of the start area signal is further away from the end signal of the audio signal obtained in the previous sampling time period than the amplitude of the start signal of the overlapping area of the current sampling time period, part of the start signal of the overlapping area of the current sampling time period is deleted so that the corrected start signal of the overlapping area of the current sampling time period is continuous with the end signal of the audio signal obtained in the previous sampling time period.
6. The audio synchronization method of distributed microphones according to claim 1, characterized in that: Before correcting the start region signal obtained in the current sampling time period based on the end signal of the audio signal obtained in the previous sampling time period, the method further includes: A plurality of change trends of the start region signal are fused to obtain a fused audio signal.
7. The audio synchronization method of distributed microphones according to claim 6, characterized in that: The fusing of multiple change trends of the starting area signal includes: One of the intermediate trend, average trend and main trend of the multiple changing trends is taken as the final trend of the multiple changing trends.
8. An audio synchronization device for a distributed microphone, characterized in that: include: An audio signal acquisition unit, configured to acquire audio signals from multiple microphones; a delay calculation unit, configured to calculate a time difference between the audio signals of the plurality of microphones and a host receiving the audio signals; a time correction unit, configured to perform time correction on the audio signals of the plurality of microphones according to the time difference; A feature acquisition unit, configured to acquire feature information of a plurality of audio signals after time correction; a signal adjustment unit, configured to adjust the multiple audio signals according to the feature information so that the multiple audio signals are normalized to an audio signal containing the most volume features, thereby obtaining a final audio signal; The signal adjustment unit is specifically configured to normalize the multiple audio signals into an audio signal containing the most volume features to obtain a normalized audio signal, wherein the normalized audio signal includes an overlapping area and a non-overlapping area, wherein the overlapping area is an area where the change trends of the audio signals are consistent, and the non-overlapping area is an area where the change trends of the audio signals diverge; Determining a starting region signal of the normalized audio signal according to the non-overlapping region; If the amplitudes of the end signal of the audio signal obtained in the previous sampling time period and the start signal of the overlapping area of the current sampling time period are continuous, the start area signal is discarded; if the amplitudes of the end signal of the audio signal obtained in the previous sampling time period and the start signal of the overlapping area of the current sampling time period are discontinuous, the start signal of the overlapping area of the current sampling time period is corrected based on the start area signal so that the corrected start signal of the overlapping area of the current sampling time period is continuous with the end signal of the audio signal obtained in the previous sampling time period.
9. A computer storage medium, characterized in that The computer storage medium stores a computer program, and when the computer program is executed, the audio synchronization method of the distributed microphone according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Voice signal de-noising and pickup processing method and apparatus, and refrigerator
CN106710601A
Apparatuses and methods for encoding or decoding a multi-channel audio signal using frame control synchronization
CN108885879A