Audio multitrack synthesis method, video shooting equipment and system based on communication network

By connecting multiple audio acquisition devices and video shooting devices through a wireless communication network, time synchronization and packet transmission of audio information, as well as attenuation synthesis, are achieved. This solves the problems of relying on manual tuning and inconvenient transmission methods in existing technologies, and realizes convenient lossless audio transmission and high sound quality.

CN115527518BActive Publication Date: 2026-04-03GUANGDONG HUAYI CULTURE MEDIA CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-19
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In existing technologies, multi-track audio synthesis relies on manual tuning, and wired and Bluetooth transmission methods are inconvenient for device movement and cannot achieve lossless audio transmission, resulting in poor sound quality.

Method used

Multiple audio acquisition devices and video shooting devices are connected through a wireless communication network to synchronize their time, and the audio information is packaged, attenuated, and synthesized. Lossless audio transmission and synthesis are achieved using the wireless communication network.

Benefits of technology

It enables convenient audio information transmission and lossless audio synthesis, ensuring sound quality, and achieves better sound quality through close-range acquisition, while reducing the maximum level of the audio signal to avoid distortion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115527518B_ABST
    Figure CN115527518B_ABST
Patent Text Reader

Abstract

This invention relates to an audio synchronization method for video recording based on a communication network, comprising: an audio acquisition device packaging audio information into data packets and sending them to a video recording device via a wireless communication network; the video recording device reconstructing each audio acquisition device's data packet into a single audio track; attenuating each audio track signal when the sum of the maximum audio level values ​​exceeds a preset value; and finally, superimposing and synthesizing the audio tracks to obtain a composite audio signal. In this invention, the audio acquisition device transmits audio information to the video recording device via a wireless communication network, facilitating convenient audio transmission and enabling lossless audio transmission to ensure sound quality; using multiple audio acquisition devices allows for close-range acquisition of various sound sources, resulting in better sound quality; and attenuating each audio track signal reduces the maximum level of the composite audio signal, preventing distortion caused by excessively high audio signal levels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of video shooting technology, and relates to an audio multitrack synthesis method, video shooting equipment and system based on a communication network. Background Technology

[0002] During video recording, situations often arise where multiple instruments are playing simultaneously, requiring individual audio capture near each sound source to obtain better sound quality. The captured audio is then combined to achieve a superior recording effect. However, multi-track audio synthesis typically involves manual mixing using a mixing console. The sound engineer amplifies, mixes, distributes, modifies, and processes the multiple input signals based on their own auditory perception to achieve optimal sound quality; this places high demands on the sound engineer. Furthermore, audio capture and video recording equipment usually transmit audio information via wired or Bluetooth transmission. However, wired transmission is inconvenient for equipment movement and suffers from significant signal loss; Bluetooth transmission has a lower speed, cannot transmit lossless audio, and is not suitable for networking. Summary of the Invention

[0003] In view of this, the purpose of the present invention is to provide an audio multitrack synthesis method, video shooting equipment and system based on a communication network.

[0004] To achieve the above objectives, the present invention provides the following technical solution:

[0005] A method for audio multitrack synthesis during video capture based on a communication network includes the following steps:

[0006] S1. Connect multiple audio acquisition devices to a video shooting device through a wireless communication network, and synchronize the multiple audio acquisition devices with the video shooting device respectively;

[0007] S2. While recording video information, the video shooting device sends instructions to each audio acquisition device to acquire audio information.

[0008] S3. Each audio acquisition device collects audio information through audio sampling, and then packages the sampled audio information into a data packet and sends it to the video shooting device through a wireless communication network.

[0009] S4. After receiving the data packets from each audio acquisition device, the video recording device reconstructs each audio acquisition device's data packet into a single audio track; and obtains the maximum audio level of each audio track.

[0010] S5. Calculate the sum of the maximum audio levels of each audio track; if the sum of the maximum audio levels is greater than the preset high-level audio threshold, proceed to step S6; otherwise, proceed to step S7.

[0011] S6. Attenuate the audio signals of each track so that the sum of the maximum audio levels of the attenuated audio signals of each track is less than or equal to the audio high-level threshold.

[0012] S7. Align the timelines of each audio track, and then superimpose and synthesize the audio tracks to obtain a synthesized audio signal. Align the synthesized audio signal with the timeline of the video information, and then synthesize it with the video information to form a video signal with audio information.

[0013] Furthermore, in step S6, the method for attenuating the audio signals of each track includes the following steps:

[0014] S61. Preset the audio attenuation level threshold;

[0015] S62. Determine the maximum audio level after attenuation of each audio signal track;

[0016] S63. Determine the attenuation ratio based on the audio attenuation level threshold, the maximum audio level of the audio signal, and the maximum audio level after attenuation; the formula for determining the attenuation ratio P is:

[0017] P = (MAX' - δ) / (MAX - δ);

[0018] Where MAX' represents the maximum audio level after the audio signal has attenuated; δ represents the audio attenuation level threshold; and MAX represents the maximum audio level of the audio signal.

[0019] S64. Attenuate the portion of each audio track signal that exceeds the audio attenuation level threshold according to the attenuation ratio; the attenuation formula is:

[0020] A' = (A - δ) × P + δ;

[0021] Where A represents the audio signal level before attenuation, and A' represents the audio signal level after attenuation.

[0022] Furthermore, in step S61, the audio attenuation level threshold for each audio track is determined based on the maximum audio level of each track's audio signal; the formula for determining the audio attenuation level threshold δ is:

[0023] δ = MAX × Q;

[0024] Where MAX represents the maximum audio level of the audio signal; Q represents a preset constant, 0 < Q < P.

[0025] Furthermore, in step S62, the method for determining the maximum audio level after attenuation of each audio track includes the following steps:

[0026] S621. Calculate the ratio R of the audio high-level threshold to the sum of the maximum audio levels of each audio track; the calculation formula is:

[0027] R = σ / ΣMAX;

[0028] Where σ represents the audio high-level threshold; MAX represents the maximum audio level of the audio signal;

[0029] ΣMAX represents the sum of the maximum audio levels of all audio signals on each track;

[0030] S622. Calculate the product S of the maximum audio level of each audio track and the ratio mentioned above; the calculation formula is:

[0031] S = MAX × R;

[0032] S623. Determine the maximum audio level MAX' of each track audio signal after attenuation based on the value of the product S, such that MAX' ≤ S.

[0033] Furthermore, an audio low-level threshold is preset, and before superimposing and synthesizing the audio signals of each track, the portion of the audio signal below the audio low-level threshold is removed.

[0034] Furthermore, the wireless communication network is a WIFI communication network, which includes a WIFI router. Both the audio acquisition device and the video shooting device are equipped with WIFI modules, and the audio acquisition device and the video shooting device are respectively connected to the WIFI router through their WIFI modules.

[0035] Furthermore, the audio acquisition device is a surround sound recording device, a high-impedance musical instrument recording device, or a recording device that actively provides phantom power.

[0036] A video shooting device based on a communication network, including

[0037] The video capture module is used to acquire video information through video capture;

[0038] The first wireless communication module is used to connect to the audio acquisition device through a wireless communication network and acquire data packets of audio information sent by the audio acquisition device.

[0039] The master synchronization module is used to keep the audio acquisition device connected to the video shooting device synchronized with the video shooting device in time.

[0040] The attenuation control module is used to obtain the maximum audio level of each audio signal track, and when the sum of the maximum audio levels of each audio signal track is greater than a preset high-level audio threshold, the audio attenuation module attenuates the audio signal of each track.

[0041] An audio attenuation module is used to recover the received audio data packets into audio signals, remove portions of each audio track that are below the low-level audio threshold, and attenuate portions of each audio track that are above the attenuation level threshold when attenuation is required; and

[0042] The audio-video synthesis module is used to align the timelines of the audio signals of each track, and then superimpose and synthesize the audio signals of each track to obtain a synthesized audio signal. Finally, the synthesized audio signal is aligned with the timeline of the video information and synthesized with the video information to form a video signal with audio information.

[0043] Furthermore, the wireless communication network is a WIFI communication network, which includes a WIFI router. Both the audio acquisition device and the video shooting device are equipped with WIFI modules, and the audio acquisition device and the video shooting device are respectively connected to the WIFI router through their WIFI modules.

[0044] A video shooting system based on a communication network includes a video shooting device and multiple audio acquisition devices; the audio acquisition devices include:

[0045] The audio acquisition module is used to acquire audio information through audio sampling and package the acquired audio information into data packets;

[0046] The second wireless communication module is used to connect to the video recording device via a wireless communication network and to send the data packets to the video recording device; and

[0047] The synchronization module is used to keep the audio acquisition device and the video shooting device in time synchronization.

[0048] In this invention, the audio acquisition device transmits audio information to the video shooting device through a wireless communication network. The audio information transmission is convenient and lossless, ensuring sound quality. Using multiple audio acquisition devices allows for close-range acquisition of various sound sources, resulting in better sound quality. By attenuating the high-level portion of each audio track, the maximum level of the synthesized audio signal can be reduced, preventing distortion caused by excessively high audio signal levels while preserving the main characteristics of each audio track. Attached Figure Description

[0049] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein:

[0050] Figure 1 This is a flowchart of a preferred embodiment of the audio multitrack synthesis method for video shooting based on a communication network according to the present invention.

[0051] Figure 2 This is a flowchart for attenuating the audio signals of each track.

[0052] Figure 3 A flowchart illustrating a method for determining the maximum audio level after attenuation of each audio track.

[0053] Figure 4 This is a schematic diagram of audio signal attenuation.

[0054] Figure 5 This is a schematic diagram of a preferred embodiment of the video shooting device based on a communication network according to the present invention.

[0055] Figure 6 This is a schematic diagram of a preferred embodiment of the video shooting system based on a communication network of the present invention. Detailed Implementation

[0056] The following specific examples illustrate the implementation of the present invention. The illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0057] like Figure 1 As shown, a preferred embodiment of the audio multitrack synthesis method for video shooting based on a communication network according to the present invention includes the following steps:

[0058] S1. Connect multiple audio acquisition devices to a video recording device via a wireless communication network, and synchronize the time of each audio acquisition device with the video recording device to ensure that each audio acquisition device is time-synchronized with the video recording device. The audio acquisition devices can be surround sound recording devices, high-impedance instrument recording devices, or recording devices that actively provide phantom power. Since the audio acquisition devices and video recording device are connected via a wireless communication network, setup is convenient. Therefore, multiple audio acquisition devices can be placed near each sound source for audio acquisition. Close-range acquisition yields better sound quality, and the acquired audio is then synthesized to achieve even better sound quality. The wireless communication network is preferably a Wi-Fi network, which includes a Wi-Fi router. Both the audio acquisition devices and the video recording device are preferably equipped with Wi-Fi modules, and each device connects to the Wi-Fi router via its Wi-Fi module. Of course, the wireless communication network can also be a 4G or 5G mobile communication network. The mobile communication network includes a mobile communication base station. Both the audio acquisition device and the video recording device are equipped with a 4G or 5G communication module, and are connected to the mobile communication base station through their respective 4G or 5G communication modules. Using a wireless communication network to transmit audio information not only makes audio transmission convenient and supports simultaneous transmission of multiple audio channels, but also enables lossless audio transmission, ensuring high sound quality.

[0059] S2. While recording video information, the video shooting equipment sends instructions to each audio acquisition device to acquire audio information.

[0060] S3. Each audio acquisition device collects audio information through audio sampling, packages the sampled audio information into data packets, and sends them to the video shooting device via a wireless communication network. Preferably, each audio acquisition device is provided with a transmission buffer, which is a transmission data storage queue; the audio acquisition device stores the sampled data packets in the transmission data storage queue, and sends all data packets stored in the transmission data storage queue to the video shooting device via a wireless communication network. Step S3 may include the following steps:

[0061] S301. Shift the data packets in each storage location of the transmit data storage queue sequentially to the next location. Assuming that the transmit data storage queue previously only stored the first data packet generated by the audio acquisition device in the first storage location, after the second data packet generated by the audio acquisition device, move data packet 1 from the first storage location of the transmit data storage queue to the second storage location, and store data packet 2 in the first storage location of the transmit data storage queue.

[0062] S302. Discard the data packet stored in the last storage position in the transmit data storage queue. When the number of data packets stored in the transmit data storage queue reaches the maximum storage capacity of the transmit data storage queue (i.e., when the last storage position of the transmit data storage queue contains a data packet), the data packet stored in the last storage position will be discarded when the data packets stored in the transmit data storage queue are moved to the next position, so as to free up the first storage position for storing newly generated data packets from the audio acquisition device.

[0063] S303. The newly generated data packets from the audio acquisition device are stored in the first storage location of the transmit data storage queue. This updates the data packets stored in the transmit data storage queue, causing the queue to discard previously stored data packets and cache newly generated data packets.

[0064] S304. All data packets stored in the transmit buffer are transmitted to the video recording device via the wireless communication network. Assuming the transmit data storage queue can store 5 data packets, all 5 data packets stored in the transmit data storage queue will be transmitted when transmitting data packets; thus, each data packet will be transmitted 5 times to avoid packet loss that could prevent the video recording device from receiving the data packet.

[0065] S4. After receiving data packets from each audio acquisition device, the video capturing device reconstructs each data packet from an audio acquisition device into a single audio signal track; and obtains the maximum audio level of each audio signal track. Preferably, the video capturing device forms a receiving buffer for each audio acquisition device. This receiving buffer is a receiving data storage queue. After receiving a data packet, the received data packets are first stored in the receiving data storage queue according to their order in the sending data storage queue of the audio acquisition device. Once the number of data packets stored in the receiving data storage queue reaches a predetermined number, the data packets are sequentially removed from the receiving data storage queue according to the first-in, first-out principle. When a data packet is missing, a storage location corresponding to the missing data packet is reserved in the receiving data storage queue. This step may specifically include the following steps:

[0066] S401. Check if there are any missing data packets stored in each received data storage queue. If there are, proceed to step S402. If there are no missing data packets, proceed to step S403.

[0067] S402. Find the missing data packet in the received data packet data storage queue and store it in the corresponding position in the received data packet data storage queue, then execute step S403; if the missing data packet is not found, then execute step S403 directly.

[0068] S403. Remove the data packet stored at the last storage location in the received data storage queue from the received buffer, and move the data packets at each storage location in the received data storage queue one storage location to the right in sequence.

[0069] S404. Detect whether there is a newly generated data packet in the received data packet (i.e., a data packet generated after the data packet stored in the second storage position of the received data storage queue). If so, store the data packet in the first storage position of the received data storage queue. If not, reserve the first storage position and mark the data packet as missing in that storage position. This ensures that the data packets stored in the transmitted data storage queue and the received data storage queue are consistent.

[0070] Because of the receive buffer that buffers the received data, when packet loss is detected, the missing data packets can be found in the subsequently received data packets, thereby completing the missing data packets and avoiding the impact of data packet loss on sound quality.

[0071] S405. The video shooting device parses the data packets removed from each receiving data storage queue into a single audio signal and obtains the maximum audio level of the audio signal.

[0072] To match and align the waveforms of multi-track audio signals, after recovering the data packets from each audio acquisition device into a single audio track, the following steps can be performed:

[0073] S411. The duration of the matching period is preset, and the envelope of each audio signal track within a matching period is calculated respectively; the duration of the matching period is preferably the duration of the audio information in a data packet.

[0074] S412. Find the time points corresponding to each peak of the envelope of each audio signal track.

[0075] S413. Align the time points of the peak values ​​corresponding to the envelopes of each audio signal track in sequence, and sum the time differences between the peak values ​​of the other corresponding envelopes at this time.

[0076] S414. Find the time point when the sum of time differences is the smallest, which is the peak value of the envelope, and use it as the alignment time point to align the alignment time points of each audio track.

[0077] Finding the alignment time point by summing the time difference using the envelope peaks can make waveform alignment more accurate.

[0078] S5. Calculate the sum of the maximum audio levels of each audio track. If the sum of the maximum audio levels is greater than the preset audio high-level threshold, proceed to step S6; otherwise, proceed to step S7. Since an audio level greater than 0 dB exceeds the maximum allowable range of quantization depth, the synthesized audio will produce clipping noise. Therefore, the audio high-level threshold is generally set to be less than 0 dB. Especially when multiple signals are input, it is necessary to calculate the maximum level that each signal and the synthesized signal can reach based on the high-level threshold and attenuation ratio. If it exceeds 0 dB, the high-level threshold needs to be readjusted.

[0079] S6. Attenuate the audio signals of each track so that the sum of the maximum audio levels of the attenuated audio signals of each track is less than or equal to the audio high-level threshold. For example... Figure 2 As shown, the method for attenuating the audio signals of each track may include the following steps:

[0080] S61. Preset the audio attenuation level threshold. Preferably, the audio attenuation level threshold for each audio track is determined based on the maximum audio level of each track; the formula for determining the audio attenuation level threshold δ is:

[0081] δ = MAX × Q;

[0082] Where MAX represents the maximum audio level of the audio signal; Q represents the preset attenuation constant, 0 < Q < P. By pre-specifying the value of the attenuation constant Q, the audio attenuation level threshold for each audio track can be automatically calculated. Alternatively, a value can be directly specified as the audio attenuation level threshold for each audio track. For example, when the specified audio attenuation level threshold is -10dB, the portion of each audio track above -10dB will be attenuated, eliminating the need to calculate the audio attenuation level threshold for each track separately.

[0083] S62. Determine the maximum audio level after attenuation of each audio track. For example... Figure 3 As shown, the method for determining the maximum audio level after attenuation of each audio track may include the following steps:

[0084] S621. Calculate the ratio R of the audio high-level threshold to the sum of the maximum audio levels of each audio track; the calculation formula is:

[0085] R = σ / ΣMAX;

[0086] Where σ represents the audio high-level threshold; MAX represents the maximum audio level of the audio signal; and ΣMAX represents the sum of the maximum audio levels of all audio tracks. The required attenuation can be determined by calculating the ratio R.

[0087] S622. Calculate the product S of the maximum audio level of each audio track and the ratio mentioned above; the calculation formula is:

[0088] S = MAX × R;

[0089] S623. Determine the maximum audio level MAX' of each track audio signal after attenuation based on the value of the product S, such that MAX' ≤ S.

[0090] S63. Determine the attenuation ratio based on the audio attenuation level threshold, the maximum audio level of the audio signal, and the maximum audio level after attenuation; the formula for determining the attenuation ratio P is:

[0091] P = (MAX' - δ) / (MAX - δ);

[0092] Where MAX' represents the maximum audio level after the audio signal has decayed; δ represents the audio decay level threshold; and MAX represents the maximum audio level of the audio signal.

[0093] When the audio attenuation level threshold δ of each audio track is determined by pre-specifying the value of the attenuation constant Q, the attenuation ratio P of each audio track is the same, and it is only necessary to calculate the attenuation ratio P of one audio track. When a value is directly specified as the audio attenuation level threshold δ of each audio track, it is necessary to calculate the attenuation ratio P of each audio track separately.

[0094] S64. Attenuate the portion of each audio track signal that exceeds the audio attenuation level threshold according to the attenuation ratio; the attenuation formula is:

[0095] A' = (A - δ) × P + δ;

[0096] Where A represents the audio signal level before attenuation, and A' represents the audio signal level after attenuation.

[0097] When attenuating audio, or at the end of attenuation, the audio signal level can suddenly increase or decrease dramatically due to the abrupt occurrence or disappearance of attenuation, generating additional noise. To avoid this, a first buffer time (onset time) can be set when the audio signal begins to attenuate. During this first buffer time, the audio signal transitions from its pre-attenuation level to the expected attenuated level via a continuous, smooth curve. Alternatively, a second buffer time (outcome time) can be set when the audio signal should end attenuation, allowing the audio signal to transition from its attenuated level back to its original, unattenuated level via a continuous, smooth curve. This smoothing curve can be a linear curve, but higher-order or other complex fitting curves can be used for specific needs. For example, if the high-level threshold is set to -6dB, the attenuation ratio is 1 / 2, the first buffer time is 10ms, and the sampling point when the original audio signal reaches -6dB is set as time 0, and the timing is based on the sampling point, if the system sampling frequency is 48000Hz, then 10ms corresponds to 480 sampling points. At the 0th sampling point, the attenuation is 0. At the 1st sampling point (i.e., time 1), the attenuation is -6dB×(1 / 480)×(1 / 2). At the nth sampling point (i.e., time n), the attenuation is -6dB×(n / 480)×(1 / 2). Then at the 480th sampling point (i.e., time 480), which is the end of the first buffer time of 10ms, the attenuation is -6dB×(480 / 480)×(1 / 2), which is the target attenuation of -3dB. If, during the first buffer time, the signal level at a certain sampling point falls below the high-level threshold, a new second buffer time is immediately initiated to recalculate the time; if, during the second buffer time, the signal level at a certain sampling point rises above the high-level threshold, a new first buffer time is immediately initiated to recalculate the time, and so on.

[0098] S7. Align the time axes of each audio track, and then superimpose and synthesize the audio tracks to obtain a composite audio signal. Next, align the composite audio signal with the time axis of the video information to create a video signal with audio information. When aligning the time axes of each audio track, alignment can be achieved using the clocks synchronized between the audio acquisition device and the video shooting device in step S1. Alternatively, if steps S411 to S414 were performed in step S4, the time axes of each audio track can be aligned by matching the time points to achieve higher alignment accuracy. To reduce the impact of background noise, an audio low-level threshold can be preset. This threshold is typically set to -50dB to -30dB. Before superimposing and synthesizing the audio tracks, remove any portions of the audio track below the audio low-level threshold. When removing portions below the audio low-level threshold, a buffer time for activation and deactivation, as described above, can be set to avoid noise caused by sudden transitions from the original signal to the deactivated low-level state or a sudden return to the original signal. For example, the audio low-level threshold can be set to -40dB. This will remove the portion of the audio signal below -40dB from each track, effectively removing background noise. Since audio acquisition equipment is typically placed near the sound source, the level of the useful audio signal will inevitably be greater than -40dB, thus preventing accidental removal. Figure 4 The diagram illustrates the attenuation of an audio signal. Input represents the audio signal before attenuation, and OUT represents the audio signal after attenuation. As can be seen, the portion of the audio signal with a moderate level is not attenuated, while the majority of the audio signal level lies in this region. This allows for better preservation of the characteristics of each audio track in the synthesized audio signal.

[0099] In this embodiment, audio information is transmitted to the video recording device via a wireless communication network. This not only ensures convenient transmission but also enables lossless audio transmission, guaranteeing sound quality. Using multiple audio acquisition devices to capture audio at close range from the sound source yields better sound quality. Attenuating the high-level components of each audio track reduces the maximum level of the synthesized audio signal while preserving the main characteristics of each track. A transmission buffer in the audio acquisition device allows for multiple transmissions of the same audio data packet, overcoming the impact of packet loss in the wireless communication network. A reception buffer in the video recording device promptly detects lost or missing audio data and allows time for re-reception and reconstruction, significantly improving sound quality in situations such as live video streaming.

[0100] This invention also discloses a video shooting device based on a communication network, such as... Figure 5As shown, a preferred embodiment of the video shooting device based on a communication network of the present invention includes a video shooting module, a first wireless communication module, a main synchronization module, an attenuation control module, an audio attenuation module, and an audio-video synthesis module. The video shooting module is used to acquire video information through video recording to facilitate the synthesis of the captured video.

[0101] The first wireless communication module is used to connect to the audio acquisition device via a wireless communication network and acquire data packets of audio information sent by the audio acquisition device. The wireless communication network is preferably a Wi-Fi network, which includes a Wi-Fi router, and the first wireless communication module is a Wi-Fi module connected to the Wi-Fi router. Alternatively, the wireless communication network can also be a 4G or 5G mobile communication network, which includes a mobile communication base station, and the first wireless communication module is either a 4G or 5G communication module connected to the mobile communication base station.

[0102] The master synchronization module is used to ensure that the audio acquisition device connected to the video shooting device maintains time synchronization with the video shooting device. The attenuation control module is used to obtain the maximum audio level of each audio track, and when the sum of the maximum audio levels of each audio track is greater than a preset high-level audio threshold, the audio attenuation module attenuates the audio signals of each track. The audio attenuation module is used to restore the received audio information data packets to audio signals, remove the portions of each audio track that are below the low-level audio threshold, and attenuate the portions of each audio track that are above the attenuation level threshold when attenuation is required. The audio-video synthesis module is used to align the time axes of each audio track, superimpose and synthesize the audio signals of each track to obtain a synthesized audio signal, and then align the synthesized audio signal with the time axis of the video information to synthesize a video signal with audio information.

[0103] This invention also discloses a video shooting system based on a communication network, such as... Figure 6 As shown, a preferred embodiment of the video shooting system based on a communication network of the present invention includes a video shooting device and multiple audio acquisition devices; the audio acquisition devices include an audio acquisition module, a second wireless communication module, and a synchronization module. The audio acquisition module is used to acquire audio information through audio sampling and package the acquired audio information into data packets. The second wireless communication module is used to connect to the video shooting device through a wireless communication network and to send the data packets to the video shooting device. The synchronization module is used to keep the audio acquisition devices and the video shooting device synchronized in time.

[0104] In this embodiment, the audio acquisition device transmits audio information to the video shooting device through a wireless communication network. The audio information transmission is convenient and lossless, ensuring sound quality. Using multiple audio acquisition devices to collect audio from the sound source at close range can achieve better sound quality. By attenuating the high-level portion of each audio track, the maximum level of the synthesized audio signal can be reduced while preserving the main characteristics of each audio track.

[0105] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for audio multitrack synthesis during video recording based on a communication network, characterized in that, Includes the following steps: S1. Connect multiple audio acquisition devices to a video shooting device via a wireless communication network, and synchronize the multiple audio acquisition devices with the video shooting device respectively; S2. While recording video information, the video shooting device sends instructions to each audio acquisition device to acquire audio information. S3. Each audio acquisition device collects audio information through audio sampling, and then packages the sampled audio information into a data packet and sends it to the video shooting device through a wireless communication network. S4. After receiving the data packets from each audio acquisition device, the video recording device reconstructs each audio acquisition device's data packet into a single audio track; and obtains the maximum audio level of each audio track. S5. Calculate the sum of the maximum audio levels of each audio track; If the sum of the maximum audio levels is greater than the preset audio high-level threshold, then proceed to step S6. Otherwise, proceed to step S7; S6. Attenuate the audio signals of each track so that the sum of the maximum audio levels of the attenuated audio signals of each track is less than or equal to the audio high-level threshold. S7. Align the time axes of each audio track, and then superimpose and synthesize the audio tracks to obtain a synthesized audio signal. Align the synthesized audio signal with the time axis of the video information, and then synthesize it with the video information to form a video signal with audio information. In step S6, the method for attenuating the audio signals of each track includes the following steps: S61. Preset the audio attenuation level threshold; S62. Determine the maximum audio level after attenuation of each audio signal track; S63. Determine the attenuation ratio based on the audio attenuation level threshold, the maximum audio level of the audio signal, and the maximum audio level after attenuation; the formula for determining the attenuation ratio P is: P = (MAX' - δ) / (MAX - δ); Where MAX' represents the maximum audio level after the audio signal has attenuated; δ represents the audio attenuation level threshold; and MAX represents the maximum audio level of the audio signal. S64. Attenuate the portion of each audio track signal that is above the audio attenuation level threshold according to the attenuation ratio; the attenuation formula is: A' = (A - δ) × P + δ; Where A represents the audio signal level before attenuation, and A' represents the audio signal level after attenuation.

2. The audio multitrack synthesis method for video shooting based on a communication network according to claim 1, characterized in that: In step S61, the audio attenuation level threshold for each audio track is determined based on the maximum audio level of each track; the formula for determining the audio attenuation level threshold δ is: δ = MAX × Q; Where MAX represents the maximum audio level of the audio signal; Q represents a preset constant, 0 < Q < P.

3. The method for audio multitrack synthesis during video recording based on a communication network according to claim 1, characterized in that: In step S62, the method for determining the maximum audio level after attenuation of each audio track includes the following steps: S621. Calculate the ratio R of the audio high-level threshold to the sum of the maximum audio levels of each audio track; the calculation formula is: R = σ / ΣMAX; Where σ represents the audio high-level threshold; MAX represents the maximum audio level of the audio signal; ΣMAX represents the sum of the maximum audio levels of all audio tracks. S622. Calculate the product S of the maximum audio level of each audio track and the ratio mentioned above; the calculation formula is: S = MAX × R; S623. Determine the maximum audio level MAX' of each track audio signal after attenuation based on the value of the product S, such that MAX' ≤ S.

4. The method for audio multitrack synthesis during video recording based on a communication network according to claim 1, characterized in that: A low-level audio threshold is preset, and before superimposing and synthesizing the audio signals of each track, the portion of the audio signal below the low-level audio threshold is removed.

5. The method for audio multitrack synthesis during video recording based on a communication network according to any one of claims 1 to 4, characterized in that: The wireless communication network is a WIFI communication network, which includes a WIFI router. Both the audio acquisition device and the video shooting device are equipped with WIFI modules, and the audio acquisition device and the video shooting device are respectively connected to the WIFI router through their WIFI modules.

6. The method for audio multitrack synthesis during video recording based on a communication network according to any one of claims 1 to 4, characterized in that, The audio acquisition device is a surround sound recording device, a high-impedance musical instrument recording device, or a recording device that actively provides phantom power.

7. A video shooting device based on a communication network, characterized in that: include The video capture module is used to acquire video information through video capture; The first wireless communication module is used to connect to the audio acquisition device through a wireless communication network and acquire data packets of audio information sent by the audio acquisition device. The master synchronization module is used to keep the audio acquisition device connected to the video shooting device synchronized with the video shooting device in time. The attenuation control module is used to obtain the maximum audio level of each audio signal track, and when the sum of the maximum audio levels of each audio signal track is greater than a preset high-level audio threshold, the audio attenuation module attenuates the audio signal of each track. The audio attenuation module is used to restore the received audio information data packets to audio signals, remove the parts of each audio track that are below the low audio level threshold, and attenuate the parts of each audio track that are above the audio attenuation level threshold when attenuation is required. as well as The audio-video synthesis module is used to align the timelines of the audio signals of each track, and then superimpose and synthesize the audio signals of each track to obtain a synthesized audio signal. The synthesized audio signal is then aligned with the timeline of the video information and synthesized with the video information to form a video signal with audio information. The method by which the audio attenuation module attenuates the portion of each audio track signal that is higher than the audio attenuation level threshold includes the following steps: S61. Preset the audio attenuation level threshold; S62. Determine the maximum audio level after attenuation of each audio signal track; S63. Determine the attenuation ratio based on the audio attenuation level threshold, the maximum audio level of the audio signal, and the maximum audio level after attenuation; the formula for determining the attenuation ratio P is: P = (MAX' - δ) / (MAX - δ); Where MAX' represents the maximum audio level after the audio signal has attenuated; δ represents the audio attenuation level threshold; and MAX represents the maximum audio level of the audio signal. S64. Attenuate the portion of each audio track signal that is above the audio attenuation level threshold according to the attenuation ratio; the attenuation formula is: A' = (A - δ) × P + δ; Where A represents the audio signal level before attenuation, and A' represents the audio signal level after attenuation.

8. The video shooting device based on a communication network according to claim 7, characterized in that: The wireless communication network is a WIFI communication network, which includes a WIFI router. Both the audio acquisition device and the video shooting device are equipped with WIFI modules, and the audio acquisition device and the video shooting device are respectively connected to the WIFI router through their WIFI modules.

9. A video shooting system based on a communication network, characterized in that: Includes the video recording device as described in claim 7 or 8 and multiple audio acquisition devices; The audio acquisition device includes: The audio acquisition module is used to acquire audio information through audio sampling and package the acquired audio information into data packets; The second wireless communication module is used to connect to the video recording device via a wireless communication network and to send the data packets to the video recording device; and The synchronization module is used to keep the audio acquisition device and the video shooting device in time synchronization.

Citation Information

Patent Citations

  • Multipath signal live broadcast method and system for local area network

    CN107197317A

  • Voice recording and reproducing device

    JP2000090574A