Audio and video acquisition equipment and audio processing chip
By combining local and wireless microphones in the audio and video acquisition device to synthesize stereo dual-channel audio data, and using AI to recognize the user's status and automatically adjust the device's working status, the problem of limited sound pickup distance and insufficient sound effect of traditional devices is solved, achieving high-quality audio and video acquisition and resource conservation.
Patent Information
- Application Number
- CN202511538510.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-24
- Publication Date
- 2026-01-27
AI Technical Summary
Traditional audio and video capture devices have limited local microphone pickup distance, which makes it impossible to effectively capture audio data in some scenarios. Furthermore, mono capture lacks stereo sound and cannot meet the needs of application scenarios with high sound quality requirements. In addition, the device cannot automatically detect the status when the user goes offline, resulting in wasted resources and the risk of privacy leaks.
Design an audio and video acquisition device that combines a local microphone and a wireless microphone. The device synthesizes stereo dual-channel audio data through an audio processing unit and automatically adjusts the device's working status, including a standby mode, to save resources and protect privacy by recognizing the user's status using AI.
It breaks through the traditional limitations of sound pickup distance, provides stereo sound, improves sound quality, and automatically enters standby mode when the user goes offline, saving resources and reducing the risk of privacy leaks.
Smart Images

Figure CN121418752A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio and video processing technology, and more specifically, to an audio and video acquisition device and an audio processing chip. Background Technology
[0002] In scenarios such as video conferencing, live streaming, and online courses, audio and video capture devices are required to collect users' audio and video data. These devices typically include a built-in local microphone and an image sensor to capture audio and video data, respectively, which are then processed and output. However, because the built-in local microphone of these devices is usually fixed in position with the device, its pickup distance is limited, restricting its application in some scenarios. Furthermore, local microphones can only capture mono audio data from a fixed location. Although it can be converted into stereo audio data through subsequent processing, the final processed audio data lacks spatial and stereo sound, making it unsuitable for scenarios with high sound quality requirements. Summary of the Invention
[0003] In view of this, this application provides an audio and video acquisition device and an audio processing chip.
[0004] According to a first aspect of this application, an audio and video acquisition device is provided, the audio and video acquisition device including an image sensor, a video processing unit, a local microphone, and an audio processing unit; the image sensor is communicatively connected to the video processing unit, the audio processing unit is communicatively connected to the local microphone, and the audio processing unit is also wirelessly connected to an external wireless microphone; The image sensor is used to acquire video data and output it to the video processing unit; The video processing unit is used to process the video data and output the processed video data; The local microphone is used to collect the first audio data and output it to the audio processing unit; The audio processing unit is used to receive the second audio data collected by the wireless microphone, combine the first audio data and the second audio data into stereo audio data with stereo effect, and output it.
[0005] According to a second aspect of this application, an audio processing chip is provided, which is applied to an audio and video acquisition device. The audio and video acquisition device includes an image sensor, a video processing chip, a local microphone, and the audio processing chip. The image sensor is communicatively connected to the video processing chip, the audio processing chip is communicatively connected to the local microphone, and the audio processing chip is also wirelessly connected to an external wireless microphone. The audio processing chip is used to receive first audio data sent by the local microphone and second audio data collected by the wireless microphone, combine the first audio data and the second audio data into stereo audio data, and output it. And for determining the current user status based on the first audio data and / or the second audio data; the user status includes an online status or an offline status; when the user status is offline, controlling at least one of the audio processing chip and the video processing chip to enter a standby state.
[0006] Based on the solution provided in this application, an audio and video acquisition device is designed. This device supports both local and wireless microphones. It acquires one audio channel via the local microphone built into the device, and simultaneously acquires another audio channel via a wireless microphone worn by the user, also integrated into the device, through an audio processing unit. The audio processing unit then combines the two audio channels into a stereo sound with realistic spatial perception. Simultaneously, video acquisition and processing can be achieved using a built-in image sensor and video processing unit. This audio and video acquisition device overcomes the limitations of traditional devices in terms of pickup distance and solves the problem of mono sound lacking stereo perception, achieving simultaneous audio and video acquisition and high-quality output.
[0007] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0008] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0009] Figure 1 This is a schematic diagram illustrating an application scenario of one embodiment of this application.
[0010] Figure 2 This is a schematic diagram of an audio / video acquisition device according to an embodiment of this application.
[0011] Figure 3 This is a schematic diagram of an audio / video acquisition device according to an embodiment of this application.
[0012] Figure 4 This is a schematic diagram of the data processing link of an audio / video acquisition device according to an embodiment of this application. Detailed Implementation
[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0014] With the widespread adoption of applications such as video conferencing, live streaming, and online classes, higher demands are being placed on the audio and video capture equipment used in these scenarios. Traditional audio and video capture equipment typically has a built-in local microphone and image sensor to collect users' audio and video data, respectively. However, the position of the built-in local microphone is fixed along with the location of the audio and video capture equipment, resulting in a short pickup distance, usually less than 3 meters, which limits its application in some scenarios. For example, in video conferencing, when multiple people are sitting around a table, the voices of participants at a distance are usually weak and cannot be clearly captured. In live streaming scenarios, the broadcaster can only broadcast from a position close to the audio and video capture equipment and cannot move to a position farther away.
[0015] Furthermore, local microphones typically capture mono audio. Although pseudo-stereo can be generated by copying it later, the lack of spatial stereo means it cannot reproduce the true sound field, resulting in poor sound quality. Traditional audio and video capture equipment is unsuitable for scenarios with high audio quality requirements, such as live streaming.
[0016] Based on this, embodiments of this application provide an audio / video acquisition device that simultaneously supports both a local microphone and a wireless microphone. It acquires one audio data stream via the local microphone built into the device, and simultaneously connects the audio data to the audio processing unit integrated within the device. A wireless microphone worn by the user wirelessly connects and acquires another audio data stream from the microphone. The audio processing unit then combines the two audio streams into a two-channel stereo sound with a realistic sense of space. Simultaneously, a built-in image sensor and video processing unit can be used to acquire and process video. This audio and video acquisition device overcomes the limitations of traditional equipment's pickup distance and solves the problem of mono sound lacking stereo effect, achieving simultaneous audio and video acquisition and high-quality output.
[0017] like Figure 1The diagram illustrates an application scenario of the audio / video acquisition device of this application. Taking video conferencing as an example, a user can wear a wireless microphone, which can wirelessly connect to the audio / video acquisition device. The audio / video acquisition device can connect to a display device (e.g., a computer), and a video conferencing app can be installed on the display device (computer). The audio / video acquisition device has a built-in image sensor and a local microphone. The image sensor can acquire the user's video data, and the device can perform various processing on the video data acquired by the image sensor, such as adding subtitles, adding watermarks, resolution conversion, and zooming. The processed video data is then encoded and output. The local microphone in the audio / video acquisition device can acquire audio data 1 and receive audio data 2 acquired by the wireless microphone. The device can perform noise reduction, timbre customization, level control, and gain adjustment on both audio data 1 and audio data 2 respectively. The two processed audio data are then combined into a single audio data 3 with stereo effects, encoded, and output. For example, the audio / video acquisition device can output the encoded audio / video data 3 and video data to a display device, which is then sent to a cloud server by the video conferencing app.
[0018] like Figure 2 The diagram shows a schematic of an audio / video acquisition device 10 in one embodiment. The device includes an image sensor 11, a video processing unit 12, a local microphone 13, and an audio processing unit 14. The image sensor 11 can be any type of sensor capable of converting light signals into electrical signals, such as a CMOS sensor or a CCD sensor. The image sensor 11 can be used to acquire raw video data, such as 4K@30FPS / 60FPS RAW video data.
[0019] The video processing unit 12 can be any chip with image processing capabilities, such as a SOC chip. The video processing unit 12 can communicate with the image sensor 11, for example, through a hardware interface, to receive the raw video data output by the image sensor 11, process the raw video data, and then output it. For example, it can perform various processing on the raw video data, such as adding subtitles, embedding logos, converting resolution, flipping images, and applying electronic zoom. It can also encode the processed raw video data before outputting it, for example, through H.264 / H.265 encoding and outputting via the UVC protocol, ensuring that the video signal conforms to various transmission and storage standards.
[0020] The local microphone 13 can be a microphone built into the audio and video acquisition device 10, which can acquire analog audio signals and convert them into digital signals. The local microphone 13 can be communicatively connected to the audio processing unit 14. For example, it can be connected to the audio processing unit 14 through a hardware interface to output the acquired audio data (hereinafter referred to as the first audio data) to the audio processing unit 14.
[0021] The audio processing unit 14 can be any chip with wireless receiving and audio signal processing capabilities, such as an FPGA chip. This chip includes a wireless module (e.g., a WiFi module, Bluetooth module) to achieve wireless connection with a user-worn wireless microphone, thereby receiving audio data (hereinafter referred to as the second audio data) transmitted by the wireless microphone. After receiving the first audio data output from the local microphone 13 and the second audio data transmitted by the wireless microphone, the audio processing unit 14 can synthesize the two into stereo audio data with a stereo effect. For example, algorithms such as phase calibration and level matching can be used to synthesize the two audio streams into truly spatially positioned stereo data, which is then output via the UAC protocol.
[0022] In some embodiments, the video processing unit 12 and the audio processing unit 14 can be integrated on a single chip. In some embodiments, the video processing unit 12 and the audio processing unit 14 can also be integrated on two separate chips, which can communicate with each other.
[0023] The audio processing unit 14 can be wirelessly connected to one or multiple wireless microphones, meaning the second audio / video data can include multiple audio data streams collected by multiple wireless microphones. The audio processing unit 14 can include multiple audio data processing channels, each of which can be used to process either the first audio data or a single second audio data stream collected by one wireless microphone.
[0024] In some embodiments, the audio processing unit and the wireless microphone can communicate based on various short-range communication protocols, such as Bluetooth, WiFi, or custom protocols.
[0025] Traditional audio and video capture devices cannot automatically detect changes in the user's online / offline status when the user ends use the device (e.g., logging off or leaving). They continue to capture and output the user's audio and video until the device is manually turned off via a client or physical button. However, in many scenarios, users may not turn off the device in time or forget to do so. In these scenarios, the device operates at full capacity throughout, especially when unused for extended periods, resulting in unnecessary resource consumption. Furthermore, in situations where users forget to turn off the device, the continuous capture of audio and video data by the audio and video capture device 10 may record irrelevant conversations or ambient noise, posing a privacy risk and making it particularly unsuitable for private scenarios such as homes and offices.
[0026] To avoid the aforementioned problems, in some embodiments, the audio processing unit 14 can detect the first audio data and / or the second audio data to analyze the current user status. For example, whether the user is currently online or offline. An online state refers to the state where the user currently needs to use the audio / video capture device 10, while an offline state refers to the state where the user does not currently need to use the audio / video capture device 10. If it is determined based on the first audio data and / or the second audio data that the user is currently offline, the audio processing unit 14 and / or the video processing unit 12 can be controlled to enter a standby state.
[0027] For example, the audio processing unit 14 can analyze the first audio data collected by the local microphone 13 and the second audio data collected by the wireless microphone in real time, detect the user's status through the built-in AI voice recognition model, or determine the user's status through some preset detection rules.
[0028] For example, in some embodiments, human voice detection can be performed on the first audio data and the second audio data to determine whether there is human voice. If no human voice is detected in the first audio data or the second audio data for a preset time, it can be determined that the user is currently offline. For example, if no human voice signal is detected in the first audio data or the second audio data for more than 5 consecutive minutes, it can be determined that the user is currently offline.
[0029] In some embodiments, specific keyword detection can be performed on the first audio data and the second audio data. If a specific keyword is detected in the first audio data or the second audio data, it is determined that the user is currently offline. The specific keyword includes words or sentences used to indicate that the user is about to log off. The specific keyword can be flexibly set based on the characteristics of the application scenario. For example, for a live streaming scenario, the specific keyword could be "logged off," "see you next time," "bye-bye," "that's all for today's live stream," "anyone else want more?" or "For those who haven't had enough, we'll notify you as soon as the live stream starts tomorrow. Goodnight for now." For a video conferencing scenario, the specific keyword could be "that's all for today's meeting," "the meeting is over," etc.
[0030] When a user is detected to be offline, the audio processing unit 14 can be automatically put into standby mode. Once in standby mode, the audio processing unit 14 can shut down some high-power functional modules or reduce the frequency and power consumption of others. For example, it can reduce the sampling frequency of the local microphone 13 or the wireless microphone, disable audio processing functions (e.g., disable AI noise reduction, EQ, ALC, gain adjustment, etc.), disable stereo synthesis, and stop outputting audio data. Of course, in some embodiments, to achieve rapid wake-up and return to normal operation, the keyword wake-up function can be retained. For example, the audio processing unit 14 can still receive the first and second audio data from the local microphone 13 and the wireless microphone respectively, and perform keyword detection. If certain specific keywords are detected, it switches back to normal operation.
[0031] In some embodiments, such as Figure 3 As shown, the audio processing unit 14 and the video processing unit 12 can communicate. When the user is detected to be offline, the audio processing unit 14 can send a state switching command to the video processing unit 12 to control the video processing unit 12 to enter a standby state. After entering the standby state, the video processing unit 12 can also shut down some high-power functional modules or reduce the frequency and power consumption of some functional modules. For example, it can pause the processing and output of video data, or reduce the frame rate of the image sensor 11, etc.
[0032] In some embodiments, in order to improve the quality and sound effects of the output audio data, the audio processing unit 14 may first perform sound effect processing on the first audio data and the second audio data respectively before combining the first audio data and the second audio data into stereo audio data. Then, the first audio data and the second audio data after sound effect processing are combined into stereo audio data.
[0033] In some embodiments, audio processing may include one or more of the following: AI noise reduction, timbre customization, automatic level control, gain adjustment, and echo cancellation, applied to the first and second audio data. Traditional audio / video acquisition devices 10 typically use a fixed noise reduction algorithm to uniformly reduce all types of noise during noise reduction. This algorithm treats all non-steady-state signals as noise and suppresses them, failing to distinguish between "human voice" and "noise" signals in the same frequency band. This approach results in significant sound quality loss after noise reduction, with a large difference between the human voice and the original sound. To avoid these problems, in some embodiments, when processing the first and second audio data, the first and / or second audio data can be input into a pre-trained AI noise reduction model. This AI noise reduction model identifies the noise type in the first and / or second audio data and then performs targeted noise reduction based on the noise type.
[0034] For example, typical noise types that can be recorded include: keyboard clicks, page turning, desk and chair sounds, mouse clicks, environmental noise (air conditioning, fans, outdoor traffic, whispers), and background music. A large amount of sample data can be collected for each type of noise. For instance, a clean human voice can be used as a label, and mixed audio data including that type of noise can be used as sample input to perform supervised training on the model. This allows the model to learn the optimal processing method for different types of noise, resulting in an AI noise reduction model. Then, the trained AI noise reduction model can be used to perform noise reduction processing on the first audio data and / or the second audio data.
[0035] In some embodiments, when the audio processing unit 14 is in standby mode, the audio processing unit 14 suspends the steps of performing sound effect processing on the first audio data and the second audio data and synthesizing stereo audio data. For example, the video processing unit 12 suspends processing such as AI noise reduction, timbre customization, automatic level control, and gain adjustment on the first audio data and the second audio data, stops the synthesis and encoding processing of the two audio data, and stops outputting the encoded audio data.
[0036] In some embodiments, when the video processing unit 12 is in standby mode, it suspends the steps of processing and outputting video data. For example, the video processing unit 12 may pause processing the video data acquired by the image sensor 11, such as adding subtitles, adding watermarks, resolution conversion, and encoding, and also stop outputting the encoded video data. In some embodiments, when it is determined that the user is currently online based on the first audio data or the second audio data, a state switching command may be sent to the video processing unit 12 to control it to switch from standby mode to normal operation mode. By having the audio processing unit 14 continuously detect the first audio data and / or the second audio data to determine whether the user is online, normal operation mode can be quickly restored after the user is detected to be online.
[0037] In some embodiments, after the audio processing unit 14 and / or the video processing unit 12 are put into standby mode, the audio processing unit 14 may continue to detect the first audio data and / or the second audio data to determine the user status. For example, if it is determined based on the first audio data or the second audio data that the user is currently online, the audio processing unit 14 may be switched from standby mode to normal operation mode.
[0038] In some embodiments, if it is determined that the user is currently online based on the first audio data or the second audio data, a state switching instruction can also be sent to the video processing unit 12 to control the video processing unit 12 to switch from the standby state to the normal working state.
[0039] This could be determined by detecting human voices, or by detecting specific keywords. These specific keywords could be words or phrases that indicate the user is online, such as "The meeting has started" or "Family, I'm back."
[0040] In some embodiments, the video processing unit 12 can also be used to detect and analyze video data to determine the user's current status. For example, if the video processing unit 12 determines that the user is currently offline based on the video data, it can control the video processing unit 12 and / or the audio processing unit 14 to enter a standby state.
[0041] In some embodiments, the video processing unit 12 can perform human detection on the video frame. If no human is detected in the video frame, it is determined that the user is currently offline.
[0042] In some embodiments, the video processing unit 12 can detect and analyze the actions of people in the video frame. After determining that a specific action of a person has been detected, it determines that the user is currently offline. This specific action can be the action of the user leaving, such as a person moving from the center of the frame to the edge and eventually leaving the frame completely (lasting more than 2 seconds). The specified action can be the user waving goodbye, etc.
[0043] In some embodiments, when detecting a user's online and offline status, the detection results of audio data by the audio processing unit 14 and the detection results of video data by the video processing unit 12 can be combined to comprehensively determine the user's status and obtain a more accurate determination result.
[0044] For example, in some embodiments, when determining the user state, the audio processing unit 14 may receive a first user state detection result sent by the video processing unit 12, which is determined by the video processing unit 12 based on video data collected by the image sensor 11. Then, a second user state detection result is determined based on the first audio data and / or the second audio data, and the final user state is determined based on the combination of the first user state detection result and the second user state detection result.
[0045] Traditional audio and video acquisition devices generally employ a fixed noise reduction strategy for audio data captured by the local microphone 13. This strategy cannot flexibly adjust the processing method according to the needs of different scenarios, leading to excessive noise reduction and resulting in sound quality loss. For example, in scenarios where environmental sound effects need to be preserved (such as outdoor live streaming or music performance recording), the fixed noise reduction processing will simultaneously filter out effective high-frequency components in the environmental sound (such as high notes from instruments and details of natural sounds), making the sound thin and distorted. Considering that wireless microphones primarily capture human voices, while the local microphone 13 can capture environmental sounds, users have different requirements for preserving environmental sounds in different scenarios. To meet the usage needs of different scenarios, in some embodiments, the local microphone 13 can be set to two working modes: a noise reduction on mode and a noise reduction off mode. The switching of the working mode of the local microphone 13 can be done manually by the user or automatically by the audio acquisition device. For example, the audio processing unit 14 can have a built-in scene recognition model. By analyzing the characteristics of the first audio data collected by the local microphone 13 and the second audio data collected by the wireless microphone, it can automatically determine the scene type. When it detects that human voices are dominant (human voices account for >60%) and there is significant environmental noise (such as in a conference room scene), it automatically switches to noise reduction on mode. When it detects abundant ambient sounds and low noise (such as in outdoor natural scenes or musical performances), it automatically switches to noise reduction off mode. Of course, it can also support users to force switching modes through physical buttons on the device, the accompanying APP, or voice commands (such as "turn on noise reduction" or "turn off noise reduction"). Among these, manual switching commands have higher priority than automatic switching commands. Alternatively, the device can also save the user's most recent manual switching preference and apply it preferentially in the same scenario (such as if the user is used to turning off noise reduction during live streaming, the device will use this setting by default in subsequent live streaming scenarios).
[0046] When the local microphone 13 is in noise reduction enabled mode, the audio processing unit 14 can perform target type processing on the first audio data. When the local microphone 13 is in noise reduction disabled mode, the audio processing unit 14 does not perform target type processing on the first audio data. The target type processing includes: performing AI noise reduction processing on the first audio data and removing the high-frequency part of the first audio data with a frequency higher than a preset frequency.
[0047] The target type processing can include two levels of processing. First, AI noise reduction can be performed on the environmental noise in the first audio data. For example, a trained AI noise reduction model can be used to reduce noise in the first audio data. Second, high-frequency components in the first audio data with frequencies higher than a preset frequency can be removed. For example, the portion above 18kHz can be removed.
[0048] With its dual-mode design, the same device can handle both scenarios—"prioritizing clear human voices" (such as in meetings) and "fully preserving ambient sounds" (such as in outdoor recording)—meeting diverse needs without the need for additional equipment. In noise-cancellation-on mode, it effectively suppresses specific noises while minimizing damage to human voices through precise filtering. In noise-cancellation-off mode, it fully preserves ambient sound details, enhancing the richness and realism of the sound. Noise cancellation is only applied when noise is detected and noise is enabled, maximizing the preservation of high-frequency details in the audio (such as instrument overtones and crisp ambient sounds).
[0049] The audio and video acquisition device of this application will be described below with reference to a specific embodiment.
[0050] To increase the pickup distance of audio and video acquisition devices and to obtain audio and video acquisition devices with stereo sound effects, this embodiment provides an audio and video acquisition device that simultaneously supports local microphones and wireless microphones. The audio and video acquisition device includes an image sensor, a video processing unit, a built-in microphone, and an audio processing unit, with the audio processing unit wirelessly connected to the wireless microphone. Figure 4 The diagram shown is a data processing link of the audio and video acquisition device in this embodiment. The following sections will describe each data processing link.
[0051] 1. Video data processing link (1) The image sensor in the audio and video acquisition device can acquire raw video data and realize the conversion of optical signals to electrical signals. For example, the acquired video data is 4K, 60FPS / 4K, 30FPS RAW data.
[0052] (2) The video processing unit in the audio and video acquisition device can configure video acquisition parameters, receive and send data (i.e., video input), and supports functions such as adding subtitles, adding image logos, pixel format / resolution conversion, and image flipping. The video processing unit can include multiple processing channels, each of which can perform different types of processing on the video data. The video processing unit can also include an extension channel for implementing electronic zoom. It can also encode the processed video data, such as H264, H265, and MJPEG encoding. The video processing unit can encode the video data into UVC format and then output it via a USB interface.
[0053] In addition, the video processing unit can also receive feedback information from the audio processing unit in the audio and video acquisition device, and switch the standby / working state of the video processing unit based on the feedback information.
[0054] 2. Wireless microphone audio data processing link (1) The wireless microphone can collect analog audio signals and convert them into digital signals. For example, it can collect 48K sampling rate, mono, 16-bit PDM data and send it to the audio processing unit of the audio and video acquisition device via Bluetooth protocol.
[0055] (2) The audio processing unit can perform audio protocol processing on the received wireless microphone audio data, such as protocol conversion processing on the received wireless microphone audio data.
[0056] (3) The audio processing unit can perform EQ (Equalizer, frequency domain processor) on the wireless microphone audio data, that is, the function of customizing the timbre of the wireless microphone audio data (such as highlighting the human voice and enhancing the low frequency).
[0057] (4) The audio processing unit can perform AI noise reduction processing on the audio data of the wireless microphone. For example, it can identify various types of noise, such as keyboard sounds, book turning sounds, table and chair moving sounds, mouse sounds, environmental noise, etc., and then filter out the noise accordingly.
[0058] (5) The audio processing unit can perform ALC (Automatic Level Control) on the audio data of the wireless microphone, monitor the amplitude of the input signal in real time, automatically adjust the gain, and prevent overload distortion (such as output popping and crackling).
[0059] (6) The audio processing unit can digitally adjust the audio data of the wireless microphone and change the amplitude of the audio signal through software algorithms to ensure stable output level and prevent sudden volume changes.
[0060] 3. Built-in microphone audio and video data processing link (1) The built-in microphone collects analog audio signals and converts them into digital signals. For example, the collected audio data is 48K sampling rate, mono, 16-bit PCM data.
[0061] (2) The audio processing unit can decode the audio data of the built-in microphone and decode the LC3 encoded audio data into PCM data.
[0062] (3) The built-in microphone has two modes: noise reduction off and noise reduction on.
[0063] In noise cancellation off mode, the built-in microphone audio data can be EQ (Equalizer, frequency domain processor), which customizes the timbre of the built-in microphone audio data (such as highlighting vocals or enhancing low frequencies). Then, the built-in microphone audio data can be subjected to ALC (Automatic Level Control), which monitors the input signal amplitude in real time and automatically adjusts the gain to prevent overload distortion (such as output popping or crackling). Finally, digital gain adjustment of the built-in microphone audio data can be performed, using software algorithms to change the audio signal amplitude to ensure stable output levels and prevent sudden volume changes.
[0064] With noise cancellation mode enabled, the built-in microphone audio data can first undergo EQ2 processing, a function that customizes the tone (such as highlighting vocals and enhancing low frequencies) while removing high-frequency noise (18-20kHz). Then, AI noise reduction processing can be applied to the built-in microphone audio data. For example, it can identify various types of noise, such as keyboard sounds, page turning, the sound of furniture being moved or tapped, mouse clicks, and ambient noise, and then filter out these noises accordingly. Afterward, the aforementioned EQ, ALC, and digital gain adjustments can be performed.
[0065] 4. Built-in microphone audio data and wireless microphone audio data detection link The audio processing unit receives data from the built-in microphone and the wireless microphone to detect whether the sound contains human voices. If human voices are detected, the video link, wireless microphone link, and built-in microphone link are notified to operate normally; if no human voices are detected within five minutes, the video link, wireless microphone link, and local microphone link are notified to enter standby mode.
[0066] Alternatively, the audio processing unit receives data from the built-in microphone and the wireless microphone to detect whether the sound contains online keywords. If not, it notifies the video link, the wireless microphone link, and the local microphone link to operate normally. If the user goes offline, it notifies the video link, the wireless microphone link, and the local microphone link to enter standby mode.
[0067] 5. Combined link of built-in microphone audio data and wireless microphone audio data (1) After processing the audio data of the built-in microphone and the audio data of the wireless microphone, the audio processing unit can merge the sound data of the two microphones into stereo dual-channel data, then encode it into UAC format data and output it.
[0068] The solutions in the above embodiments can be freely combined to obtain new solutions when there is no conflict. Due to space limitations, they will not be listed one by one here.
[0069] Furthermore, this application embodiment also provides an audio processing chip, which is applied to an audio and video acquisition device. The audio and video acquisition device includes an image sensor, a video processing chip, a local microphone, and the audio processing chip. The image sensor is communicatively connected to the video processing chip, the audio processing chip is communicatively connected to the local microphone, and the audio processing chip is also wirelessly connected to an external wireless microphone. The audio processing chip is used to receive first audio data sent by the local microphone and second audio data collected by the wireless microphone, combine the first audio data and the second audio data into stereo audio data, and output it. And for determining the current user status based on the first audio data and / or the second audio data; the user status includes an online status or an offline status; when the user status is offline, controlling at least one of the audio processing chip and the video processing chip to enter a standby state.
[0070] In some embodiments, the audio processing chip is configured to determine the current user state based on the first audio data and / or the second audio data, specifically for: If no human voice is detected in the first audio data or the second audio data for a preset duration, the user's status is determined to be offline; and / or If a specific keyword is detected in the first audio data or the second audio data, the user's status is determined to be offline; wherein, the specific keyword includes words or sentences used to indicate that the user is about to go offline.
[0071] In some embodiments, the audio processing chip is configured to determine the user state based on the first audio data and / or the second audio data, specifically for: The video processing chip receives a first user status detection result, which is determined by the video processing chip based on video data collected by the image sensor. The second user status detection result is determined based on the first audio data and / or the second audio data; The user status is determined based on the first user status detection result and the second user status detection result.
[0072] In some embodiments, the first user status detection result includes the current user status being offline, and the video processing chip determining the current user status as offline if no person is detected in the video data; and / or If the video processing chip detects a specified action of a person in the video data, it determines that the current user is offline.
[0073] In some embodiments, after the audio processing chip controls at least one of the audio processing chip and the video processing chip to enter a standby state, the audio processing chip is further configured to: If it is determined that the user is currently online based on the first audio data and / or the second audio data, control at least one of the audio processing chip and the video processing chip to switch from standby state to normal operation state.
[0074] In some embodiments, the audio processing chip is used to synthesize the first audio data and the second audio data into stereo audio data, specifically for: The first audio data and the second audio data are processed with sound effects respectively. The first audio data after sound effect processing and the second audio data after sound effect processing are combined into stereo audio data with stereo effect.
[0075] In some embodiments, the sound processing includes one or more of the following: The first audio data and the second audio data are subjected to AI noise reduction processing, timbre customization processing, automatic level control processing, and automatic gain adjustment processing.
[0076] In some embodiments, when the audio processing chip is in standby mode, the audio processing chip suspends the steps of performing sound effect processing on the first audio data and the second audio data and synthesizing the stereo audio data.
[0077] In some embodiments, when the video processing chip is in standby mode, the video processing chip suspends the steps of processing and outputting the video data.
[0078] In some embodiments, the local microphone's operating mode includes a noise reduction on mode and a noise reduction off mode. When the local microphone is in the noise reduction on mode, the audio processing chip performs target-type processing on the first audio data. When the local microphone is in the noise reduction off mode, the audio processing chip does not perform the target-type processing on the first audio data. The target-type processing includes: performing AI noise reduction processing on the environmental noise in the first audio data and removing high-frequency components in the first audio data whose frequencies are higher than a preset frequency.
[0079] The specific processing details of the audio processing chip can be found in the description of the audio and video acquisition equipment mentioned above, and will not be repeated here.
[0080] Accordingly, this application also provides a computer storage medium storing a program that, when executed by a processor, implements the method in any of the above embodiments.
[0081] The embodiments of this application may take the form of a computer program product implemented on one or more storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing program code. Computer-usable storage media include permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information may be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to: phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, read-only optical discs, read-only memory (CD-ROM), digital versatile optical discs (DVD) or other optical storage, magnetic tape, disks or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0082] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0083] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0084] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0085] The methods and apparatus provided in the embodiments of this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this application should not be construed as a limitation of this application.
Claims
1. An audio / video acquisition device, characterized in that, The audio and video acquisition device includes an image sensor, a video processing unit, a local microphone, and an audio processing unit; the image sensor is communicatively connected to the video processing unit, the audio processing unit is communicatively connected to the local microphone, and the audio processing unit is also wirelessly connected to an external wireless microphone. The image sensor is used to acquire video data and output it to the video processing unit; The video processing unit is used to process the video data and output the processed video data; The local microphone is used to collect the first audio data and output it to the audio processing unit; The audio processing unit is used to receive the second audio data collected by the wireless microphone, combine the first audio data and the second audio data into stereo audio data with stereo effect, and output it.
2. The audio and video acquisition device according to claim 1, characterized in that, The audio processing unit is also used for: The current user status is determined based on the first audio data and / or the second audio data; the user status includes online status or offline status. When the user is offline, at least one of the audio processing unit and the video processing unit is controlled to enter standby mode.
3. The audio and video acquisition device according to claim 2, characterized in that, The audio processing unit is used to determine the current user status based on the first audio data and / or the second audio data, specifically for: If no human voice is detected in the first audio data or the second audio data for a preset duration, the user's status is determined to be offline; and / or If a specific keyword is detected in the first audio data or the second audio data, the user's status is determined to be offline; wherein, the specific keyword includes words or sentences used to indicate that the user is about to go offline.
4. The audio and video acquisition device according to claim 2, characterized in that, The audio processing unit is used to determine the user's state based on the first audio data and / or the second audio data, specifically for: The system receives a first user status detection result sent by the video processing unit, wherein the first user status detection result is determined by the video processing unit based on the video data. The second user status detection result is determined based on the first audio data and / or the second audio data; The user status is determined based on the first user status detection result and the second user status detection result.
5. The audio and video acquisition device according to claim 4, characterized in that, The first user status detection result includes the following: if the current user status is offline, and the video processing unit does not detect any person in the video data, then the current user status is determined to be offline; and / or If the video processing unit detects a specified action of a person in the video data, it determines that the current user is offline.
6. The audio and video acquisition device according to claim 2, characterized in that, After the audio processing unit controls at least one of the audio processing unit and the video processing unit to enter a standby state, the audio processing unit is further configured to: If it is determined that the user is currently online based on the first audio data and / or the second audio data, control at least one of the audio processing unit and the video processing unit to switch from standby state to normal working state.
7. The audio and video acquisition device according to claim 1 or 2, characterized in that, The audio processing unit is used to synthesize the first audio data and the second audio data into stereo audio data with stereo effects, specifically for: The first audio data and the second audio data are processed with sound effects respectively. The first audio data after sound effect processing and the second audio data after sound effect processing are combined into stereo audio data with stereo effect.
8. The audio and video acquisition device according to claim 7, characterized in that, The sound effects processing includes one or more of the following: The first audio data and the second audio data are subjected to AI noise reduction processing, timbre customization processing, automatic level control processing, and automatic gain adjustment processing.
9. The audio and video acquisition device according to claim 7, characterized in that, When the audio processing unit is in standby mode, the audio processing unit suspends the step of performing sound effect processing on the first audio data and the second audio data and synthesizing the stereo audio data.
10. The audio and video acquisition device according to claim 2, characterized in that, in, When the video processing unit is in standby mode, the video processing unit suspends the steps of processing and outputting the video data.
11. The audio and video acquisition device according to claim 1, characterized in that, The local microphone has two operating modes: noise reduction on and noise reduction off. When the local microphone is in noise reduction on mode, the audio processing unit performs target-type processing on the first audio data. When the local microphone is in noise reduction off mode, the audio processing unit does not perform the target-type processing on the first audio data. The target-type processing includes performing AI noise reduction processing on the environmental noise in the first audio data and removing high-frequency components in the first audio data that are higher than a preset frequency.
12. An audio processing chip, the audio processing chip being applied to an audio and video acquisition device, the audio and video acquisition device comprising an image sensor, a video processing chip, a local microphone, and the audio processing chip; the image sensor being communicatively connected to the video processing chip, the audio processing chip being communicatively connected to the local microphone, and the audio processing chip also being wirelessly connected to an external wireless microphone; The audio processing chip is used to receive first audio data sent by the local microphone and second audio data collected by the wireless microphone, combine the first audio data and the second audio data into stereo audio data, and output it. And for determining the current user status based on the first audio data and / or the second audio data; the user status includes an online status or an offline status; when the user status is offline, controlling at least one of the audio processing chip and the video processing chip to enter a standby state.
13. The audio processing chip according to claim 12, characterized in that, The audio processing chip is used to determine the user's state based on the first audio data and / or the second audio data, specifically for: The video processing chip receives a first user status detection result, which is determined by the video processing chip based on video data collected by the image sensor. The second user status detection result is determined based on the first audio data and / or the second audio data; The user status is determined based on the first user status detection result and the second user status detection result.
14. The audio processing chip according to claim 12, characterized in that, The local microphone has two operating modes: noise reduction on and noise reduction off. When the local microphone is in noise reduction on mode, the audio processing chip performs target-type processing on the first audio data. When the local microphone is in noise reduction off mode, the audio processing chip does not perform the target-type processing on the first audio data. The target-type processing includes: performing AI noise reduction processing on the environmental noise in the first audio data and removing high-frequency components in the first audio data that are higher than a preset frequency.