Pickup method and terminal equipment
Patent Information
- Application Number
- CN202580002848.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-05-07
- Filing Date
- 2025-04-27
- Publication Date
- 2026-01-09
AI Technical Summary
Existing terminal devices suffer from a mismatch between video and audio when shooting from a distance, resulting in poor audio quality.
The system uses a combination of millimeter-wave radar and microphone for sound pickup. The pickup method is selected based on the shooting distance and zoom level. The millimeter-wave radar is used to pick up audio from a distance, while the microphone is used to pick up audio from a distance. The millimeter-wave radar is used to determine the angle and distance of the target object for signal transmission and reception.
It enables the matching of video and audio during long-distance video shooting, improving audio quality and user experience.
Smart Images

Figure CN121312151A_ABST
Abstract
Description
Sound pickup methods and terminal equipment
[0001] This application claims priority to Chinese patent application filed on May 7, 2024, with application number 202410557998.9 and entitled "Method and Terminal Device for Sound Pickup", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of terminal technology, and in particular to a method for sound pickup and a terminal device. Background Technology
[0003] With the development of the internet, more and more users are recording their lives in the form of videos. Videos can capture richer scene details, and their dynamic visuals and sound reproduction enrich people's daily lives. Sharing videos on social media has also become an increasingly popular sharing method for users.
[0004] A high-quality video requires not only high-definition visuals but also matching high-quality audio. However, most current consumer electronics products focus on improving picture quality but neglect the accompanying audio pickup technology. This can lead to situations where a device can use a telephoto lens to capture high-definition images from a greater distance but cannot pick up the corresponding sound, resulting in a mismatch between the video and audio. Summary of the Invention
[0005] This application provides a method and terminal device for sound pickup, which uses millimeter-wave radar to pick up sound from distant images in video shooting scenarios, thus solving the problem of mismatch between video images and audio at long distances.
[0006] In the sound pickup method of this application, when shooting a distant scene, a millimeter-wave radar installed in the terminal device picks up the distant audio, allowing the terminal device to overcome the limitation of the sound pickup distance and provide the user with a video that matches the audio. Specifically, when the user uses the terminal device to shoot a distant scene, they can choose to use a microphone array or a millimeter-wave radar for sound pickup based on the current zoom level or the distance manually set by the user. When millimeter-wave radar is selected for sound pickup, the target object and its angle in the shooting scene can be determined first based on preset characteristics. Then, the millimeter-wave radar emits a beam towards that angle to detect the distance between the target object and the terminal device. Next, the transmission power of the radar wave signal is determined based on the distance, and the signal is emitted in the direction of the target object. Based on the received echo signal modulated by the target object, the audio signal of the target object can be obtained through processing, that is, the audio corresponding to the video scene can be obtained.
[0007] In a first aspect, a method for sound pickup is provided, applied to a terminal device, the terminal device including a microphone and millimeter-wave radar, the method comprising:
[0008] Get the zoom level of the current recording, or the pickup distance input by the user;
[0009] The sound pickup method is obtained based on the zoom level or the pickup distance. The sound pickup method includes using the millimeter-wave radar alone, using the microphone alone, and using both the microphone and the millimeter-wave radar simultaneously.
[0010] The first audio signal picked up by the millimeter-wave radar and / or the second audio signal picked up by the microphone are processed according to the aforementioned sound pickup method to obtain the processed audio result.
[0011] Both the microphone and millimeter-wave radar in the terminal device can be used for sound pickup. Specifically, the microphone can be used to pick up audio at close range, while the millimeter-wave radar can be used to pick up audio at long range. The distinction between close range and long range can be flexibly made based on the hardware performance of the terminal device (such as the sound pickup performance of the microphone and millimeter-wave radar).
[0012] According to the sound pickup method provided in this implementation, by using a microphone and millimeter-wave radar in combination, the terminal device can not only pick up audio at close range, but also at long distance.
[0013] In conjunction with the first aspect, in some implementations of the first aspect, the terminal device further includes a camera, and the distance between the millimeter-wave radar and the camera is less than or equal to a first distance.
[0014] Understandably, when using millimeter-wave radar to pick up audio from a distance, it is necessary to determine the angle of the target object emitting the sound based on the video footage captured by the camera so that the millimeter-wave radar can determine the beam direction. Therefore, placing the millimeter-wave radar close to the rear camera can reduce the error of obtaining the beam direction based on the angle of the target object in the video footage, thereby improving the accuracy of target object detection and the sound pickup quality.
[0015] In conjunction with the first aspect, in some implementations of the first aspect, the terminal device is a mobile phone, the camera is a rear camera in the mobile phone, and the millimeter-wave radar is located in the camera area corresponding to the rear camera.
[0016] In conjunction with the first aspect, in certain implementations of the first aspect, the sound pickup method is obtained based on the zoom magnification, specifically including:
[0017] When the zoom factor is less than the first factor, the sound pickup method is to use a microphone alone for sound pickup;
[0018] When the zoom factor is greater than or equal to the first factor and less than the second factor, the sound pickup method is to simultaneously use the microphone and the millimeter-wave radar to pick up sound;
[0019] When the zoom factor is greater than or equal to the second factor, the sound pickup method is to use millimeter-wave radar for sound pickup alone; wherein, the value corresponding to the second factor is greater than the value corresponding to the first threshold.
[0020] Understandably, when the zoom level is greater than or equal to the second zoom level, it corresponds to the terminal device recording a video of a distant scene. At this time, it is necessary to obtain the sound that matches the video of the distant scene, so millimeter-wave radar can be used to pick up the sound separately.
[0021] In conjunction with the first aspect, in certain implementations of the first aspect, the sound pickup method is obtained based on the sound pickup distance, specifically including:
[0022] When the pickup distance is less than the first threshold, the pickup method is to use a microphone alone for pickup.
[0023] When the pickup distance is greater than or equal to the first threshold and less than the second threshold, the pickup method is to simultaneously use the microphone and the millimeter-wave radar to pick up sound.
[0024] When the pickup distance is greater than or equal to the second threshold, the pickup method is to use millimeter-wave radar for pickup alone; wherein, the second threshold is greater than the first threshold.
[0025] Understandably, when the sound pickup distance is greater than or equal to the second threshold, it corresponds to the terminal device recording a video at a distance. At this time, it is necessary to obtain the sound that matches the video recording at a distance, so millimeter-wave radar can be used to pick up the sound alone.
[0026] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes:
[0027] The system receives a first operation from the user, which is used to select a sound pickup mode, including an automatic sound pickup mode and a manual sound pickup mode.
[0028] When the sound pickup mode is automatic sound pickup mode, the zoom magnification is obtained, and the sound pickup method is obtained according to the zoom magnification;
[0029] When the sound pickup mode is manual sound pickup mode, a second operation input by the user is received, the second operation being used to select the sound pickup distance;
[0030] The pickup method is obtained based on the pickup distance.
[0031] In conjunction with the first aspect, in some implementations of the first aspect, when the sound pickup method involves using the millimeter-wave radar alone or using both the microphone and the millimeter-wave radar simultaneously, the method further includes:
[0032] Get the currently recorded video frame;
[0033] Based on the preset first feature information of the sound-emitting object and the video frame, the target angle of the sound-emitting object relative to the terminal device is obtained;
[0034] Set the beam pointing of the millimeter-wave radar to the target angle;
[0035] The second distance between the sound-emitting object and the terminal device is determined by the beam;
[0036] The transmission power of the millimeter-wave radar signal is determined based on the second distance;
[0037] According to the transmission power, a transmission signal is sent to the target angle, and the coverage area of the transmission signal includes the sound-emitting object.
[0038] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes:
[0039] Acquire the echo signal, wherein the echo signal is the signal after the transmitted signal is modulated by the vibrating part of the sound-producing object;
[0040] Demodulate the echo signal to recover the first audio signal corresponding to the sound source.
[0041] In conjunction with the first aspect, in certain implementations of the first aspect, demodulating the echo signal to recover the first audio signal corresponding to the sound-producing object specifically includes:
[0042] The echo signal is demodulated, and the first audio signal corresponding to the sound source is recovered based on the first phase of the transmitted signal and the second phase of the echo signal.
[0043] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes:
[0044] The first audio signal is subjected to noise reduction and audio super-resolution processing to obtain the third audio signal corresponding to the millimeter-wave radar.
[0045] In conjunction with the first aspect, in some implementations of the first aspect, when the sound pickup method involves simultaneously utilizing the microphone and the millimeter-wave radar for sound pickup, the method further includes:
[0046] The second audio signal is subjected to noise reduction and audio super-resolution processing to obtain the fourth audio signal corresponding to the microphone;
[0047] The third audio signal and the fourth audio signal are spliced together in the channel dimension to obtain the spliced dual-channel signal;
[0048] The dual signals are fused using a self-attention network to obtain the fused audio result. The fusion includes dynamically assigning different weights to the audio signal corresponding to the millimeter-wave radar and the audio signal corresponding to the microphone in the dual signals.
[0049] Secondly, a terminal device is provided, including:
[0050] processor;
[0051] Memory;
[0052] The memory stores a computer program, which includes instructions that, when executed by the processor, cause the terminal device to perform the method as described in any of the implementations of the first aspect above.
[0053] Thirdly, a chip system is provided, the chip system including a processing circuit, a receiving pin, and a transmitting pin; wherein the receiving pin, the transmitting pin, and the processing circuit communicate with each other through an internal connection path, and the processing circuit executes the method as described in any implementation of the first aspect above to control the receiving pin to receive signals and control the transmitting pin to transmit signals.
[0054] Fourthly, a computer-readable storage medium is provided that stores computer-executable program instructions, which, when executed on a computer, cause the computer to perform the method as described in any of the implementations of the first aspect above.
[0055] Fifthly, a computer program product is provided, the computer program product including computer program code, which, when run on a computer, causes the computer to perform the method as described in any of the implementations of the first aspect above. Attached Figure Description
[0056] Figure 1 is a schematic diagram of an application scenario for a sound pickup method provided in an embodiment of this application.
[0057] Figure 2 is a schematic structural diagram of a terminal device 100 provided in an embodiment of this application.
[0058] Figure 3 is a software structure block diagram of a terminal device 100 provided in an embodiment of this application.
[0059] Figures 4A and 4B are schematic diagrams showing the installation location of a millimeter-wave radar in a terminal device according to an embodiment of this application.
[0060] Figure 5 is a schematic diagram of the structure of a millimeter-wave radar in a terminal device according to an embodiment of this application.
[0061] Figures 6A to 6D are schematic diagrams of the GUI that may be involved in the implementation of a sound pickup method provided in the embodiments of this application.
[0062] Figure 7 is a flowchart of the underlying interaction process of a sound pickup method provided in an embodiment of this application.
[0063] Figure 8 is a schematic diagram of determining the sound source based on video footage according to an embodiment of this application.
[0064] Figures 9A and 9B are schematic diagrams illustrating how to determine the location of a sound-emitting object according to an embodiment of this application.
[0065] Figure 10 is a schematic flowchart of another sound pickup method provided in an embodiment of this application.
[0066] Figures 11A and 11B are schematic diagrams of audio signals before and after noise reduction processing and before and after super-resolution processing, provided in an embodiment of this application.
[0067] Figures 12A and 12B are schematic diagrams of some signal splicing and fusion provided in the embodiments of this application.
[0068] Figure 13 is a schematic flowchart of another sound pickup method provided in the embodiments of this application. Detailed Implementation
[0069] It should be noted that the terminology used in the implementation section of the embodiments of this application is only used to explain the specific embodiments of this application and is not intended to limit this application. In the description of the embodiments of this application, unless otherwise stated, " / " means "or", for example, A / B can mean A or B; "and / or" in this document is merely a description of the relationship between associated obstacles, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. In addition, in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more, "at least one" or "one or more" means one, two or more.
[0070] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature.
[0071] References to "one embodiment" or "some embodiments" as used in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0072] The technical solution of this application will be described in detail below with reference to specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0073] As described in the background section, most current terminal devices achieve audio pickup through microphone arrays. However, microphone arrays have limitations in pickup distance, resulting in poor pickup performance at long distances. With increasing demands for video recording, long-distance video shooting performance has become a key performance indicator for users. Relying solely on microphone arrays cannot provide audio quality that matches the video image quality, leading to a poor user experience.
[0074] In view of this, this application provides a method for sound pickup. When shooting a distant scene, a millimeter-wave radar installed in a terminal device picks up the distant audio, allowing the terminal device to overcome the limitation of sound pickup distance and provide the user with video that matches the audio. Specifically, when a user uses a terminal device to shoot a distant scene, they can choose to use a microphone array or a millimeter-wave radar for sound pickup based on the current zoom level or a distance manually set by the user. When millimeter-wave radar is selected for sound pickup, the target object and its angle in the shooting scene can be determined first based on preset characteristics. Then, the millimeter-wave radar emits a beam towards that angle to detect the distance between the target object and the terminal device. Next, the transmission power of the radar wave signal is determined based on the distance, and the signal is emitted in the direction of the target object. Based on the received echo signal modulated by the target object, the audio signal of the target object can be obtained through processing, that is, the audio corresponding to the video scene can be obtained.
[0075] The sound pickup method provided in this application can be applied to various terminal devices with sound pickup functions, such as mobile phones, cameras, tablet computers, personal computers (PCs), televisions, head-mounted displays (HMDs), augmented reality (AR) devices, mixed reality (MR) devices, in-vehicle electronic devices, laptop computers, monitoring equipment, robots, in-vehicle terminals, etc. For ease of description, the following embodiments use mobile phones as an example of terminal devices, but this application does not limit the specific type of terminal device.
[0076] The sound pickup method provided in this application can be applied to video shooting scenarios, especially long-distance video shooting scenarios. For ease of understanding, a possible application scenario is described below with reference to Figure 1. For example, Figure 1 is a schematic diagram illustrating an application scenario to which the sound pickup method provided in this application is applicable.
[0077] In video shooting scenarios, when the target object is close to the phone, the phone can simultaneously capture the video footage using the camera and pick up the target object's sound using a microphone array. In this case, due to the close proximity, a microphone array alone is sufficient for good sound pickup, eliminating the need for millimeter-wave radar.
[0078] As the shooting distance increases, the distance between the target object and the mobile phone also increases. When the distance increases to a certain point, if only a microphone array is used for sound pickup, the pickup distance of the microphone may limit the sound pickup effect, resulting in problems such as low audio volume or missing audio content. At this point, a combination of microphone array and millimeter-wave radar can be used to compensate for the microphone's inability to effectively pick up audio from a greater distance.
[0079] As the shooting distance increases, the distance between the target object and the mobile phone also increases. At this point, the video image mainly focuses on the target object at a distance, and the audio matching the captured image is also emitted by the target object at a distance. Therefore, the primary focus is on picking up the sound from the target object at a distance. Considering the power consumption and performance of the mobile phone, only millimeter-wave radar can be used to pick up and process the sound from the target object at a distance, without processing the audio collected by the microphone array.
[0080] In one possible implementation scenario, when picking up sound at a long distance, millimeter-wave radar can determine the angle and distance of the target object based on video footage, thereby determining its own sound pickup angle (or pickup range). The video capture range can include the sound pickup range; that is, the millimeter-wave radar can focus on the location of the target object and a small area adjacent to it within the video capture range, thus avoiding the collection of too much invalid sound (such as ambient noise) and improving the quality of sound pickup.
[0081] It should be noted that the long distance in the embodiments of this application can be set according to the sound pickup performance of the microphone array and / or millimeter-wave radar, and this application does not limit the specific value of the long distance.
[0082] As exemplarily shown in FIG2, it is a schematic structural diagram of a terminal device 100 provided in an embodiment of this application.
[0083] It should be noted that the structure of the terminal device 100 shown in the embodiment of Figure 2 is only an example. In actual applications, the terminal device 100 may have more or fewer components, and this application embodiment does not limit this.
[0084] The terminal device may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0085] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the terminal device. In other embodiments of this application, the terminal device may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0086] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0087] The controller can serve as the nerve center and command center of the terminal device. Based on the instruction opcode and timing signals, the controller generates operation control signals to control the fetching and execution of instructions.
[0088] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0089] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0090] It is understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are merely illustrative and do not constitute a structural limitation on the terminal device. In other embodiments of this application, the terminal device may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.
[0091] The charging management module 140 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 receives charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 receives wireless charging input via the wireless charging coil of the terminal device. While charging the battery 142, the charging management module 140 can also supply power to the terminal via the power management module 141.
[0092] The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, providing power to the processor 110, internal memory 121, external memory, display screen 194, camera 193, and wireless communication module 160, etc. The power management module 141 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 141 may also be located within the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may be located in the same device.
[0093] The wireless communication function of the terminal device can be implemented through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor, and baseband processor.
[0094] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the terminal device can be used to cover one or more communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with a tuning switch.
[0095] The mobile communication module 150 can provide solutions for wireless communication applications including 2G / 3G / 4G / 5G on terminal devices. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 150 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be housed in the same device.
[0096] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through an audio device (not limited to speaker 170A, receiver 170B, etc.) or displays images or videos through the display screen 194. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 110 and may be housed in the same device as the mobile communication module 150 or other functional modules.
[0097] The wireless communication module 160 can provide solutions for wireless communication applications on terminal devices, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.
[0098] In some embodiments, antenna 1 of the terminal device is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, enabling the terminal device to communicate with networks and other devices via wireless communication technology. The wireless communication technology may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time-Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies. The GNSS may include Global Positioning System (GPS), Global Navigation Satellite System (GLONASS), BeiDou Navigation Satellite System (BDS), Quasi-Zenith Satellite System (QZSS), and / or Satellite Based Augmentation Systems (SBAS).
[0099] The terminal device implements display functions through a GPU, a display screen 194, and an application processor. The display screen 194 is used to display images, videos, etc.
[0100] Terminal devices can achieve shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.
[0101] Digital signal processors (DSPs) are used to process digital signals, including digital image signals and other digital signals. For example, when a terminal device selects a frequency, a DSP performs Fourier transforms on the frequency energy. Video codecs are used to compress or decompress digital video. Neural processing units (NPUs) are neural network (NN) processors that, by borrowing from the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, rapidly process input information and can continuously learn on their own.
[0102] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the terminal device. The external memory card communicates with the processor 110 through the external memory interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external memory card. The internal memory 121 can be used to store computer executable program code, which includes instructions.
[0103] The terminal device can implement audio functions such as music playback and recording through an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, and an application processor.
[0104] The pressure sensor 180A senses pressure signals and converts them into electrical signals. The gyroscope sensor 180B determines the motion posture of the terminal device. The magnetic sensor 180D includes a Hall effect sensor. The terminal device can use the magnetic sensor 180D to detect the opening and closing of a flip cover. The accelerometer sensor 180E detects the magnitude of the terminal device's acceleration in various directions (typically three axes). When the terminal device is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the terminal's posture, applied to screen orientation switching, pedometers, and other applications. The proximity sensor 180G may include, for example, a light-emitting diode (LED) and a photodetector, such as a photodiode. The LED can be an infrared LED. The terminal device emits infrared light through the LED. The ambient light sensor 180L senses ambient light intensity. The terminal device can adaptively adjust the brightness of the display screen 194 based on the sensed ambient light intensity. The fingerprint sensor 180H collects fingerprints. The temperature sensor 180J detects temperature. The touch sensor 180K is also called a "touch panel". Touch sensor 180K can be placed on display screen 194. The touch sensor 180K and display screen 194 together form a touch screen, also known as a "touchscreen". Touch sensor 180K is used to detect touch operations applied to or near it. Bone conduction sensor 180M can acquire vibration signals.
[0105] In addition, the terminal device also includes a barometric pressure sensor 180C and a distance sensor 180F. The barometric pressure sensor 180C is used to measure air pressure. In some embodiments, the terminal device calculates altitude using the air pressure value measured by the barometric pressure sensor 180C to assist in positioning and navigation.
[0106] A distance sensor 180F is used to measure distance. The terminal device can measure distance via infrared or laser. In some embodiments, during a shooting scene, the terminal device can utilize the distance sensor 180F to measure distance for rapid focusing.
[0107] For example, the software system of the smart terminal 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application embodiment uses the layered architecture Android system as an example to illustrate the software structure of the smart terminal 100. Figure 3 is a software structure block diagram of a terminal device 100 provided in an embodiment of this application.
[0108] A layered architecture divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: application layer, application framework layer, Android runtime, system libraries, kernel layer, hardware abstraction layer (HAL), and hardware layer.
[0109] The application layer can include a series of application packages. As shown in Figure 3, application packages can include applications such as camera, calendar, map, WLAN, music, SMS, Bluetooth, video, social networking, gallery, and navigation.
[0110] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions. As shown in Figure 3, the application framework layer may include a window manager, content provider, phone manager, resource manager, notification manager, view system, etc.
[0111] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.
[0112] Content providers store and retrieve data, making that data accessible to applications. This data may include videos, images, audio, made and received phone calls, browsing history and bookmarks, phone books, etc.
[0113] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.
[0114] The phone manager is used to provide communication functions for the smart terminal 100. For example, it manages call status (including connection, hang-up, etc.).
[0115] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.
[0116] The notification manager allows applications to display notifications in the status bar. These can be used to convey informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of download completion or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating the device, or flashing indicator lights.
[0117] The Android Runtime consists of core libraries and a virtual machine. The Android runtime is responsible for the scheduling and management of the Android system.
[0118] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.
[0119] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as obstacle lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0120] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.
[0121] The Surface Manager is used to manage the display subsystem and provides the blending of 2D and 3D layers for multiple applications.
[0122] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.
[0123] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0124] A 2D graphics engine is a graphics engine for 2D drawing.
[0125] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera sensor driver, audio driver, and sensor driver.
[0126] As exemplarily shown in Figures 4A and 4B, this is a schematic diagram illustrating the installation location of a millimeter-wave radar in a terminal device according to an embodiment of this application.
[0127] The terminal device in this embodiment (hereinafter, a mobile phone, is used as an example) may be equipped with a microphone or a microphone array including multiple microphones, millimeter-wave radar, and a camera. The camera may include a front-facing camera installed inside the front casing of the phone and a rear-facing camera installed inside the rear casing of the phone. Considering the characteristic that microphones have good close-range sound pickup but poor long-range sound pickup, the microphone or microphone array in the mobile phone can be used to pick up audio at close range. When the pickup distance is greater than a certain range, millimeter-wave radar can be used for sound pickup.
[0128] In some embodiments, the millimeter-wave radar can be fixedly installed on the inside or outside of the back cover of the mobile phone, and it can be located at any position in the back cover area of the mobile phone.
[0129] Alternatively, in some other embodiments, the millimeter-wave radar can be positioned within an area less than a first distance from the rear camera. For example, as shown in Figure 4A, taking a circular rear camera area (or lens area or deco area) as an example, the millimeter-wave radar can be positioned within an annular area adjacent to the lens area, with the center of the circular lens area as the origin and a radius of the first distance D (the shaded area shown in Figure 4A). In this case, the millimeter-wave radar is located near the deco area of the phone, and its distance from the camera (or telephoto camera) can be less than the first distance D.
[0130] Alternatively, in some other embodiments, the millimeter-wave radar can be positioned within the lens area of the phone's rear camera. For example, as shown in Figure 4B, still taking a circular rear camera area as an example, the millimeter-wave radar can be positioned inside the lens area where the rear camera is located (i.e., inside the deco area). In this case, the distance between the millimeter-wave radar and the rear camera, especially the telephoto camera, is very short, and the angle of the target object relative to the camera and its angle relative to the millimeter-wave radar will not differ too much, facilitating subsequent sound pickup of the target object by the millimeter-wave radar.
[0131] It should be noted that when using millimeter-wave radar to pick up audio from a distance, it is necessary to determine the angle of the target object emitting the sound based on the video footage captured by the camera so that the millimeter-wave radar can determine the beam direction. Therefore, placing the millimeter-wave radar close to the rear camera (especially the telephoto camera), as shown in Figure 4A, where the distance between the millimeter-wave radar and the telephoto camera is less than the first distance D, or as shown in Figure 4B, placing the millimeter-wave radar in the lens area of the rear camera, can reduce the error in obtaining the beam direction based on the angle of the target object in the video footage, thereby improving the accuracy of target object detection and sound pickup quality.
[0132] In some embodiments, as shown in Figure 5, the millimeter-wave radar can be deployed on a printed circuit board (PCB) or a flexible printed circuit board (FPCB) located inside the back cover of the mobile phone. The PCB or FPCB can be glued to the inside of the back cover, specifically to the area of the rear camera (e.g., inside the raised area where the rear lens is located or inside the glass cover of the deco area). The millimeter-wave radar can connect its antenna to a connector via a wire, and then connect the antenna to a socket on the radar chip in the mobile phone's motherboard via the connector. Alternatively, the millimeter-wave radar can be directly deployed on the radar chip in the mobile phone's motherboard. This radar chip can be connected to other modules in the motherboard via wires to enable data transmission and power supply between the millimeter-wave radar, the radar chip, and other modules in the motherboard. Alternatively, the millimeter-wave radar can be deployed on the outside of the mobile phone casing. In this case, for aesthetic purposes, decorative elements can be added to cover the surface of the millimeter-wave radar.
[0133] In some embodiments, a millimeter-wave radar may include one or more (i.e., 1 to n, where n is an integer greater than 1) transmitting antennas and one or more receiving antennas. The transmitting and receiving antennas can be fixed to pads on the millimeter-wave radar PCB or FPCB board, such as by soldering. Alternatively, the transmitting and receiving antennas in the millimeter-wave radar can be shared instead of being separately configured; this application does not limit this approach.
[0134] To improve the flexibility of sound pickup functionality in terminal devices, the sound pickup method provided in this application embodiment may include multiple control modes, such as automatic mode and manual mode (or professional mode). In automatic mode, the phone can automatically adjust the sound pickup method based on the current shooting distance. For example, when the shooting distance is less than a first distance threshold, only the microphone array is used for sound pickup; when the shooting distance is greater than or equal to the first threshold and less than a second threshold, a combination of microphone array and millimeter-wave radar is used for sound pickup; when the shooting distance is greater than the second threshold, millimeter-wave radar is used for sound pickup.
[0135] As an example, Table 1 below shows an example of the correspondence between shooting distance (or sound pickup distance) and sound pickup method during shooting.
[0136] Table 1
[0137] The values corresponding to the first threshold and the second threshold increase sequentially. Their specific values can be flexibly set according to the hardware performance of the mobile phone, such as one or more of the sound pickup performance of the microphone array, the camera performance, the sound pickup performance of the millimeter-wave radar, etc. This application embodiment does not limit this.
[0138] It should be noted that the shooting distance in this application can refer to the distance between the camera and the subject being photographed when the mobile phone takes a picture, that is, the actual measured distance at the time of shooting. However, in the sound pickup method provided in the embodiments of this application, the mobile phone determines which method to use for sound pickup not only based on the actual measured shooting distance. Considering that in a typical shooting scenario, the shooting distance is proportional to the zoom distance, in some embodiments, the mobile phone can also determine which sound pickup method to use based on the zoom distance at the time of shooting. For example, if the zoom distance at the time of shooting is small, it means that the current shooting distance may be relatively close, and in this case, only a microphone array can be used for sound pickup; if the zoom magnification at the time of shooting is large (e.g., greater than the first magnification), it means that the current shooting distance is large, and using only a microphone array for sound pickup may easily result in poor sound pickup effect, so in this case, a combination of a microphone array and millimeter-wave radar can be used for sound pickup; if the zoom magnification at the time of shooting is greater than a certain threshold (e.g., greater than the second magnification), it means that the current shooting is at a long distance, and in this case, only millimeter-wave radar can be used for sound pickup.
[0139] As an example, Table 2 below shows an example of the correspondence between zoom distance and sound pickup method during shooting.
[0140] Table 2
[0141] The values corresponding to the first multiplier and the second multiplier increase sequentially. The specific values can be flexibly set according to the hardware performance of the mobile phone, such as one or more of the following: the adaptability of the microphone array, the performance of the camera, the sound pickup performance of the millimeter-wave radar, etc. For example, the first multiplier can be set to 3, and the second multiplier can be set to 8, etc. This application embodiment does not limit this.
[0142] To better understand the different pickup mode control in the pickup method provided in the embodiments of this application, the pickup process under different pickup modes will be described below with reference to the accompanying drawings.
[0143] For example, as shown in Figures 6A to 6D, these are schematic diagrams of the graphical user interface (GUI) that may be involved in the implementation of a sound pickup method provided in the embodiments of this application.
[0144] In some embodiments, the audio pickup mode control can be set in the camera application. For example, as shown in Figure 6A, this is the video recording preview interface of a mobile phone. This video recording preview interface may also include an audio pickup mode selection control 601 and other shooting parameter adjustment controls. Furthermore, the video recording preview may include a video recording preview frame, a zoom control 602, shooting mode controls (such as aperture mode, night mode, portrait mode, video mode, photo mode, professional mode, and more), a video recording start / stop control, a camera rotation control, and a quick access to the photo album control.
[0145] In some embodiments, when the mobile phone receives an operation (such as a click) input from the user to the microphone mode selection control 601, the mobile phone can display a microphone mode display interface as shown in Figure 6B. This microphone mode display interface can include the microphone modes supported by the mobile phone and a description of each microphone mode. For example, the microphone mode display interface can show that the microphone modes currently supported by the mobile phone include automatic mode 603 and manual mode 604. Automatic mode can refer to "in automatic mode, the mobile phone can automatically switch the microphone mode according to the zoom distance, actual shooting distance, etc."; manual mode can refer to "in manual mode, the mobile phone can determine the microphone mode according to the distance adjusted by the user."
[0146] In some embodiments, when the user selects automatic mode for audio pickup, the mobile phone can display the automatic mode interface as shown in Figure 6C, and in response to the user's selection of automatic mode for audio pickup, the mobile phone can automatically adjust the audio pickup method according to the zoom level of the current recording. For example, as shown in Figure 6C, the automatic audio pickup mode interface may also include a sound spectrum diagram corresponding to the currently picked-up audio, which can be used to indicate the spectral changes of the current audio.
[0147] In some embodiments, when the user selects manual mode for audio pickup, the mobile phone can display a manual mode interface as shown in Figure 6D. In response to the user's selection of automatic mode for audio pickup, the mobile phone can adjust the pickup method according to the currently selected pickup distance. For example, as shown in Figure 6D, the manual pickup mode interface may also include a sound spectrum diagram corresponding to the currently picked-up audio, which can be used to indicate the spectrum changes of the current audio. The interface may also include a pickup distance adjustment bar 605, which can receive the user's adjustment operation. This adjustment operation can be used to adjust the mobile phone's pickup distance, allowing the mobile phone to select the appropriate pickup method based on the pickup distance. The adjustment operation may include sliding the pickup distance adjustment control to the left (towards) or to the right (towards) in the pickup distance adjustment bar. The pickup distance adjustment control may be a circular control as shown in Figure 6D. Furthermore, the pickup distance here may correspond to the shooting distance in this embodiment, such as the shooting distance in Table 1 above.
[0148] It should be noted that the GUI interfaces and the information included in the interfaces shown in Figures 6A to 6D are merely examples. In practical applications, the GUI interface may include more or less information, and this application embodiment does not limit this. Furthermore, the embodiments in Figures 6A to 6D only illustrate the setting of the sound pickup mode selection control in the preview interface of the camera app. In practical applications, the sound pickup mode selection control may also be set in the "More" directory of the camera app, or in the "Pro Mode" of the camera app, or in the "Settings App" of the phone, etc. This application embodiment does not limit this.
[0149] As exemplarily shown in Figure 7, this is a flowchart illustrating the underlying interaction process of a sound pickup method provided in an embodiment of this application. Specifically, it may include the following steps:
[0150] S701 allows users to select the pickup mode through a user interface (UI).
[0151] The UI here can be a sound pickup mode selection interface. For example, this UI interface can correspond to the sound pickup mode selection interface shown in Figure 6B.
[0152] In some embodiments, users can select either automatic or manual sound pickup mode by clicking on the UI interface. For example, a user can click on the automatic mode 603 control shown in Figure 6B to select automatic sound pickup mode; or, a user can click on the manual mode 604 control shown in Figure 6B to select manual sound pickup mode.
[0153] It should be understood that the method of determining the sound pickup mode through user input selection is merely an example. In practical applications, the sound pickup mode can also be determined in other ways, such as the user selecting the sound pickup mode through voice input, etc. This application embodiment does not limit this.
[0154] S702, the UI sends a sound pickup mode selection operation to the sound pickup algorithm module.
[0155] In some embodiments, the audio pickup algorithm module may be located in the hardware abstraction layer (HAL). This audio pickup algorithm module can be used to control the switching of audio pickup modes and to process the audio data picked up by the microphone and / or millimeter-wave radar.
[0156] In some embodiments, the sound pickup algorithm module can determine whether the currently used sound pickup mode is automatic or manual based on the sound pickup mode selection operation sent by the UI interface. If the current sound pickup mode is automatic, the sound pickup algorithm module can execute step S703, which involves obtaining the zoom level sent by the recording app.
[0157] Optionally, when the sound pickup mode is in automatic mode, before executing step S703, the sound pickup algorithm module can also send zoom magnification query information to the recording app to query the zoom magnification of the current recording.
[0158] Optionally, when the audio pickup mode is set to automatic, the UI can also send the user's selection of the audio pickup mode to the recording app, informing the recording app which audio pickup mode the user has selected. When the recording app receives confirmation that the user has selected automatic audio pickup mode, it can send the current zoom level to the audio pickup algorithm module.
[0159] In some embodiments, when the zoom level of the recording changes, the recording app can send the changed zoom level to the audio pickup algorithm module in real time.
[0160] It should be understood that when the zoom level changes, the recording app sends the changed zoom level to the audio pickup algorithm module in real time. This allows the audio pickup algorithm module to adjust the audio pickup method in a timely manner according to the current zoom level, such as switching the audio pickup method from microphone pickup to millimeter-wave radar pickup, thereby ensuring the audio quality of the recording.
[0161] In some embodiments, the sound pickup algorithm module can determine which sound pickup method to use based on the received zoom level and the correspondence between zoom level and sound pickup method. The correspondence between zoom level and sound pickup method can be found in Table 2 above, and will not be repeated here.
[0162] In some embodiments, when the sound pickup mode is manual, the sound pickup algorithm module can determine which sound pickup mode to use based on the sound pickup distance selected by the user and the correspondence between the sound pickup distance and the sound pickup mode. The sound pickup distance can correspond to the shooting distance. The method for the user to select the sound pickup distance can be found in the description of the embodiment in Figure 6D above. Furthermore, the correspondence between the zoom magnification and the sound pickup mode can be found in Table 1 above, and will not be repeated here.
[0163] It should be noted that in manual mode, the audio pickup algorithm module does not need to query the recording app or obtain the zoom level of the current recording. Instead, it can obtain the pickup distance input by the user through the UI interface. Specifically, the user can input the pickup distance through the UI interface (as shown in Figure 6D), and the UI interface can send this pickup distance to the audio pickup algorithm module (this step is not shown in Figure 7), so that the audio pickup algorithm module can determine the pickup method based on the pickup distance.
[0164] It should also be noted that, regardless of whether it is in automatic or manual mode, when the recording mode is turned on, the microphone pickup function of the electronic device can be turned on by default. That is, after the recording starts, the microphone also starts to collect audio and can execute step S704, that is, the microphone sends the microphone pickup data to the pickup algorithm module.
[0165] In some embodiments, if the sound pickup algorithm module determines that the sound pickup method is through a microphone, then after the sound pickup algorithm module receives the sound pickup data sent by the microphone, it can process the sound pickup data to obtain an audio output result that matches the video recording screen. Then, step S709 is executed, that is, the sound pickup algorithm module sends the audio processing result to the recording App.
[0166] In some embodiments, if the sound pickup algorithm module determines that the sound pickup method is sound pickup via millimeter-wave radar, or sound pickup via millimeter-wave radar, then the sound pickup algorithm module may execute the following step S705.
[0167] S705, the sound pickup algorithm module sends millimeter-wave radar sound pickup indication information to the millimeter-wave radar chip.
[0168] The millimeter-wave radar indication information is used to instruct the millimeter-wave radar chip to control the millimeter-wave radar to pick up sound.
[0169] In some embodiments, in response to audio pickup indication information, the millimeter-wave radar chip can first determine the beam direction based on the currently recorded image. It should be understood that the beam here can be used to locate the source of the sound.
[0170] In some embodiments, when it is determined that sound pickup needs to be performed using millimeter-wave radar, the sound pickup algorithm module can first determine the angle of the sound-emitting object based on the current video footage, and then determine the beam direction of the millimeter-wave radar. Specifically, the process of determining the angle of the sound-emitting object may include: acquiring a video footage captured by a camera; performing feature recognition on the video footage to determine features that match preset sound-emitting object features, and determining the sound-emitting object based on the matching features; determining the position of the sound-emitting object in a preset reference coordinate system; and determining the angle of the sound-emitting object relative to the electronic device (or camera) based on the position of the sound-emitting object in the reference coordinate system. Here, the reference coordinate system can be, for example, a two-dimensional coordinate system or a three-dimensional coordinate system, and its origin can be, for example, a fixed position within the electronic device, such as the location of the camera or the millimeter-wave radar.
[0171] As exemplarily shown in Figure 8, this application provides a schematic diagram of determining a sound-emitting object based on a video recording. By setting a reference coordinate system with the electronic device as the reference, when it is determined that a sound-emitting object exists in the video recording, its position in the reference coordinate system can be determined based on the position of the sound-emitting object in the video recording, and the angle of the sound-emitting object relative to the electronic device can be determined based on that position.
[0172] In some embodiments, the electronic device can pre-acquire features that characterize the sound-producing object and store these features in local space. For example, if the sound-producing object is a person, the corresponding features can be any information that can characterize the human body, such as body parts (e.g., limbs, head), standing posture, body shape, facial features, etc. Then, the electronic device can acquire information matching the human body features from the recorded video and determine the position of the human body in the recorded video, thereby obtaining the angle of the human body relative to the electronic device.
[0173] In some embodiments, after determining the angle of the sound-emitting object, the millimeter-wave radar can adjust the beam direction so that the beam is directed at the angle of the sound-emitting object, and use the beam to detect the distance between the sound-emitting object and the electronic device.
[0174] In some embodiments, the method by which millimeter-wave radar determines the distance between a sound-emitting object and an electronic device may include: the millimeter-wave radar detecting, through the transmitted beam, whether there is an object within the current beam coverage area that matches the characteristics of the sound-emitting object, such as whether there is an object with vibration characteristics; if so, determining that the object is the sound-emitting object. Then, using the transmitted beam and the signal reflected back from the sound-emitting object, the distance between the sound-emitting object and the electronic device can be determined.
[0175] It should be understood that once the angle of the sound-emitting object relative to the electronic device and the distance between the sound-emitting object and the electronic device are determined, the position of the sound-emitting object can be determined.
[0176] The S706 millimeter-wave radar transmits millimeter-wave signals via a transmitting antenna.
[0177] This millimeter-wave signal can correspond to the transmitted signal of a millimeter-wave radar.
[0178] In some embodiments, after the millimeter-wave radar determines the distance between the sound-emitting object and the electronic device, the millimeter-wave radar can adjust its transmission power according to the distance of the sound-emitting object, so that the transmission power matches the distance of the sound-emitting object, thereby improving the sound pickup. For example, the transmission power of the millimeter-wave radar can be directly proportional to the distance of the sound-emitting object, that is, the greater the distance between the sound-emitting object and the electronic device, the greater the transmission power can be.
[0179] In some embodiments, an electronic device can transmit millimeter-wave signals to the angle of the sound-emitting object using millimeter-wave radar. The coverage area of the millimeter-wave signal can include the sound-emitting object, as shown in Figures 9A and 9B. Figure 9A shows a top view of the millimeter-wave signal coverage area, and Figure 9B shows a side view of the millimeter-wave signal coverage area. Referring to the figures, the coverage area of the millimeter-wave signal emitted by the millimeter-wave radar can be larger than the area occupied by the sound-emitting object; that is, the coverage area of the millimeter-wave signal only needs to include the sound-emitting object. Furthermore, the circuit structure for sound pickup by the millimeter-wave radar can be an existing circuit structure; for example, see the enlarged circuit diagrams shown in Figures 9A and 9B, which will not be described in detail in this embodiment.
[0180] It should be noted that the radar waveform corresponding to the millimeter-wave signal emitted by the millimeter-wave radar for the sound-emitting object can be flexibly set, and the embodiments of this application do not limit this.
[0181] The S707 is a millimeter-wave radar chip that receives echo signals transmitted by the receiving antenna.
[0182] In some embodiments, when the millimeter-wave signal (or transmitted signal) emitted by the millimeter-wave radar reaches the sound-emitting object, it will be modulated by the vibration information of the sound-emitting object, and the modulated echo signal will be returned to the receiving antenna. At this time, the millimeter-wave radar can obtain the echo signal modulated by the vibration information of the sound-emitting object through the receiving antenna.
[0183] It should be noted that the modulation here specifically refers to the following: When the millimeter-wave radar transmits a millimeter-wave signal, the transmitted signal has a transmission phase. When the transmitted signal reaches the sound-producing object, the vibrating part of the object (such as the human throat) causes a phase change in the transmitted signal. This phase change represents the vibration phenomenon (i.e., the sound production phenomenon) of the vibrating object. Subsequently, the echo signal received by the millimeter-wave radar also carries this phase change. In other words, the echo signal is a signal modulated by the vibrating part of the sound-producing object, and this phase change information represents the vibration characteristics of the sound-producing object.
[0184] After receiving the echo signal, the millimeter-wave radar sends the sound pickup data of the millimeter-wave radar to the sound pickup algorithm module.
[0185] The acoustic data picked up by the millimeter-wave radar here can correspond to the echo signal and / or transmitted signal data mentioned above.
[0186] In some embodiments, the audio pickup algorithm module can be used to process the audio data picked up by millimeter-wave radar. Considering that the audio quality acquired by millimeter-wave radar may not be high enough, after acquiring the audio data from the millimeter-wave radar, the audio pickup algorithm module can enhance the audio signal through denoising and super-resolution algorithms. The specific audio processing procedure will be described in more detail below and will not be elaborated here.
[0187] In some embodiments, if the sound pickup method is a hybrid method, i.e., sound pickup using both a microphone and millimeter-wave radar, then after the sound pickup algorithm module acquires the sound pickup data sent by the microphone and the sound pickup data sent by the millimeter-wave radar, the audio signals acquired by the microphone and the audio signals acquired by the millimeter-wave radar can be fused. Specifically, neural networks can be used to process the two audio signals, such as convolutional neural networks (CNN), recurrent neural networks (RNN), and self-attention mechanisms (transformers). The processed audio signal can have better quality and can be used as the final output.
[0188] Then, step S709 can be executed to send the audio processing result to the recording app.
[0189] In addition, after receiving the audio data from the millimeter-wave radar and / or the microphone, the audio pickup algorithm module can also perform spectrum analysis on the audio data acquired by the millimeter-wave radar and / or the audio data acquired by the microphone, and execute step S710, that is, send the real-time audio spectrum analysis results to the UI interface.
[0190] In some embodiments, the UI interface can display the real-time spectrum changes corresponding to the current sound pickup, as shown in the sound spectrum diagram in Figure 6C or Figure 6D.
[0191] According to the sound pickup method provided in the embodiments of this application, when shooting a distant scene, the millimeter-wave radar installed in the terminal device picks up the distant audio, so that the terminal device breaks through the limitation of the sound pickup distance and provides the user with a video that matches the picture and audio, thereby improving the video quality and enhancing the user's recording experience.
[0192] The following, in conjunction with the accompanying diagrams, provides a more detailed explanation of the sound pickup and audio processing processes involved in hybrid sound pickup methods and standalone millimeter-wave radar sound pickup methods.
[0193] As exemplarily shown in Figure 10, this is a schematic flowchart of another sound pickup method provided in an embodiment of this application. Specifically, it may include the following steps:
[0194] S1001, Beam Scan Positioning.
[0195] Specifically, this step can include the following processes:
[0196] S1001a, determine the beam direction based on the video footage.
[0197] In some embodiments, before determining the beam direction based on the video frame, the electronic device may also acquire the sound pickup method, such as determining the sound pickup method based on the current zoom level or the sound pickup distance selected by the user. If the sound pickup method includes sound pickup via millimeter-wave radar (e.g., a hybrid sound pickup method or a method using millimeter-wave radar alone), then the currently recorded video frame is acquired, and the beam direction is determined based on the video frame. The method for determining the beam direction based on the video frame can be found in the description in the embodiment shown in Figure 7 above, and will not be repeated here.
[0198] S1001b, transmit beam to detect the sound source.
[0199] In some embodiments, when the sound-emitting object is determined to be located at a first angle relative to the electronic device, the millimeter-wave radar can send a beam in a first direction.
[0200] It should be noted that the beam transmitted in this step can be used to detect the distance between the sound-emitting object and the electronic device. Specifically, the millimeter-wave radar can detect whether there is an object within the beam's coverage area that matches the preset vibration characteristics. If so, it indicates that there is a sound-emitting object within the beam's coverage area, and the distance between the sound-emitting object and the electronic device can then be obtained.
[0201] S1001c adjusts the transmission power according to the distance of the detected sound source.
[0202] The transmission power here is used to pick up sound from the source of the sound.
[0203] It should be understood that the greater the distance between the sound-emitting object and the electronic device, the greater the required transmission power. Therefore, when the distance between the sound-emitting object and the electronic device is determined (which may correspond to the second distance in this application), the transmission power can be determined based on this distance. Different distances can correspond to different transmission powers, that is, there can be a preset mapping relationship (such as a direct proportional relationship) between the distance and the transmission power. The millimeter-wave radar chip can determine the transmission power based on the distance between the sound-emitting object and the electronic device and the mapping relationship.
[0204] S1002, acquire the first audio signal.
[0205] In some embodiments, the first audio signal may refer to the sound signal acquired by millimeter-wave radar, that is, the sound signal corresponding to the sound-emitting object at a distance.
[0206] Specifically, this step may further include the following processes:
[0207] S1002a, a millimeter-wave radar that transmits millimeter-wave signals.
[0208] Millimeter wave signals can also be described as transmitted signals.
[0209] In some embodiments, the method by which the millimeter-wave radar transmits millimeter-wave signals may include: transmitting millimeter-wave signals with a target transmission power and a first phase, wherein the target transmission power may be a transmission power determined based on the distance between the sound-emitting object and the electronic device, and the first phase is used to direct the transmitted millimeter-wave signal in the direction of the sound-emitting object (which may be referred to as the target direction).
[0210] S1002b, the sound-emitting object modulates a millimeter-wave signal.
[0211] In some embodiments, when a millimeter-wave signal transmitted in a first phase reaches a sound-generating object, the vibrating part of the sound-generating object (such as a human throat) causes a phase change in the millimeter-wave signal, such as a change from the first phase to the second phase. This phase change represents the vibration phenomenon (i.e., the sound generation phenomenon) of the sound-generating object. Subsequently, when the millimeter-wave radar receives the echo signal with the second phase (i.e., the signal modulated by the vibrating part of the sound-generating object), it can obtain the vibration characteristics of the sound-generating object based on the phase change information and recover the audio signal of the sound-generating object (corresponding to the first audio signal).
[0212] S1002c, a millimeter-wave radar that receives echo signals.
[0213] S1002d demodulates the echo signal to recover the first audio signal of the sound-producing object.
[0214] In some embodiments, the millimeter-wave radar chip can demodulate the echo signal based on the first phase and the second phase to obtain the first audio signal corresponding to the sound-emitting object.
[0215] For example, the process of acquiring the first audio signal of the sound-producing object may include:
[0216] (1) The millimeter-wave radar transmits a single-frequency continuous wave or a linear frequency modulated signal. Taking the single-frequency continuous wave as an example, the single-frequency signal transmitted by this millimeter-wave radar can be represented by the following formula (1-1):
[0217] Among them, S t Indicates the transmitted signal; A is the amplitude of the transmitted signal; f c The carrier frequency for transmitting the signal; The phase of the transmitted signal (which may correspond to the first phase).
[0218] (2) Taking a human as the source of the sound, when a human speaks, the throat vibrates, and this vibration causes a phase change in the signal. The echo signal can be represented by the following formula (1-2): Position change; γ represents wavelength.
[0219] (3) The received echo signal is mixed, and the mixed signal is low-pass filtered to obtain the baseband signal. This baseband signal can be represented by the following formula (1-3):
[0220] Among them, S b (t) represents the baseband signal; C is the signal amplitude of the baseband signal; Δφ represents the phase change; This indicates the phase change caused by throat vibration; γ represents the wavelength.
[0221] In some embodiments, when Δφ is When the value is an odd multiple of , the vibration signal can be approximated by the following formula (1-4):
[0222] Where x(t) represents the vibration signal (which may correspond to the first audio signal).
[0223] S1003, enhance the first audio signal and obtain the enhanced third audio signal.
[0224] Specifically, this step may further include the following processes:
[0225] S1003a performs noise reduction processing on the first audio signal.
[0226] S1003b performs audio super-resolution processing on the first audio signal after denoising.
[0227] It should be noted that since the lost audio from millimeter-wave radar may contain some interference, noise reduction methods can be used to remove the noise from the audio. Furthermore, considering that the audio frequency captured by millimeter-wave radar may be low, speech super-resolution methods can be used to recover the high-frequency signals of the sound, thereby achieving better speech quality and intelligibility. The noise reduction methods used for removing noise and the super-resolution methods used for recovering the high-frequency signals of the sound can be any existing feasible methods, and this application does not limit them.
[0228] For example, the audio received by millimeter-wave radar without noise reduction and the signal frequency without super-resolution processing can be seen in Figure 11A; the audio signal after noise reduction and the signal frequency after super-resolution processing can be seen in Figure 11B. Through actual measurements of the audio signals without noise reduction and without super-resolution processing, as well as the audio signals after noise reduction and after super-resolution processing, it can be concluded that the audio signal after noise reduction is a cleaner representation of the sound signal of the emitting object, and the frequency after super-resolution processing contains more high-frequency signals.
[0229] S1004, process the second audio signal and obtain the processed fourth audio signal.
[0230] The second audio signal is the sound signal picked up by the microphone.
[0231] In some embodiments, when recording begins, the microphone can be turned on by default to pick up sound. The microphone can send the picked-up second audio signal to the sound pickup algorithm module; then, if the current sound pickup method includes sound pickup through the microphone, the sound pickup algorithm module can process the second audio signal and obtain the processed fourth audio signal.
[0232] In some embodiments, the sound pickup algorithm module may process the second audio signal in ways such as noise reduction and / or super-resolution processing. These processes are used to obtain higher quality audio signals, and the specific processing methods are not limited in the embodiments of this application.
[0233] S1005 fuses the third and fourth audio signals.
[0234] It should be noted that this step is optional. This step can be performed when the sound pickup method is a hybrid method (i.e., sound pickup using both a microphone and millimeter-wave radar). If a microphone is used for sound pickup instead of millimeter-wave radar, or vice versa, this step can be skipped.
[0235] Specifically, this step can include the following processes:
[0236] S1005a splices the third and fourth audio signals in the channel dimension.
[0237] The S1005b employs a self-attention mechanism to fuse dual signals.
[0238] For example, the process of splicing the third and fourth audio signals and fusing the two signals through a self-attention network can be seen in the flowcharts shown in Figures 12A and 12B.
[0239] In some embodiments, splicing the third and fourth audio signals in the channel dimension can specifically include: splicing one signal corresponding to the millimeter-wave radar (such as a 1×n third audio signal) and one signal corresponding to the microphone (such as a 1×n fourth audio signal) in the channel dimension to obtain a 2×n signal. That is, splicing the one-dimensional microphone signal and the one-dimensional millimeter-wave signal in the channel dimension into a two-dimensional signal.
[0240] In some embodiments, after acquiring the spliced signal, it can be input into a self-attention network for joint processing. The self-attention network can acquire the signal shared by the microphone signal and the millimeter-wave radar signal, and enhance this shared signal. Specifically, the processing method of the self-attention network is as follows: during fusion, weights can be dynamically assigned to the two audio signals according to the characteristics of the audio signals from the millimeter-wave radar and the microphone. For example, assuming that the audio signal corresponding to the millimeter-wave radar is of higher quality at a certain time period, then a larger weight can be assigned to the audio signal corresponding to the millimeter-wave radar, while a smaller weight can be assigned to the audio signal corresponding to the microphone; while in another frequency band, if the audio signal corresponding to the microphone is of higher quality, then a larger weight can be assigned to the audio signal corresponding to the microphone, while a smaller weight can be assigned to the audio signal corresponding to the millimeter-wave radar. Simply put, the self-attention network can acquire the more important (or higher quality) audio signal, assign a larger weight to the more important audio signal, and assign a smaller weight to the other less important audio signal, and then sum these two signals together to achieve the effect of highlighting the more important audio.
[0241] In some embodiments, after fusing the two signals using a self-attention mechanism, the fused audio signal can be obtained, which is the processed audio result.
[0242] S1005c stores the processed audio results.
[0243] S1006, Audio Output.
[0244] In some embodiments, the audio output may specifically refer to the audio result being processed by the sound pickup algorithm module and then output to the recording app. After obtaining the audio result, the recording app can combine the audio result with the video footage to obtain a complete video.
[0245] According to the sound pickup method provided in the embodiments of this application, when shooting a distant scene, the millimeter-wave radar installed in the terminal device picks up the distant audio, and the microphone picks up the audio at close range, so that the terminal device breaks through the limitation of the sound pickup distance and provides the user with a video that matches the picture and audio, thereby improving the video quality and enhancing the user's recording experience.
[0246] It should be noted that, for ease of description, this application uses a mobile phone as an example for the terminal device. However, in practical applications, the type of terminal device is not limited to this; for example, it can also be a tablet computer, a wearable device, etc. When the terminal device is a tablet computer, the millimeter-wave radar can also be positioned at a distance less than the first distance from the camera in the tablet computer. Furthermore, in the descriptions of the sound-emitting object in this application, people are often used as an example; however, in practical applications, the type of sound-emitting object is not limited to this, and any sound-emitting object can be used. This application does not impose any limitations on this.
[0247] For example, Figure 13 shows a schematic flowchart of another sound pickup method provided in an embodiment of this application. The execution entity of this process can be a terminal device including a microphone and a millimeter-wave radar, and specifically may include the following steps:
[0248] S1301, obtain the zoom level of the current recording or the pickup distance input by the user.
[0249] S1302, the sound pickup method is obtained according to the zoom magnification or the sound pickup distance. The sound pickup method includes using millimeter-wave radar to pick up sound alone, using a microphone to pick up sound alone, and using a microphone and millimeter-wave radar to pick up sound simultaneously.
[0250] S1303, process the first audio signal picked up by the millimeter-wave radar and / or the second audio signal picked up by the microphone according to the sound pickup method, and obtain the processed audio result.
[0251] In some embodiments, the terminal device includes a microphone and a millimeter-wave radar, specifically including: the terminal device includes a camera; the distance between the millimeter-wave radar and the camera is less than or equal to a first distance.
[0252] In some embodiments, the terminal device is a mobile phone, the camera is a rear camera in the mobile phone, and the millimeter-wave radar is located in the camera area corresponding to the rear camera.
[0253] In some embodiments, obtaining the sound pickup method based on the zoom magnification specifically includes: when the zoom magnification is less than a first magnification, obtaining the sound pickup method as using a microphone alone; when the zoom magnification is greater than or equal to the first magnification and less than a second magnification, obtaining the sound pickup method as using both the microphone and the millimeter-wave radar simultaneously; when the zoom magnification is greater than or equal to the second magnification, obtaining the sound pickup method as using only the millimeter-wave radar; wherein the value corresponding to the second magnification is greater than the value corresponding to the first threshold.
[0254] In some embodiments, obtaining the sound pickup method based on the sound pickup distance specifically includes: when the sound pickup distance is less than a first threshold, obtaining the sound pickup method as using a microphone alone; when the sound pickup distance is greater than or equal to the first threshold and less than a second threshold, obtaining the sound pickup method as using both the microphone and the millimeter-wave radar simultaneously; when the sound pickup distance is greater than or equal to the second threshold, obtaining the sound pickup method as using only the millimeter-wave radar; wherein, the second threshold is greater than the first threshold.
[0255] In some embodiments, the method further includes: receiving a first operation input by a user, the first operation being used to select a sound pickup mode, the sound pickup mode including an automatic sound pickup mode and a manual sound pickup mode; when the sound pickup mode is an automatic sound pickup mode, obtaining the zoom magnification and obtaining a sound pickup method based on the zoom magnification; when the sound pickup mode is a manual sound pickup mode, receiving a second operation input by a user, the second operation being used to select the sound pickup distance; and obtaining a sound pickup method based on the sound pickup distance.
[0256] In some embodiments, when the sound pickup method involves using the millimeter-wave radar alone or using both the microphone and the millimeter-wave radar simultaneously, the method further includes: acquiring a currently recorded video frame; acquiring a target angle of the sound-emitting object relative to the terminal device based on preset first feature information of the sound-emitting object and the video frame; setting the beam direction of the millimeter-wave radar to the target angle; determining a second distance between the sound-emitting object and the terminal device using the beam; determining the transmission power of the millimeter-wave radar transmission signal based on the second distance; and transmitting a transmission signal towards the target angle according to the transmission power, wherein the coverage area of the transmission signal includes the sound-emitting object.
[0257] In some embodiments, the method further includes: acquiring an echo signal, the echo signal being a signal modulated by the transmitted signal via a vibration portion of the sound-producing object; and demodulating the echo signal to recover the first audio signal corresponding to the sound-producing object.
[0258] In some embodiments, demodulating the echo signal to recover the first audio signal corresponding to the sound-producing object specifically includes: demodulating the echo signal and recovering the first audio signal corresponding to the sound-producing object based on the first phase of the transmitted signal and the second phase of the echo signal.
[0259] In some embodiments, the method further includes: performing noise reduction processing and audio super-resolution processing on the first audio signal to obtain a third audio signal corresponding to the millimeter-wave radar.
[0260] In some embodiments, when the sound pickup method involves simultaneously using the microphone and the millimeter-wave radar, the method further includes: performing noise reduction and audio super-resolution processing on the second audio signal to obtain a fourth audio signal corresponding to the microphone; concatenating the third audio signal and the fourth audio signal in the channel dimension to obtain a concatenated dual-channel signal; and fusing the dual-channel signal using a self-attention network to obtain a fused audio result, wherein the fusion includes dynamically assigning different weights to the audio signal corresponding to the millimeter-wave radar and the audio signal corresponding to the microphone in the dual-channel signal.
[0261] According to the sound pickup method provided in the embodiments of this application, when shooting a distant scene, a millimeter-wave radar installed in the terminal device picks up the distant audio, enabling the terminal device to overcome the limitation of the sound pickup distance and provide the user with a video that matches the audio. Specifically, when the user uses the terminal device to shoot a distant scene, they can choose to use a microphone array or a millimeter-wave radar for sound pickup based on the current zoom level or the distance manually set by the user. When millimeter-wave radar is selected for sound pickup, the target object and its angle in the shooting scene can be determined first based on preset characteristics. Then, the millimeter-wave radar emits a beam towards that angle to detect the distance between the target object and the terminal device. Next, the transmission power of the radar wave signal is determined based on the distance, and the signal is emitted in the direction of the target object. Based on the received echo signal modulated by the target object, the audio signal of the target object can be obtained through processing, that is, the audio corresponding to the video scene can be obtained.
[0262] Based on the same technical concept, embodiments of this application also provide a terminal device, including a processor; a memory; the memory stores a computer program, the computer program including instructions, which, when executed by the processor, cause the terminal device to perform one or more steps of any of the above methods.
[0263] Based on the same technical concept, this application embodiment also provides a chip system, the chip system including: a processing circuit, a receiving pin, and a transmitting pin; wherein, the receiving pin, the transmitting pin, and the processing circuit communicate with each other through an internal connection path, and the processing circuit executes one or more steps of any of the above methods to control the receiving pin to receive signals and control the transmitting pin to transmit signals.
[0264] Based on the same technical concept, embodiments of this application also provide a computer-readable storage medium storing computer-executable program instructions, which, when executed on a computer, cause the computer or processor to perform one or more steps of any of the above methods.
[0265] Based on the same technical concept, embodiments of this application also provide a computer program product containing instructions, the computer program product including computer program code, which, when run on a computer, causes the computer or processor to perform one or more steps of any of the above methods.
[0266] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0267] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
[0268] The above description is merely a specific implementation of the embodiments of this application, but the protection scope of the embodiments of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the embodiments of this application should be covered within the protection scope of the embodiments of this application. Therefore, the protection scope of the embodiments of this application should be determined by the protection scope of the claims.
Claims
1. A method for sound pickup, characterized in that, Applied to a terminal device, the terminal device including a microphone and millimeter-wave radar, the method includes: Get the zoom level of the current recording, or the pickup distance input by the user; The sound pickup method is obtained based on the zoom level or the pickup distance. The sound pickup method includes using the millimeter-wave radar alone, using the microphone alone, and using both the microphone and the millimeter-wave radar simultaneously. The first audio signal picked up by the millimeter-wave radar and / or the second audio signal picked up by the microphone are processed according to the aforementioned sound pickup method to obtain the processed audio result.
2. The method according to claim 1, characterized in that, The terminal device also includes a camera, and the distance between the millimeter-wave radar and the camera is less than or equal to a first distance.
3. The method according to claim 2, characterized in that, The terminal device is a mobile phone, the camera is the rear camera of the mobile phone, and the millimeter-wave radar is located in the camera area corresponding to the rear camera.
4. The method according to any one of claims 1-3, characterized in that, The sound pickup method is obtained based on the zoom level, specifically including: When the zoom factor is less than the first factor, the sound pickup method is to use a microphone alone for sound pickup; When the zoom factor is greater than or equal to the first factor and less than the second factor, the sound pickup method is to simultaneously use the microphone and the millimeter-wave radar to pick up sound; When the zoom factor is greater than or equal to the second factor, the sound pickup method is to use millimeter-wave radar for sound pickup alone; wherein, the value corresponding to the second factor is greater than the value corresponding to the first factor.
5. The method according to any one of claims 1-3, characterized in that, The pickup method is obtained based on the pickup distance, specifically including: When the pickup distance is less than the first threshold, the pickup method is to use a microphone alone for pickup. When the pickup distance is greater than or equal to the first threshold and less than the second threshold, the pickup method is to simultaneously use the microphone and the millimeter-wave radar to pick up sound. When the pickup distance is greater than or equal to the second threshold, the pickup method is to use millimeter-wave radar for pickup alone; wherein, the second threshold is greater than the first threshold.
6. The method according to any one of claims 1-5, characterized in that, The method further includes: The system receives a first operation from the user, which is used to select a sound pickup mode, including an automatic sound pickup mode and a manual sound pickup mode. When the sound pickup mode is automatic sound pickup mode, the zoom magnification is obtained, and the sound pickup method is obtained according to the zoom magnification; When the sound pickup mode is manual sound pickup mode, a second operation input by the user is received, the second operation being used to select the sound pickup distance; The pickup method is obtained based on the pickup distance.
7. The method according to any one of claims 1-6, characterized in that, When the sound pickup method involves using the millimeter-wave radar alone or using both the microphone and the millimeter-wave radar simultaneously, the method further includes: Get the currently recorded video frame; Based on the preset first feature information of the sound-emitting object and the video frame, the target angle of the sound-emitting object relative to the terminal device is obtained; Set the beam pointing of the millimeter-wave radar to the target angle; The second distance between the sound-emitting object and the terminal device is determined by the beam; The transmission power of the millimeter-wave radar signal is determined based on the second distance; According to the transmission power, a transmission signal is sent to the target angle, and the coverage area of the transmission signal includes the sound-emitting object.
8. The method according to claim 7, characterized in that, The method further includes: Acquire the echo signal, wherein the echo signal is the signal after the transmitted signal is modulated by the vibrating part of the sound-producing object; Demodulate the echo signal to recover the first audio signal corresponding to the sound source.
9. The method according to claim 8, characterized in that, Demodulating the echo signal to recover the first audio signal corresponding to the sound source specifically includes: The echo signal is demodulated, and the first audio signal corresponding to the sound-emitting object is obtained based on the first phase of the transmitted signal and the second phase of the echo signal.
10. The method according to claim 9, characterized in that, The method further includes: The first audio signal is subjected to noise reduction and audio super-resolution processing to obtain the third audio signal corresponding to the millimeter-wave radar.
11. The method according to claim 9, characterized in that, When the sound pickup method involves simultaneously using the microphone and the millimeter-wave radar for sound pickup, the method further includes: The second audio signal is subjected to noise reduction and audio super-resolution processing to obtain the fourth audio signal corresponding to the microphone; The third audio signal and the fourth audio signal are spliced together in the channel dimension to obtain the spliced dual-channel signal; The dual signals are fused using a self-attention network to obtain the fused audio result. The fusion includes dynamically assigning different weights to the audio signal corresponding to the millimeter-wave radar and the audio signal corresponding to the microphone in the dual signals.
12. A terminal device, characterized in that, include: processor; Memory; The memory stores a computer program, the computer program including instructions that, when executed by the processor, cause the terminal device to perform the method as described in any one of claims 1 to 11.
13. A chip system, characterized in that, The chip system includes a processing circuit, a receiving pin, and a transmitting pin; wherein the receiving pin, the transmitting pin, and the processing circuit communicate with each other through an internal connection path, and the processing circuit executes the method as described in any one of claims 1 to 11 to control the receiving pin to receive signals and control the transmitting pin to transmit signals.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable program instructions that, when executed on a computer, cause the computer to perform the method as described in any one of claims 1 to 11.