Display devices and methods for acquiring voice signals

By determining the coordinates of the sound source using an infrared imaging module and a radar module, and acquiring audio data by combining it with a microphone array, a high-precision speech signal is synthesized by adjusting the timing. This solves the problem of low speech signal accuracy caused by the low sampling rate of the display device, and improves the accuracy of speech recognition.

CN118972653BActive Publication Date: 2025-10-28HISENSE VISUAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410946360.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-15
Publication Date
2025-10-28
Estimated Expiration
2044-07-15

AI Technical Summary

Technical Problem

Existing display devices, with their low sampling rates, struggle to generate high-precision speech signals, resulting in low speech recognition accuracy.

Method used

By combining an infrared imaging module and a radar module, the coordinates of the sound source are determined, and audio data is acquired through a microphone array. The start time point of the audio data is adjusted using time difference to synthesize a high-precision speech signal.

Benefits of technology

It improves the accuracy of voice signals in display devices, enhances the accuracy of voice wake-up and recognition, and reduces the computational load on devices due to high sampling frequencies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118972653B_ABST
    Figure CN118972653B_ABST
Patent Text Reader

Abstract

This application provides a display device and a method for acquiring voice signals. The method responds to a user-input voice command by acquiring sound source coordinates, determining a first and second coordinate of the sound source based on an infrared imaging module, and determining a third coordinate of the sound source based on a radar module. It acquires first and second audio data, determines the time difference between the first and second audio data based on the sound source coordinates, and adds a first starting time point to the time difference to obtain a second starting time point. It then modifies the starting time point of the first audio data based on the second starting time point to obtain updated first audio data. Finally, it synthesizes the updated first and second audio data to obtain third audio data. The method uses an infrared imaging module and a radar module to determine the sound source coordinates and time difference, then aligns multiple audio data streams based on the time difference, and synthesizes the aligned multiple audio data streams to obtain a high-precision voice signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of display device technology, and in particular to a display device and a method for acquiring voice signals. Background Technology

[0002] Display devices are intelligent devices capable of presenting a user interface and supporting user interaction. Taking smart TVs as an example, smart TVs are television products based on Internet application technology, possessing an open operating system and chip, and an open application platform. They enable two-way human-computer interaction and integrate multiple functions such as audio-visual entertainment and data to meet diverse and personalized user needs. Display devices can collect user-input voice signals and achieve human-computer interaction and voice wake-up functions through voice recognition.

[0003] To improve the accuracy of speech recognition on display devices, these devices need to acquire high-precision speech signals. Therefore, a high sampling frequency is required to sample the user's input speech signal. This can be achieved by adding a phase shifter to the analog-to-digital converter and then calculating the phase difference used for data synthesis, thereby increasing the sampling rate.

[0004] However, the method of adding a phase shifter to the analog-to-digital converter and then calculating the phase difference for data synthesis suffers from phase error and algorithm error, and the calculation process is unstable, resulting in poor effect on improving the sampling rate and difficulty in generating high-precision speech signals. Summary of the Invention

[0005] This application provides a display device and a method for acquiring voice signals, in order to solve the problem that the voice signals acquired by the display device are of low accuracy due to the low sampling rate.

[0006] In a first aspect, some embodiments of this application provide a display device, including: a display, an infrared imaging module, a radar module, a microphone array, and a controller. The display is configured to display a user interface; the infrared imaging module is configured to acquire temperature information; the radar module is configured to emit electromagnetic waves and receive reflected waves, the reflected waves being electromagnetic waves reflected by a target within the environment of the display device; the microphone array is configured to acquire audio data, the microphone array including at least a first microphone and a second microphone; the controller is configured to:

[0007] In response to a user's voice command, the coordinates of the sound source are obtained. The coordinates of the sound source include a first coordinate, a second coordinate, and a third coordinate. The first and second coordinates are determined based on the temperature information collected by the infrared imaging module. The third coordinate is determined based on the sound source reflection time of the transmitted wave and the reflected wave collected by the radar module.

[0008] Acquire first audio data and second audio data; wherein, the first audio data is the audio data corresponding to the voice command collected by the first microphone; and the second audio data is the audio data corresponding to the voice command collected by the second microphone.

[0009] The time difference between the first audio data and the second audio data is determined based on the sound source coordinates;

[0010] The first starting time point is added to the time difference to obtain the second starting time point; wherein, the first starting time point is the starting time point of the first audio data;

[0011] Modify the start time of the first audio data according to the second start time to obtain updated first audio data;

[0012] The updated first audio data and the second audio data are synthesized to obtain the third audio data.

[0013] Secondly, some embodiments of this application also provide a method for acquiring voice signals, applied to the display device provided in the first aspect, wherein the display device includes: a display, an infrared imaging module, a radar module, a microphone array, and a controller. The method includes:

[0014] In response to a user's voice command, the coordinates of the sound source are obtained. The coordinates of the sound source include a first coordinate, a second coordinate, and a third coordinate. The first and second coordinates are determined based on temperature information collected by the infrared imaging module. The third coordinate is determined based on the sound source reflection time of the transmitted and reflected waves collected by the radar module.

[0015] Acquire first audio data and second audio data; wherein, the first audio data is the audio data corresponding to the voice command collected by the first microphone; and the second audio data is the audio data corresponding to the voice command collected by the second microphone.

[0016] The time difference between the first audio data and the second audio data is determined based on the sound source coordinates;

[0017] The first starting time point is added to the time difference to obtain the second starting time point; wherein, the first starting time point is the starting time point of the first audio data;

[0018] Modify the start time of the first audio data according to the second start time to obtain updated first audio data;

[0019] The updated first audio data and the second audio data are synthesized to obtain the third audio data.

[0020] As can be seen from the above technical solutions, some embodiments of this application provide a display device and a method for acquiring voice signals. The method responds to a user-input voice command by acquiring the coordinates of a sound source, determining a first and second coordinate of the sound source based on an infrared imaging module, and determining a third coordinate of the sound source based on a radar module. It acquires first and second audio data, determines the time difference between the first and second audio data based on the sound source coordinates, and adds a first starting time point to the time difference to obtain a second starting time point. It also modifies the starting time point of the first audio data based on the second starting time point to obtain updated first audio data. Finally, it synthesizes the updated first and second audio data to obtain third audio data. The method uses an infrared imaging module and a radar module to determine the sound source coordinates and the time difference, then aligns multiple audio data streams based on the time difference, and synthesizes the aligned multiple audio data streams to obtain a high-precision voice signal. Attached Figure Description

[0021] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a schematic diagram illustrating an operational scenario between a display device and a control device provided in some embodiments of this application;

[0023] Figure 2 This is a schematic diagram of the hardware configuration of a display device provided in some embodiments of this application;

[0024] Figure 3 This is a schematic diagram of the software configuration of a display device provided in some embodiments of this application;

[0025] Figure 4 This is a schematic diagram of a voice signal acquisition method provided in some embodiments of this application;

[0026] Figure 5 This is a schematic diagram illustrating a method for determining sound source coordinates using an infrared imaging module, provided in some embodiments of this application.

[0027] Figure 6 This is a schematic diagram illustrating the controller connection method of a display device provided in some embodiments of this application;

[0028] Figure 7 This is a schematic diagram of a target person detected by an infrared imaging module provided in some embodiments of this application;

[0029] Figure 8 A schematic diagram illustrating a method for determining the first and second coordinates of a command sound source according to some embodiments of this application;

[0030] Figure 9 This is a schematic diagram illustrating a method for determining sound source coordinates using a radar module, provided in some embodiments of this application.

[0031] Figure 10 This is a schematic diagram illustrating a method for calculating time difference based on sound source coordinates, provided in some embodiments of this application.

[0032] Figure 11 This is a schematic diagram showing the positional relationship between the microphone array and the radar module provided in some embodiments of this application;

[0033] Figure 12 This is a schematic diagram of audio data synthesis provided for some embodiments of this application;

[0034] Figure 13 This is an overall flowchart of a speech signal acquisition method provided in some embodiments of this application. Detailed Implementation

[0035] The embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described below do not represent all embodiments consistent with this application. They are merely examples of systems and methods consistent with some aspects of this application as detailed in the claims.

[0036] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.

[0037] The terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar or related objects or entities, and do not necessarily imply a specific order or sequence, unless otherwise specified. It should be understood that such terms are interchangeable where appropriate.

[0038] The terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device that includes a range of components is not necessarily limited to all of the components that are clearly listed, but may include other components that are not clearly listed or that are inherent to such product or device.

[0039] The term "module" refers to any known or subsequently developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functions associated with that element.

[0040] In this embodiment, the display device 200 generally refers to a device with screen display and data processing capabilities. For example, the display device 200 includes, but is not limited to, smart TVs, mobile terminals, computers, monitors, advertising screens, wearable devices, virtual reality devices, augmented reality devices, etc.

[0041] Figure 1 This is a schematic diagram illustrating an operational scenario between a display device and a control device provided in some embodiments of this application. For example... Figure 1 As shown, a user can operate the display device 200 via touch operation, a mobile terminal 300, and a control device 100. The control device 100 receives user input commands and converts them into control commands that the display device 200 can recognize and respond to. For example, the control device 100 can be a remote control, a stylus, a gamepad, etc.

[0042] The mobile terminal 300 can function as a control device for human-computer interaction between the user and the display device 200. It can also function as a communication device for establishing a communication connection with the display device 200 and exchanging data. In some embodiments, the mobile terminal 300 can have software applications installed on it and communicate with the display device 200 via network communication protocols to achieve one-to-one control and data communication. Furthermore, it can transmit audio and video content displayed on the mobile terminal 300 to the display device 200 for synchronized display.

[0043] In some embodiments, the mobile terminal 300 or other electronic devices may also simulate the functions of the control device 100 by running an application that controls the display device 200.

[0044] like Figure 1 The diagram also shows that the display device 200 communicates with the server 400 via various communication methods. This allows the display device 200 to communicate via a local area network (LAN), a wireless local area network (WLAN), and other networks.

[0045] Display device 200 can provide broadcast television reception function, and can also be equipped with intelligent network television function that provides computer support, including but not limited to network television, smart television, Internet Protocol television (IPTV), etc.

[0046] Figure 2 Provided for some embodiments of this application Figure 1 Hardware configuration block diagram of display device 200.

[0047] In some embodiments, the display device 200 may include at least one of a tuner 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a memory, a power supply, and a user input interface.

[0048] In some embodiments, detector 230 is used to acquire signals from the external environment or to interact with the external environment. For example, detector 230 includes a light receiver, a sensor for acquiring ambient light intensity; or, detector 230 includes an image acquisition device, such as a camera, which can be used to acquire external environmental scenes, user attributes, or user interaction gestures; or, detector 230 includes a sound acquisition device, such as a microphone, for receiving external sounds; or, detector 230 includes an infrared imaging module for acquiring temperature distribution images; or, detector 230 includes millimeter-wave radar for identifying the distance, speed, etc., of a target.

[0049] In some embodiments, the display 260 includes display function components for presenting images and driving components for driving image display. The display 260 is used to receive and display image signals output from the controller 250. For example, the display 260 can be used to display video content, image content, menu control interface components, and user control UI interfaces, etc.

[0050] In some embodiments, the communication device 220 is a component used to communicate with external devices or the server 400 according to various communication protocol types. The display device 200 may have multiple communication devices 220 depending on the supported communication methods. For example, when the display device 200 supports wireless network communication, it may have a communication device 220 with WiFi functionality. When the display device 200 supports Bluetooth connectivity, it needs to have a communication device 220 with Bluetooth functionality.

[0051] The communication device 220 enables the display device 200 to communicate with external devices or the server 400 via wireless or wired connections. Wired connections utilize data cables, interfaces, or other components to connect the display device 200 to external devices. Wireless connections utilize wireless signals or wireless networks. The display device 200 can directly establish a connection with external devices or indirectly through gateways, routers, or other connection devices.

[0052] In some embodiments, the controller 250 may include at least one of a central processing unit, a video processor, an audio processor, a graphics processor, and a power processor, and a first to an nth interface for input / output. The controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in memory. The controller 250 controls the overall operation of the display device 200.

[0053] In some embodiments, the controller 250 and the tuner 210 may be located in different separate devices, that is, the tuner 210 may also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.

[0054] In some embodiments, a user can input user commands through a graphical user interface (GUI) displayed on a display 260, and the user input interface receives user input commands through the graphical user interface (GUI).

[0055] In some embodiments, the audio output device 270 can be a built-in speaker of the display device 200 or an external audio output device connected to the display device 200. For the external audio output device connected to the display device 200, the display device 200 may also be provided with an external audio output terminal, through which the audio output device can be connected to the display device 200 to output sound from the display device 200.

[0056] In some embodiments, the user input interface 280 can be used to receive instructions from user input.

[0057] In some embodiments, to enable user interaction, the display device 200 may run an operating system. The operating system is a computer program used to manage and control the hardware and software resources of the display device 200. The operating system can control the display device 200 to provide a user interface; for example, the operating system can directly control the display device 200 to provide a user interface, or it can provide a user interface by running an application. The operating system also allows users to interact with the display device 200.

[0058] It should be noted that the operating system can be a native operating system based on a specific operating platform, a third-party operating system that is deeply customized based on a specific operating platform, or an independent operating system specifically developed for the display device 200.

[0059] An operating system can be divided into different modules or levels based on the functions it implements, for example... Figure 3As shown, in some embodiments, the system is divided into four layers, from top to bottom: the Applications layer (referred to as the "Application Layer"), the Application Framework layer (referred to as the "Framework Layer"), the System Library layer, and the Kernel layer.

[0060] In some embodiments, the application layer provides services and interfaces for applications, enabling the display device 200 to run applications and interact with the user based on the applications. The application layer may contain at least one application, which may be a built-in Windows program, system settings program, or clock program of the operating system; or it may be an application developed by a third-party developer. In specific implementations, the application packages in the application layer are not limited to the examples above.

[0061] The framework layer provides application programming interfaces (APIs) and a programming framework for applications. The application framework layer includes predefined functions. It acts as a central processing unit, determining the actions taken by applications within the application layer. Through the API, applications can access system resources and obtain system services during execution.

[0062] like Figure 3 As shown, the application framework layer in this embodiment includes a view system, managers, and content providers. The view system designs and implements the application's interface and interactions, and includes lists, grids, text boxes, and buttons. The managers include at least one of the following modules: an activity manager for interacting with all running activities in the system; a location manager for providing system services or applications with access to system location services; a package manager for retrieving various information related to application packages currently installed on the device; a notification manager for controlling the display and clearing of notification messages; and a window manager for managing icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.

[0063] In some embodiments, the Activity Manager manages the lifecycle of individual applications and common navigation and back functions, such as controlling application exit, opening, and back actions. The Window Manager manages all window programs, such as obtaining the screen size, determining if a status bar is present, locking the screen, capturing the screen, and controlling changes to the display window, such as shrinking the display window, shaking the display, or distorting the display.

[0064] In some embodiments, the system runtime library layer can provide support for the framework layer. When the framework layer is used, the operating system runs the instruction library contained in the system runtime library layer, such as the C / C++ instruction library, to implement the functions to be performed by the framework layer.

[0065] In some embodiments, the kernel layer is a functional layer situated between the hardware and software of the display device 200. The kernel layer can implement functions such as hardware abstraction, multitasking, and memory management. For example, ... Figure 3 As shown, hardware drivers can be configured in the kernel layer. The drivers included in the kernel layer can be at least one of the following: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver, etc.

[0066] It should be noted that the above examples are merely a simple division of operating system functions and do not limit the specific form of the operating system of the display device 200 in this application embodiment. Depending on the function of the display device, the type of operating system, and other factors, the number of levels and the specific level type of the operating system may be expressed in other forms.

[0067] In some embodiments, the display device 200 includes an infrared imaging module. The infrared imaging module is communicatively connected to the controller 250 and detects the temperature distribution of the environment surrounding the display device 200 in real time through non-contact measurement. The core components of the infrared imaging module include an infrared lens, an infrared detector, and a signal processor. The infrared lens collects infrared radiation emitted by objects and focuses it onto the infrared detector; the infrared detector receives the infrared radiation focused by the infrared lens and converts it into an electrical signal; the infrared detector includes infrared complementary metal-oxide-semiconductor (CMOS), infrared charge-coupled device (CCD), etc.; the signal processor amplifies, adjusts, filters, and performs analog-to-digital (AD) conversion on the electrical signal output from the infrared detector to generate thermal image information, which is then transmitted to the controller 250 of the display device 200.

[0068] In some embodiments, the display device 200 includes a radar module, which is communicatively connected to the controller 250. The radar module processes and analyzes received signals to detect the location of a target object. The radar module includes a transmitter, a receiver, an antenna, and a signal processing system. The transmitter, consisting of a frequency-stabilized oscillator, a microwave power amplifier, and a drive circuit, is responsible for transmitting high-frequency or ultra-high-frequency signals. The receiver, consisting of a front-end amplifier, an intermediate frequency amplifier, a detector, and an echo suppressor, is responsible for receiving echo signals and amplifying, demodulating, and restoring the received echo signals to the original information signal. The antenna concentrates the high-frequency current output from the transmitter and radiates it outwards in a certain direction, and receives echo signals, inputting the echo signals to the receiver for processing. The signal processing system is responsible for receiving, processing, and analyzing the radar echo signals.

[0069] In some embodiments, the display device 200 includes a microphone array comprising at least two microphones. The microphone array is communicatively connected to the controller 250 for acquiring audio data input by the user to the display device 200. Because the microphone array includes multiple microphones, it can simultaneously capture audio data from multiple sound sources, enabling it to efficiently capture and process audio data in complex environments such as meetings and presentations, avoiding sound overlap or noise interference. Furthermore, because the microphone array receives audio data over a wide frequency range, it achieves excellent audio capture performance.

[0070] Based on the aforementioned display device 200, the display device 200 can respond to the user's input voice command, acquire the user's input audio data, and process the audio data to obtain a high-precision voice signal, thereby improving the accuracy of the display device 200's voice wake-up and voice recognition functions.

[0071] In some embodiments, the display device 200 acquires user-input voice signals with frequencies ranging from 100 Hz to 8000 Hz. Higher sampling frequencies result in higher accuracy of the acquired voice signal. To reduce the computational load on the display device 200 while acquiring high-precision voice signals, a method can be used to synthesize multiple low-sampling-frequency audio data into a high-precision voice signal. For example, two low-frequency audio data streams can be acquired and synthesized to obtain a high-precision voice signal.

[0072] In some embodiments, the display device 200 acquires multiple audio data streams via a microphone array consisting of multiple microphones. Each microphone in the microphone array acquires one audio data stream. Because the multiple microphones in the microphone array are located at different positions on the display device 200, the sound from the same sound source arrives at each microphone at different times. This results in a time difference between the multiple audio data streams acquired by the display device 200 through the microphone array, causing errors in the subsequent synthesis of the multiple audio data streams. Therefore, it is necessary to detect the location of the sound source, determine the specific time difference between each audio data stream, and adjust the starting time point of the audio data used for synthesis based on the time difference to eliminate synthesis errors.

[0073] In some embodiments, the display device 200 can detect the location of the sound source of the input voice signal via a radar module. The radar module can be a millimeter-wave radar module. Millimeter-wave radar has characteristics such as short wavelength and strong robustness, resulting in high imaging resolution and resistance to clutter such as dust and smoke. It also boasts advantages such as high resolution, strong penetration, high accuracy, and strong anti-interference capability. Furthermore, because millimeter-wave radar employs Frequency Modulated Continuous Wave (FMCW) technology, it can simultaneously detect the distance and velocity of multiple targets, achieving multi-target detection.

[0074] The method by which the display device 200 detects the location of a sound source using millimeter-wave radar includes: controlling the millimeter-wave radar to generate a radar signal waveform, such as FMCW, through a signal generator, and modulating the signal to the millimeter-wave frequency band (e.g., 76GHz-77GHz) through multi-stage frequency conversion modulation processing; radiating the millimeter-wave frequency band signal into the ambient space where the display device 200 is located through an antenna; receiving the reflected signal reflected back by the target (sound source) through the antenna; mixing the reflected signal with the transmitted signal to obtain a complex baseband signal; performing analog-to-digital (AD) conversion on the complex baseband signal to obtain discrete complex baseband signal data; and performing a Fast Fourier Transform (FFT) on the discrete complex baseband signal data. The range spectrum is formed by performing a range transform (FFT) operation to obtain the target's range information. An FFT transformation is then performed on the range spectrum to obtain the target's velocity information. Incoherent accumulation processing is applied to the data from each channel of the millimeter-wave radar to improve the signal-to-noise ratio and detection accuracy. Constant false alarm rate (CFAR) detection is then performed on the incoherently accumulated data to distinguish between target signals and interference signals. Algorithms such as Digital Beamforming (DBF) or Multiple Signal Classification (MUSIC) are used to estimate the angle of the target signal to obtain its azimuth information. Finally, the range, velocity, and angle information are comprehensively processed, and three-dimensional spatial positioning algorithms (such as triangulation and least squares methods) are used to calculate the specific location of the sound source in three-dimensional space.

[0075] The method of detecting the sound source location by millimeter-wave radar module in the above embodiments has the advantage of high accuracy, but the signal processing is complex and the detection speed is slow, making it impossible for the display device 200 to acquire high-precision voice signals under low load.

[0076] In some embodiments, the display device 200 can also acquire high-precision voice signals according to the following voice signal acquisition method to solve the problem of low accuracy of the acquired voice signals due to the low sampling rate of the display device 200. To satisfy the implementation of the voice signal acquisition method, the display device 200 should at least include a display 260, an infrared imaging module, a radar module, a microphone array, and a controller 250. The display 260 is configured to display a user interface; the infrared imaging module is configured to acquire temperature information; the radar module is configured to emit electromagnetic waves and receive reflected waves, the reflected waves being electromagnetic waves reflected by targets within the environment where the display device is located; the microphone array is configured to acquire audio data, and the microphone array includes at least a first microphone and a second microphone; as... Figure 4 As shown, the controller 250 is configured to execute the program steps corresponding to the voice signal acquisition method, including the following:

[0077] S100: Responds to user-inputted voice commands and obtains the coordinates of the sound source.

[0078] In some embodiments, the display device 200 responds to a user-inputted voice command, acquires corresponding audio data, and determines the location of the sound source issuing the voice command based on the corresponding audio data. The sound source location is determined by the coordinates of the sound source in three-dimensional space, i.e., the sound source coordinates, which include a first coordinate, a second coordinate, and a third coordinate. For example, the sound source coordinates of the first sound source are (x1, y1, z1); and the sound source coordinates of the second sound source are (x2, y2, z2).

[0079] In some embodiments, the first and second coordinates of the sound source are coordinates determined based on temperature information acquired by the infrared imaging module, such as... Figure 5 As shown, the first and second coordinates are determined as follows:

[0080] S501: Send a first acquisition command to the infrared imaging module so that the infrared imaging module can acquire temperature information in response to the first acquisition command.

[0081] In some embodiments, such as Figure 6 As shown, the controller 250 of the display device 200 is communicatively connected to the infrared imaging module. When the display device 200 detects a user's input voice signal, in order to determine the location of the sound source of the input voice signal, the controller 250 sends a first acquisition command to the infrared imaging module to collect temperature information of the environment where the display device 200 is located, so as to control the infrared imaging module to respond to the first acquisition command and collect temperature information. For example, when a first user issues a first voice command in front of the display device 200, the display device 200 detects the first voice command, and the controller 250 controls the infrared imaging module to collect the temperature information of the environment where the display device 200 is located as the first temperature information. The first temperature information includes the temperature information of the first user who issued the first voice command.

[0082] In some embodiments, to improve the positioning accuracy of the command sound source, the infrared imaging module can communicate with the intelligent voice module to quickly locate the command sound source that emitted the voice signal. For example, when the intelligent voice module detects a voice command, it first preliminarily determines the position and direction of the sound source, and then sends the preliminary positioning information (such as the direction angle of the sound source) to the infrared imaging module. The infrared imaging module adjusts the imaging area and focal length according to the preliminary positioning information, prioritizing the area in the preliminary positioning direction.

[0083] S502: Generate a temperature distribution image based on temperature information.

[0084] In some embodiments, targets with temperatures above absolute zero (-273.15°C) emit infrared radiation; the higher the target's temperature, the stronger the emitted infrared radiation. Therefore, the infrared imaging module can calculate the target's temperature by detecting the infrared radiation intensity of all targets in the environment where the display device 200 is located, all of which have temperatures above absolute zero, and generate a temperature distribution image based on the target's temperature. For example, the infrared detector in the infrared imaging module receives the infrared radiation emitted by targets in the environment where the display device 200 is located and converts it into an electrical signal; the signal processor processes the electrical signal, converting it into a digital signal and forming a pixel image to generate a temperature distribution image reflecting the target's temperature distribution. The higher the target's temperature, the higher the brightness of the corresponding pixel color; the lower the target's temperature, the lower the brightness of the corresponding pixel color.

[0085] S503: Determine the location of the command sound source based on the temperature distribution image.

[0086] In some embodiments, after the infrared imaging module generates a temperature distribution image, it can obtain the distribution range of the target temperature in the temperature distribution image. The target temperature is the temperature of the person issuing the voice command, and the location of the command sound source is determined based on the distribution range of the target temperature. For example, if the temperature of the person issuing the voice command is higher than the temperature of targets such as walls and the ground, when the infrared imaging module detects the presence of a person in the environment where the display device 200 is located, the pixel color brightness of the person on the temperature distribution image is higher than that of other targets. This is used to determine the location of the person, i.e., the location of the sound source issuing the voice command. Figure 7 As shown, the pixel color brightness of the human target is higher than that of the surrounding environment, thus the location of the human target and its outline information can be determined.

[0087] More specifically, to improve the accuracy of sound source location, the position of a person on the temperature distribution image can be determined based on the precise value of human body temperature. For example, the range of human body surface temperature is 36℃-37.5℃. After the infrared imaging module generates a temperature distribution image, the pixels representing 36℃-37.5℃ are used as pixels representing the person to determine the sound source location and its contour information.

[0088] In some embodiments, after the display device 200 obtains the distribution range of the target temperature, it is also necessary to generate the specific coordinates of the command sound source based on the distribution range of the target temperature. To calculate the specific coordinates of the command sound source, it is first necessary to obtain the geometric data of the infrared imaging module, including at least one of the field of view, focal length, and resolution, and determine the position of the command sound source based on the distribution range and the geometric data.

[0089] It should be noted that the target of the voice command is a person, but the pixel range of a person is large, making it impossible to accurately locate the coordinates of the sound source. Since a person issues voice commands through their lips, it is necessary to determine the accurate coordinates of the person's lips—that is, the accurate coordinates of the command sound source—based on the target person's outline information and the geometric data from the infrared imaging module. Figure 8 As shown, the methods for determining the accurate coordinates of the command sound source include:

[0090] The position of the pixel representing the command sound source in the temperature distribution image is determined (S801). The infrared imaging module is geometrically calibrated, and the intrinsic parameters (such as focal length, principal point coordinates, etc.) and extrinsic parameters (such as position, attitude, etc.) of the camera are obtained through image processing algorithms (S802). The field of view of the infrared imaging module is calculated based on the camera's calibration parameters and imaging principle (S803). Using the pixel position of the target's lips in the temperature distribution image, the camera's intrinsic and extrinsic parameters, and the field of view, the abscissa and ordinate of the command sound source are obtained through coordinate transformation and geometric calculation (S804).

[0091] The horizontal coordinate of the command sound source calculated using the above method is the first coordinate determined by the infrared imaging module, and the vertical coordinate is the second coordinate determined by the infrared imaging module.

[0092] In some embodiments, to further determine the precise location of the command sound source, it is also necessary to obtain coordinates representing the distance between the command sound source and the display device 200, i.e., the third coordinates of the command sound source. The third coordinates are determined based on the sound source reflection times of the transmitted and reflected waves collected by the radar module. For example... Figure 9 As shown, the third coordinate is determined as follows:

[0093] S901: Receives reflected waves.

[0094] In some embodiments, such as Figure 6 As shown, the controller 250 of the display device 200 is communicatively connected to the millimeter-wave radar module, and can control the millimeter-wave radar module to transmit millimeter waves and receive reflected waves from targets. During the operation of the display device 200, the controller 250 controls the millimeter-wave radar module to transmit millimeter waves in all directions of the environment in which the display device 200 is located, and receives reflected waves from targets in all directions. For example, the display device 200 controls the millimeter-wave radar to transmit millimeter waves to human targets, wall targets, and ground targets in the current environment, and receives the first reflected wave reflected by the human target, the second reflected wave reflected by the wall target, and the third reflected wave reflected by the ground target.

[0095] S902: Extract waveform features from reflected waves.

[0096] In some embodiments, after receiving a reflected wave reflected by multiple targets, the display device 200 extracts waveform features from the reflected wave to obtain information such as the phase, frequency, and amplitude of the reflected wave, which is used to calculate the coordinates of the targets that formed the reflected wave.

[0097] S903: Calculate the coordinates of the first target and the second target based on the phase difference between the transmitted wave and the reflected wave.

[0098] In some embodiments, after the display device 200 acquires the phase of the reflected wave, it calculates the phase difference between the reflected wave and the transmitted wave reflected by different targets, based on the phase of the transmitted wave. After obtaining the phase difference, the abscissa and ordinate of the target forming each reflected wave are calculated based on the phase difference. Specifically, the distance between the target and the millimeter-wave radar is calculated using the phase difference; the propagation direction of the reflected wave is inferred by using a direction estimation algorithm by analyzing the phase difference of the reflected waves received by different antenna elements in the array; and the abscissa and ordinate of the target forming the reflected wave are calculated by combining the coordinate system, distance, and direction of the millimeter-wave radar. The abscissa is the coordinate of the first target, and the ordinate is the coordinate of the second target.

[0099] S904: If the coordinates of the first target are the same as the first coordinate, and the coordinates of the second target are the same as the second coordinate, then obtain the transmission time of the transmitted wave and the reception time of the reflected wave.

[0100] In some embodiments, the first target coordinates calculated from the reflected wave reflected by the command sound source are the same as the first coordinate, and the second target coordinates are the same as the second coordinate. Therefore, the reflected wave that meets the condition that the first target coordinates are the same as the first coordinate and the second target coordinates are the same as the second target is marked as the reflected wave of the command target. After marking the reflected wave of the command target, the reception time of the reflected wave of the command target and the emission time of the corresponding transmitted wave are obtained for subsequent reflection time calculation. For example, if the abscissa of the first reflected wave reflected by the human target is the same as the first coordinate and the ordinate is the same as the second coordinate, then the human target is the command sound source. Therefore, the reception time of the first reflected wave is obtained as 6.67 × 10⁻⁶. -9 The corresponding transmission time of the transmitted wave is 0 seconds (S).

[0101] S905: Calculate the sound source reflection time based on the transmission and reception times.

[0102] In some embodiments, after acquiring the reception time of the reflected wave reflected by the command sound source and the emission time of the corresponding emitted wave, the display device 200 subtracts the emission time from the reception time to obtain the sound source reflection time. For example, if the target person is the command sound source, the reception time of the first reflected wave reflected by the target person is 6.67 × 10⁻⁶. -9If the emission time of the corresponding emitted wave is 0s, then the sound source reflection time is 6.67 × 10⁻⁶. -9 S-0S=6.67×10 -9 S.

[0103] S906: Multiply the sound source reflection time by the propagation speed of the reflected wave to obtain the reflection distance.

[0104] In some embodiments, after obtaining the sound source reflection time, the reflection distance is obtained by multiplying the sound source reflection time by the propagation speed of the reflected wave. It should be understood that the propagation speed of millimeter waves emitted by millimeter-wave radar is approximately equal to the speed of light, i.e., 3 × 10⁻⁶. 8 m / s, therefore, if the sound source reflection time is 6.67 × 10 m / s, -9 S, the reflection distance is 6.67 × 10 -9 S×3×10 8 m / S, approximately equal to 2 meters (m).

[0105] S907: Generate the third coordinate based on the reflection distance value.

[0106] In some embodiments, after acquiring the reflection distance, the display device 200 generates a third coordinate based on the value of the reflection distance. It should be noted that the reflection distance calculated in the above steps is the sum of the distance from the millimeter-wave radar transmitted to the target and the distance from which the reflected wave is reflected back to the millimeter-wave radar. Therefore, the value of the generated third coordinate should be half of the reflection distance. For example, if the reflection distance is 2m, then the value of the generated third coordinate is 1m.

[0107] The third coordinate of the command sound source can be determined by the above method, and the first, second and third coordinates together constitute the sound source coordinates of the command sound source, that is, the precise location of the command sound source.

[0108] S200: Acquire the first audio data and the second audio data.

[0109] In some embodiments, such as Figure 6 As shown, the display device 200 includes a microphone array, which is communicatively connected to the controller 250. The microphone array can detect voice signals in the environment where the display device 200 is located in real time and transmit them to the display device 200. The microphone array includes at least a first microphone and a second microphone. The display device 200 acquires the audio data corresponding to the voice command through the first microphone as first audio data, and acquires the audio data corresponding to the voice command through the second microphone as second audio data. Therefore, the display device 200 can acquire dual-channel audio data for subsequent audio synthesis to obtain a high-precision voice signal.

[0110] It should be understood that the microphone array of the display device 200 may also include a third microphone and a fourth microphone, and acquire third audio data through the third microphone and fourth audio data through the fourth microphone. That is, the display device 200 can acquire four or more audio data. In this embodiment of the application, the number of microphones in the microphone array of the display device 200 is not specifically limited, as long as the multiple audio data acquired by the microphone array can be synthesized into a high-precision voice signal.

[0111] S300: Determine the time difference between the first audio data and the second audio data based on the coordinates of the sound source.

[0112] In some embodiments, after the display device 200 obtains the coordinates of the sound source sending the voice command and multiple audio data (first audio data and second audio data), it can determine the time difference between the first audio data and the second audio data using the sound source coordinates. For example... Figure 10 As shown, methods for determining time differences include:

[0113] S1001: Calculate the first distance between the command sound source and the first microphone based on the sound source coordinates, and calculate the second distance between the command sound source and the second microphone based on the sound source coordinates.

[0114] In some embodiments, after the display device 200 obtains the coordinates of the sound source, it can calculate the first distance between the command sound source and the first microphone and the second distance between the command sound source and the second microphone based on the positional relationship between the command sound source, the radar module, the first microphone and the second microphone, and then calculate the time when the voice command issued by the command sound source reaches the first microphone and the second microphone based on the first distance and the second distance respectively.

[0115] In some embodiments, the first distance and the second distance can be calculated by the following method: obtaining the first straight-line distance between the first microphone and the radar module; performing geometric operations on the sound source coordinates and the first straight-line distance to obtain the first distance; obtaining the second straight-line distance between the second microphone and the radar module; and performing geometric operations on the sound source coordinates and the second straight-line distance to obtain the second distance.

[0116] like Figure 11 As shown, the display device 200 is equipped with a first microphone 1101, a second microphone 1102, and a millimeter-wave radar module 1103. The distance between the millimeter-wave radar module 1103 and the first microphone 1101 is 30 millimeters (mm), and the distance between the first microphone 1101 and the second microphone 1102 is 70 millimeters. The first microphone 1101, the second microphone 1102, and the millimeter-wave radar module 1103 are all arranged parallel to each other on the same horizontal line of the display device 200. Figure 11Taking the display device 200 as an example, assuming the voice command source is located 1m (1000mm) directly in front of the millimeter-wave radar module 1201, the coordinates of the voice source are (0,0,1000). First, the first straight-line distance between the first microphone 1101 and the millimeter-wave radar module 1103 is obtained, which is 30mm. Then, geometric calculations are performed on the coordinates of the voice source and the first straight-line distance. Given that the first straight-line distance between the first microphone 1101 and the millimeter-wave radar module 1103 is 30mm, and based on the voice source coordinates (0,0,1000), the distance between the voice command source and the millimeter-wave radar module 1103 is 1000mm. Since the first microphone 1101, the millimeter-wave radar module 1103, and the command sound source form a triangular structure, and given the lengths of the two sides of the triangular structure and the positional relationship between the first microphone 1101, the millimeter-wave radar module 1103, and the command sound source (in this embodiment, since the command sound source is located directly in front of the millimeter-wave radar module 1103, the three form a right-angled triangle structure), the first distance between the first microphone 1101 and the command sound source can be calculated to be 1000.45mm according to the Pythagorean theorem.

[0117] Similarly, when calculating the second distance, the second straight-line distance between the second microphone 1102 and the millimeter-wave radar module 1103 is first obtained, which is 30mm + 70mm = 100mm. Then, geometric operations are performed on the sound source coordinates (0,0,1000) and the second straight-line distance of 100mm to obtain the second distance between the second microphone 1102 and the command sound source as 1004.99mm.

[0118] It should be understood that even if the triangular structure formed by the first microphone 1101 or the second microphone 1102, the millimeter-wave radar module 1103, and the command sound source is not a right triangle, the distance between the first microphone 1101 or the second microphone 1102 and the command sound source can still be calculated based on the angular relationship between the three. It is not only possible to calculate the distance between the first microphone 1101 or the second microphone 1102 and the command sound source based on the triangular structure. This application does not impose any specific limitation on this.

[0119] S1002: Divide the first distance by the speed of sound propagation to obtain the first time when the first microphone collects the first audio data.

[0120] In some embodiments, after the display device 200 obtains the first distance between the first microphone 1101 and the command sound source, it can calculate the first time for the first microphone 1101 to collect the first audio data according to the formula t = d / v for speed, time, and distance. For example, if the first distance is 1000.45mm and the speed of sound is 340m / s, then the first time is 0.0029425s.

[0121] S1003: Divide the second distance by the speed of sound propagation to obtain the second time when the second microphone collects the second audio data.

[0122] In some embodiments, after the display device 200 obtains the second distance between the second microphone 1102 and the command sound source, it can calculate the second time for the second microphone 1102 to acquire the second audio data using the same method. For example, if the second distance is 1004.99 mm and the speed of sound is 340 m / s, then the second time is 0.0029559 s.

[0123] S1004: Calculate the time difference based on the second time and the first time.

[0124] In some embodiments, the time difference between the first audio data and the second audio data acquired by the display device 200 can be obtained by subtracting the second time from the first time. For example, if the first time is 0.0029425S and the second time is 0.0029559S, then the time difference is 0.0000134S.

[0125] S400: Add the first starting time point to the time difference to obtain the second starting time point.

[0126] In some embodiments, after the display device 200 acquires the first audio data and the second audio data, it needs to align the first audio data and the second audio data for subsequent synthesis operations. Specifically, the starting time point (i.e., the first starting time point) of the first audio data is added to the time difference to obtain a second time point that is the same as the starting time point of the second audio data.

[0127] S500: Modify the start time point of the first audio data according to the second start time point to obtain updated first audio data.

[0128] In some embodiments, the first audio data is truncated, and the second starting time point is used as the starting time point of the first audio data to obtain updated first audio data. At this time, the updated first audio data and the second audio data are aligned and can be synthesized.

[0129] S600: Synthesize and update the first and second audio data to obtain the third audio data.

[0130] In some embodiments, the display device 200 aligns audio data from different channels before performing a synthesis operation to obtain a high-precision speech signal. First, it acquires the first spectrum of the updated first audio data and the second spectrum of the second audio data; second, it adds the first and second spectra to obtain a third spectrum; finally, it acquires the audio data corresponding to the third spectrum to obtain the third audio data. The spectrum diagram of the synthesized third audio data from the updated first and second audio data is shown below. Figure 12 As shown, the spectrum of the third audio data is the sum of the spectrum of the first audio data and the spectrum of the second audio data, therefore, the third audio data has higher accuracy.

[0131] In some embodiments, after the display device 200 synthesizes the updated first audio data and the second audio data into third audio data, the authenticity of the third audio data can be verified. Since audio data has higher reproduction accuracy at low sampling rates, low-frequency data is used for authenticity verification. Specifically, verification data is extracted from the updated first audio data or the second audio data; the verification data is data with a frequency lower than a frequency threshold in the updated first audio data or the second audio data. Data to be verified is extracted from the third audio data; the data to be verified is data with a frequency lower than a frequency threshold in the third audio data, and the frequency of the data to be verified is the same as that of the verification data. In this embodiment, the frequency threshold is 150Hz. If the data to be verified is the same as the verification data, it indicates that the synthesized high-precision audio data has high authenticity. A wake-up command can be generated based on the third audio data, and in response to the wake-up command, the display 260 can be controlled to display the user interface.

[0132] If the data to be verified differs from the verification data, it indicates that the updated first audio data and the second audio data are not fully aligned. The alignment operation needs to be performed again on the updated first and second audio data, specifically as follows: Determine the second time difference between the updated first and third audio data based on the sound source coordinates; add the second starting time point to the second time difference to obtain the third starting time point; modify the starting time point of the updated first audio data based on the third starting time point to obtain the second updated audio data; synthesize the second updated audio data and the second audio data to obtain the updated third audio data; finally, verify the authenticity of the updated third audio data again. If the verification is successful, it means that the updated third audio data can be used for subsequent voice wake-up and other operations on the display device; if the verification fails, the alignment operation is performed again using the above method.

[0133] In some embodiments, when the environment in which the display device 200 is located includes interference audio from other sound sources in addition to the voice commands issued by the command sound source, the display device 200 will acquire multiple synthesized voice signals, including interference audio, according to the above-described voice signal acquisition method, affecting the accuracy of the acquired voice signal. Therefore, the display device 200 can also use the above-described method for verifying the authenticity of the third audio data to remove the synthesized interference audio. For example, by verifying the low-frequency data of the first synthesized audio with the low-frequency data of the original audio data, and the verification result shows that the error rate between the two is greater than the error threshold (e.g., 50%), the first synthesized audio is determined to be interference audio, reducing the impact of interference audio on the accuracy of the high-precision voice signal synthesized by the display device 200.

[0134] like Figure 13 As shown, the display device 200 controls the infrared imaging module to acquire temperature information and determines the human body contour based on the temperature information. Based on the human body contour and the geometric features of the infrared imaging module, the location of the lips is determined as the direction of the sound source. Simultaneously, the display device 200 also controls the radar module to acquire distance information and combines the distance information with the sound source direction to determine accurate sound source coordinates. Then, multiple audio data streams are acquired through a microphone array. The time difference between the multiple audio data streams is calculated using the sound source coordinates and the positional relationship between the microphone array and the radar module. Alignment is then performed on the multiple audio data streams based on the time difference. The aligned multiple audio data streams are then synthesized to obtain synthesized audio data. Finally, the authenticity of the synthesized audio data is verified. Upon successful verification, a high-precision voice signal is obtained.

[0135] Based on the aforementioned display device 200, some embodiments of this application also provide a method for acquiring voice signals, including:

[0136] In response to a user's voice command, the coordinates of the sound source are obtained. The coordinates of the sound source include a first coordinate, a second coordinate, and a third coordinate. The first and second coordinates are determined based on temperature information collected by the infrared imaging module. The third coordinate is determined based on the sound source reflection time of the transmitted and reflected waves collected by the radar module.

[0137] Acquire first audio data and second audio data; wherein, the first audio data is the audio data corresponding to the voice command collected by the first microphone; and the second audio data is the audio data corresponding to the voice command collected by the second microphone.

[0138] The time difference between the first audio data and the second audio data is determined based on the coordinates of the sound source.

[0139] The first starting time point is added to the time difference to obtain the second starting time point; wherein, the first starting time point is the starting time point of the first audio data;

[0140] Modify the start time of the first audio data according to the second start time point to obtain updated first audio data;

[0141] The updated first audio data and the second audio data are synthesized to obtain the third audio data.

[0142] As can be seen from the above technical solutions, some embodiments of this application provide a display device and a method for acquiring voice signals. The method responds to a user-input voice command by acquiring the coordinates of a sound source, determining a first and second coordinate of the sound source based on an infrared imaging module, and determining a third coordinate of the sound source based on a radar module. It acquires first and second audio data, determines the time difference between the first and second audio data based on the sound source coordinates, and adds a first starting time point to the time difference to obtain a second starting time point. It also modifies the starting time point of the first audio data based on the second starting time point to obtain updated first audio data. Finally, it synthesizes the updated first and second audio data to obtain third audio data. The method uses an infrared imaging module and a radar module to determine the sound source coordinates and the time difference, then aligns multiple audio data streams based on the time difference, and synthesizes the aligned multiple audio data streams to obtain a high-precision voice signal.

[0143] The same or similar parts among the various embodiments in this specification can be referred to mutually, and will not be repeated here.

[0144] Those skilled in the art will clearly understand that the techniques in the embodiments of the present invention can be implemented using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions in the embodiments of the present invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or certain parts of the embodiments of the present invention.

[0145] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

[0146] For ease of explanation, the above description has been provided in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Various modifications and variations can be obtained based on the above teachings. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, thereby enabling those skilled in the art to better utilize the described embodiments and various different variations of embodiments suitable for specific use considerations.

Claims

1. A display device, characterized in that, include: The monitor is configured to display the user interface; The infrared imaging module is configured to acquire temperature information; The radar module is configured to transmit electromagnetic waves and receive reflected waves, wherein the reflected waves are electromagnetic waves that are reflected by targets in the environment where the display device is located after the transmitted waves are reflected by the transmitted waves. A microphone array configured to acquire audio data, the microphone array including at least a first microphone and a second microphone; The controller is configured as follows: In response to a user's voice command, the coordinates of the sound source are obtained. The coordinates of the sound source include a first coordinate, a second coordinate, and a third coordinate. The first and second coordinates are determined based on the temperature information collected by the infrared imaging module. The third coordinate is determined based on the sound source reflection time of the transmitted wave and the reflected wave collected by the radar module. Acquire first audio data and second audio data; wherein, the first audio data is the audio data corresponding to the voice command collected by the first microphone; and the second audio data is the audio data corresponding to the voice command collected by the second microphone. The time difference between the first audio data and the second audio data is determined based on the coordinates of the sound source. The first starting time point is added to the time difference to obtain the second starting time point; wherein, the first starting time point is the starting time point of the first audio data; Modify the start time of the first audio data according to the second start time point to obtain updated first audio data; The updated first audio data and the second audio data are synthesized to obtain the third audio data.

2. The display device according to claim 1, characterized in that, The controller is configured to acquire the coordinates of the sound source as follows: Send a first acquisition command to the infrared imaging module so that the infrared imaging module acquires temperature information in response to the first acquisition command; A temperature distribution image is generated based on the temperature information; The location of the command sound source is determined based on the temperature distribution image, wherein the command sound source is the sound source that issues the voice command; the horizontal coordinate of the command sound source is the first coordinate, and the vertical coordinate of the command sound source is the second coordinate.

3. The display device according to claim 2, characterized in that, The controller is configured to determine the location of the command sound source based on the temperature distribution image, specifically as follows: Obtain the distribution range of the target temperature in the temperature distribution image; Obtain the geometric data of the infrared imaging module, wherein the geometric data includes at least one of the field of view, focal length, and resolution; The location of the command sound source is determined based on the distribution range and geometric data.

4. The display device according to claim 2, characterized in that, The controller is configured to acquire the coordinates of the sound source as follows: Receive the reflected wave; Waveform features are extracted from the reflected wave, the waveform features including at least one of phase, frequency and amplitude; The coordinates of the first target and the coordinates of the second target are calculated based on the phase difference between the transmitted wave and the reflected wave; wherein, the coordinates of the first target are the abscissa of the target reflecting the transmitted wave, and the coordinates of the second target are the ordinate of the target reflecting the transmitted wave. If the first target coordinates are the same as the first coordinates, and the second target coordinates are the same as the second coordinates, then the transmission time of the transmitted wave and the reception time of the reflected wave are obtained. The sound source reflection time is calculated based on the transmission time and the reception time. Multiply the sound source reflection time by the propagation speed of the reflected wave to obtain the reflection distance; A third coordinate is generated based on the value of the reflection distance.

5. The display device according to claim 4, characterized in that, The controller is specifically configured to determine the time difference between the first audio data and the second audio data based on the sound source coordinates as follows: Calculate the first distance between the command sound source and the first microphone based on the sound source coordinates, and calculate the second distance between the command sound source and the second microphone based on the sound source coordinates; Divide the first distance by the speed of sound to obtain the first time when the first microphone collects the first audio data; Divide the second distance by the speed of sound to obtain the second time when the second microphone collects the second audio data; The time difference is calculated based on the second time and the first time.

6. The display device according to claim 5, characterized in that, The controller performs the calculation of a first distance between the command sound source and the first microphone based on the sound source coordinates, specifically configured as follows: Obtain the first straight-line distance between the first microphone and the radar module; Perform geometric operations on the sound source coordinates and the first straight-line distance to obtain the first distance; The controller performs the calculation of a second distance between the command sound source and the second microphone based on the sound source coordinates, specifically configured as follows: Obtain the second straight-line distance between the second microphone and the radar module; Perform geometric operations on the sound source coordinates and the second straight-line distance to obtain the second distance.

7. The display device according to claim 1, characterized in that, The controller performs the synthesis of the updated first audio data and the second audio data to obtain the third audio data, specifically configured as follows: Obtain the first spectrum of the updated first audio data and the second spectrum of the updated second audio data; The first spectrum and the second spectrum are added together to obtain the third spectrum; Obtain the audio data corresponding to the third spectrum to obtain the third audio data.

8. The display device according to claim 1, characterized in that, After the controller performs the synthesis of the updated first audio data and the second audio data to obtain the third audio data, it is further configured to: Verification data is extracted from the updated first audio data or the second audio data, wherein the verification data is data in the updated first audio data or the second audio data whose frequency is lower than a frequency threshold; Extract the data to be verified from the third audio data, wherein the data to be verified is the data in the third audio data whose frequency is lower than the frequency threshold; If the data to be verified is the same as the data to be verified, a wake-up command is generated based on the third audio data, and in response to the wake-up command, the display is controlled to show the user interface.

9. The display device according to claim 8, characterized in that, The controller is also configured to: If the data to be verified is different from the data to be verified, then the second time difference between updating the first audio data and the third audio data is determined based on the sound source coordinates; Add the second starting time point to the second time difference to obtain the third starting time point; The starting time point for updating the first audio data is modified according to the third starting time point to obtain the second updated audio data; The second updated audio data and the second audio data are combined to obtain the updated third audio data.

10. A method for acquiring speech signals, characterized in that, Applied to the display device according to any one of claims 1-9; the method includes: In response to a user's voice command, the coordinates of the sound source are obtained. The coordinates of the sound source include a first coordinate, a second coordinate, and a third coordinate. The first and second coordinates are determined based on temperature information collected by the infrared imaging module. The third coordinate is determined based on the sound source reflection time of the transmitted and reflected waves collected by the radar module. Acquire first audio data and second audio data; wherein, the first audio data is the audio data corresponding to the voice command collected by the first microphone; and the second audio data is the audio data corresponding to the voice command collected by the second microphone. The time difference between the first audio data and the second audio data is determined based on the coordinates of the sound source. The first starting time point is added to the time difference to obtain the second starting time point; wherein, the first starting time point is the starting time point of the first audio data; Modify the start time of the first audio data according to the second start time point to obtain updated first audio data; The updated first audio data and the second audio data are synthesized to obtain the third audio data.

Citation Information

Patent Citations

  • Distributed audio capture and mixing controlling

    CN110089131A

  • Display device, server and voice instruction recognition method

    CN118283339A