Display device and method for controlling display device

By driving the collaborative work of multiple verification engines and wake word verification servers in parallel, the problem of efficient management of speech recognition services on display devices after software updates is solved, the wake word verification time is shortened, and scalability is improved.

CN121889769APending Publication Date: 2026-04-17LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LG ELECTRONICS INC
Filing Date
2023-09-27
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing display devices struggle to efficiently manage and use the wake word verification engine after software updates, resulting in long wake word verification times and a lack of scalability for speech recognition services.

Method used

The display device drives multiple verification engines in parallel, detects user voice through sensors, verifies whether the voice is a wake-up command using a controller, and makes a final determination by combining the verification engine results from the wake-up word verification server.

Benefits of technology

It enables efficient management of speech recognition verification even after software updates, shortens wake word verification time, and provides scalability for the verification engine.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121889769A_ABST
    Figure CN121889769A_ABST
Patent Text Reader

Abstract

Provided is a display device providing a voice recognition function, the display device may include: a sensor for detecting a user's voice; the controller is used for verifying whether the user voice is an instruction for waking up a voice recognition function or not, the controller drives a plurality of verification engines in parallel to verify the user voice, and a result object of the verification engine which completes user voice verification firstly is output preferentially, and then the result object of the verification engine which completes user voice verification firstly is output preferentially. And determining whether the user voice is an instruction for waking up the voice recognition function according to the results of the plurality of verification engines.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a display device and a control method for the display device. Background Technology

[0002] Recently, a hot topic in the field of multimedia devices such as mobile phones and televisions (TVs) is the new form factor. Form factor refers to the structural form of a product.

[0003] The reason why innovation in form factor is becoming increasingly important in the display industry is that, due to the increased mobility of consumers and the rapid development of device integration and intelligence, users are no longer bound by typical form factors that are customized for specific usage environments. They desire a form factor that can be used freely and conveniently without being constrained by usage environment conditions.

[0004] For example, breaking the conventional wisdom that televisions must be viewed horizontally, portrait-oriented TVs are gradually becoming more widespread. Portrait-oriented TVs are designed to switch screen orientations, reflecting the characteristics of Generation Z who are accustomed to watching content on mobile devices. They allow for easy browsing of images on social media or shopping websites, and even reading comments while watching videos, making them very convenient. Their advantages are further highlighted when portrait-oriented TVs are linked to smartphones via near-field communication (NFC) based mirroring functionality. Viewing regular television or movies can be switched horizontally.

[0005] As another example, rollable televisions and foldable smartphones share a common feature: they both utilize 'flexible displays'. Flexible displays, as the name suggests, refer to flexible electronic components. To achieve flexibility, thinness is paramount. Only when the substrate used to receive information and convert it into light is thin and flexible can its performance be ensured to remain undamaged over time.

[0006] Flexibility means that even when subjected to impact, it is not significantly affected. In flexible displays, the adhesive layers continuously bear pressure during bending or folding. To prevent internal damage under this pressure, a property that combines excellent durability with the ability to deform flexibly under pressure is needed.

[0007] Flexible displays are achieved, for example, based on OLEDs (Organic Light Emitting Diodes). OLEDs are displays that utilize organic light-emitting materials. Organic materials are more flexible than inorganic materials such as metals. Moreover, OLEDs are more competitive than other displays because of their thinner substrates. In the case of conventional LCD substrates, there are limitations in thinning due to the additional need for liquid crystal and glass.

[0008] As televisions evolve into new form factors, there is a growing demand for televisions that can be easily moved indoors and outdoors. This demand has been particularly strong in recent years due to the increased time spent at home caused by the coronavirus pandemic, leading to a surge in demand for second televisions. Furthermore, the rise in outdoor camping and other recreational activities further necessitates a new type of television that is easily portable and mobile.

[0009] This type of television differs from existing televisions in that it is driven by software called an operating system, which provides a user interface through the operating system and applications or programs.

[0010] Furthermore, it is becoming increasingly common to control televisions or run applications and programs on televisions via user voice commands.

[0011] When a television is manufactured, its operating system is specified for each model, or it is delivered to the user with the latest version available at the time of manufacture. With the introduction of various television body styles, multiple versions of the operating system are also offered, and as time goes on, updates to the operating system software are required.

[0012] As operating system software is updated, decisions or solutions need to be developed regarding how to manage and use the software (i.e., the verification engine) used to perform wake-word verification of user voice to provide voice command functionality or voice recognition-based services. This is because the recognition results of user voice depend on hardware conditions such as the microphone and the distance between the microphone and the speaker.

[0013] On the other hand, the following description is not limited to television sets, but also applies to all image or audio output devices that support voice commands based on speech recognition. Therefore, the term "display device" is used instead of "television set". Summary of the Invention

[0014] The problem that the invention aims to solve

[0015] To address the aforementioned problems, this invention proposes a scheme for managing and using software (i.e., a verification engine) that performs wake word verification based on user voice to activate speech recognition.

[0016] More specifically, the present invention proposes a specific solution in which settings are specified in a server for the software (i.e., the verification engine) used to perform wake word verification for the display device, and the verification engine is driven or used.

[0017] The problems to be solved by the present invention are not limited to those described above. Those skilled in the art should be able to clearly understand other problems not mentioned based on the following description.

[0018] Methods for solving problems

[0019] A display device is proposed that provides a voice recognition function. The display device may include: a sensor for detecting user voice; and a controller for verifying whether the user voice is an instruction for activating the voice recognition function. The controller drives multiple verification engines in parallel to verify the user voice, outputs the results of the verification engines that first complete the user voice verification, and determines whether the user voice is an instruction for activating the voice recognition function based on the results of the multiple verification engines.

[0020] A system including a display device and a wake-word verification server, wherein the display device may include: a sensor for detecting user voice; and a controller for verifying whether the user voice is an instruction for waking up a voice recognition function, the wake-word verification server for receiving the user voice from the display device, and including a controller for verifying whether the user voice is an instruction for waking up a voice recognition function, the controller of the display device controlling multiple verification engines of the display device or the wake-word verification server to drive in parallel to verify the user voice, and determining whether the user voice is an instruction for waking up the voice recognition function based on the results of the verification engines embedded in the display device and the verification engines embedded in the wake-word verification server.

[0021] A method for speech recognition is proposed, performed by a display device including: a sensor for detecting user speech; and a controller for verifying whether the user speech is an instruction to activate the speech recognition function. The method may include the following steps: recognizing the user speech; driving multiple verification engines in parallel to initiate the user speech verification process; and first outputting the results of the verification engines that have completed the user speech verification first, but determining whether the user speech is an instruction to activate the speech recognition function based on the results of the multiple verification engines.

[0022] The solutions to the problems described are only a part of the embodiments of the present invention. For those skilled in the art, various embodiments reflecting the technical features of the present invention can be derived and understood from the following detailed description of the present invention.

[0023] Invention Effects

[0024] The present invention has the following effects.

[0025] In this invention, even if the software of the display device is updated, the voice recognition verification engine can be managed independently, thus enabling efficient execution of wake word verification for enabling voice recognition services.

[0026] Furthermore, this invention drives multiple verification engines in parallel to perform wake word verification for enabling speech recognition services, thereby shortening the time required for wake word verification.

[0027] Furthermore, this invention drives multiple verification engines belonging to different entities in parallel to perform wake word verification for enabling speech recognition services, thereby shortening the time required for wake word verification and providing scalability for the algorithms used in the verification engines. Attached Figure Description

[0028] The accompanying drawings, which are included as part of the detailed description for ease of understanding of the invention, will provide embodiments of the invention and, together with the detailed description, illustrate the technical ideas of the invention.

[0029] Figure 1 It is a block diagram used to illustrate the various structures of a display device.

[0030] Figure 2 This is a diagram illustrating a display device provided according to an embodiment of the present invention.

[0031] Figure 3 The present invention illustrates the wake word recognition or wake-up process for waking up a speech recognition service.

[0032] Figure 4 The present invention illustrates the wake word recognition and verification process for wake-up voice recognition services.

[0033] Figure 5 A flowchart of the wake word recognition and verification process for wake-up speech recognition services provided by the present invention is shown.

[0034] Figure 6 A flowchart of the wake word recognition and verification process for wake-up speech recognition services provided by the present invention is shown.

[0035] Figure 7 A flowchart of the wake word recognition and verification process for wake-up speech recognition services provided by the present invention is shown.

[0036] Figure 8 A block diagram of a controller provided by the present invention for performing wake word recognition and verification for a wake-up speech recognition service is shown.

[0037] Figure 9 A block diagram of the system provided by the present invention for performing wake word recognition and verification for wake-up speech recognition services is shown. Detailed Implementation

[0038] Hereinafter, embodiments disclosed in this specification will be described in detail with reference to the accompanying drawings. However, identical or similar components will be assigned the same reference numerals regardless of the reference numerals, and their repeated descriptions will be omitted. The suffixes "module" and "part" used in connection with the components described below are assigned or used interchangeably merely for ease of writing and do not inherently have a distinguishing meaning or function. Furthermore, when describing the embodiments disclosed in this specification, detailed descriptions of relevant prior art will be omitted if it is determined that such detailed descriptions may obscure the spirit of the embodiments disclosed in this specification. Moreover, the accompanying drawings are only for the purpose of facilitating a good understanding of the embodiments disclosed in this specification. The technical ideas disclosed in this specification are not limited by the drawings and should be understood to cover all modifications, equivalents, or substitutions included within the scope of the present invention.

[0039] Terms such as "first," "second," etc., which include ordinal numbers, can be used to describe various constituent elements, but the constituent elements are not limited by the terms. The terms are only used to distinguish one constituent element from another.

[0040] When it is mentioned that a constituent element is "connected" or "linked" to another constituent element, it should be understood as being able to be directly connected or linked to its other constituent element, but there may be other constituent elements in between. Conversely, when it is mentioned that a constituent element is "directly connected" or "directly linked" to another constituent element, it should be understood as there are no other constituent elements in between.

[0041] Unless the context clearly indicates otherwise, singular expressions include plural expressions.

[0042] It should be understood that the terms "comprising" or "having" in this application are intended to specify the presence of features, figures, steps, actions, constituent elements, components or combinations thereof described in the specification, rather than to preclude the presence or additional possibilities of one or more other features or figures, steps, actions, constituent elements, components or combinations thereof.

[0043] In the following description, although it is referred to as display device 100, the name of the display device may vary, such as television or multimedia device, and the scope of the invention is not limited by its name.

[0044] Figure 1 This is a block diagram illustrating the various structures of a display device 100 provided in one embodiment of the present invention.

[0045] The display device 100 may include a broadcast receiver 110, an external device interface 171, a network interface 172, a storage unit 140, a user input interface 173, an input unit 130, a control unit 180, a display module 150, an audio output unit 160, and / or a power supply unit 190.

[0046] The broadcast receiver 110 may include a tuner unit 111 and a demodulation unit 112.

[0047] On the other hand, unlike the accompanying drawings, the display device 100 may also include only the external device interface 171 and the network interface 172, which are part of the broadcast receiver 110, the external device interface 171, and the network interface 172. That is, the display device 100 may also exclude the broadcast receiver 110.

[0048] The tuner unit 111 can select a broadcast signal from the broadcast signals received via an antenna (not shown) or a cable (not shown) that corresponds to a channel selected by the user or all channels already stored. The tuner unit 111 can convert the selected broadcast signal into an intermediate frequency signal or a baseband image or voice signal.

[0049] For example, when the selected broadcast signal is a digital broadcast signal, the tuner unit 111 can convert it into a digital intermediate frequency (DIF) signal; when the selected broadcast signal is an analog broadcast signal, it can convert it into an analog baseband image or voice signal (CVBS / SIF). That is, the tuner unit 111 can process both digital and analog broadcast signals. The analog baseband image or voice signal (CVBS / SIF) output from the tuner unit 111 can be directly input to the control unit 180.

[0050] On the other hand, the tuner unit 111 can sequentially select all broadcast signals of broadcast channels stored by the channel storage function from the received broadcast signals and convert them into intermediate frequency signals or baseband image or voice signals.

[0051] On the other hand, the tuner unit 111 can have multiple tuners to receive broadcast signals from multiple channels. Alternatively, it can also have a single tuner that can simultaneously receive broadcast signals from multiple channels.

[0052] The demodulation unit 112 can receive the digital intermediate frequency (DIF) signal converted in the tuner unit 111 to perform demodulation. After performing demodulation and channel decoding, the demodulation unit 112 can output a streaming signal (TS). At this time, the streaming signal can be a signal multiplexed from video signals, voice signals, or data signals.

[0053] The streaming signal output from the demodulation unit 112 can be input to the control unit 180. After performing demultiplexing, video / audio signal processing, etc., the control unit 180 can output video through the display module 150 and output audio through the audio output unit 160.

[0054] The sensing unit 120 refers to a device that senses changes within the display device 100 or senses external changes. For example, it may include at least one of a proximity sensor, an illumination sensor, a touch sensor, an infrared sensor (IR sensor), an ultrasonic sensor, an optical sensor (e.g., a camera), a voice sensor (e.g., a microphone), a battery gauge, and an environmental sensor (e.g., a hygrometer, a thermometer, etc.).

[0055] The control unit 180 can be controlled to check the status of the display device 100 based on the information collected in the sensor unit 120, and notify the user or make automatic adjustments to maintain the best status when a problem occurs.

[0056] Furthermore, based on information such as the audience or ambient light level sensed by the sensor unit, the content, image quality, and size of the image provided to the display module 150 can be differentiated to provide the best viewing environment. As smart display devices continue to advance, the functions integrated into the display devices are becoming more and more numerous, and the number of sensor units 20 is also increasing accordingly.

[0057] The input unit 130 may be located on one side of the main body of the display device 100. For example, the input unit 130 may include a touchpad, physical buttons, etc. The input unit 130 may receive various user commands related to the operation of the display device 100, and may transmit control signals corresponding to the input commands to the control unit 180.

[0058] Recently, as the bezels of display devices 100 have become smaller, more and more display devices 100 are designed to minimize the physical physical form of the button-type input section 130 exposed to the outside. Instead, minimal physical buttons can be provided on the back or side, and user input can be received via a touchpad or the user input interface section 173 described later through a remote control device 200.

[0059] The storage unit 140 can store programs for processing and controlling various signals within the control unit 180, and can also store processed image, voice, or data signals. For example, the storage unit 140 can store application programs designed to perform various tasks that the control unit 180 can handle, and selectively provide a portion of the stored application programs when the control unit 180 issues a request.

[0060] The programs stored in the storage unit 140 are not particularly limited as long as they can be executed by the control unit 180. The storage unit 140 can also perform functions for temporarily storing image, voice, or data signals received from external devices via the external device interface unit 171. The storage unit 140 can store information related to a designated broadcast channel through channel storage functions such as a channel mapping table.

[0061] Figure 1 An embodiment is shown in which the storage unit 140 and the control unit 180 are arranged separately, but the scope of the present invention is not limited thereto, and the storage unit 140 may also be included within the control unit 180.

[0062] The storage unit 140 may include at least one of volatile memory (e.g., DRAM (Dynamic Random Access Memory), SRAM (Static Random Access Memory), SDRAM (Synchronous Dynamic Random Access Memory) etc.) or non-volatile memory (e.g., flash memory, hard disk drive (HDD), solid-state drive (SSD) etc.).

[0063] The display module 150 can convert image signals, data signals, OSD (On-Screen Display) signals, control signals, or image signals, data signals, control signals, etc., processed in the control unit 180 or received from the interface unit 171 into drive signals. The display module 150 may include a display panel having multiple pixels.

[0064] The multiple pixels of the display panel can have RGB subpixels. Alternatively, the multiple pixels of the display panel can also have RGBW subpixels. The display module 150 can convert the image signals, data signals, OSD signals, control signals, etc., processed in the control unit 180 to generate drive signals for the multiple pixels.

[0065] The display module 150 can be a PDP (Plasma Display Panel), LCD (Liquid Crystal Display), OLED (Organic Light Emitting Diode), flexible display module, etc., and can also be a 3D display module. The 3D display module 150 can be divided into glasses-free type and glasses type.

[0066] The display device 100 includes: a display module that occupies most of the area of ​​the front surface; and a housing that covers the back, sides, etc. of the display module to encapsulate the display module.

[0067] Recently, display device 100 can utilize flexible display module 150, such as LED (Light Emitting Diodes) or OLED (Organic Light Emitting Diodes), to further realize curved images from a flat surface.

[0068] Previously, LCDs, which are not self-emissive, relied on backlight units to obtain light. A backlight unit is a device that equally supplies light from a light source to the liquid crystal on its front surface. While thinner backlight units have become possible for thinner LCDs, it has become difficult to implement them using flexible materials. When the backlight unit is bent, it becomes difficult to supply light evenly to the liquid crystal, resulting in problems such as variations in screen brightness.

[0069] Since LEDs or OLEDs are composed of pixels that emit their own light, they do not require a backlight unit and can be bent. Furthermore, because each element is self-emissive, changes in their positional relationship with adjacent elements do not affect their brightness, thus enabling the creation of a bendable display module 150 using LEDs or OLEDs.

[0070] OLED (Organic Light Emitting Diodes) panels were officially launched in mid-2010 and quickly replaced LCDs in the small and medium-sized display market. OLEDs are displays made using the self-emissive phenomenon of light emission when an electric current flows through fluorescent organic compounds. Their image response speed is faster than that of LCDs, so there is almost no image retention when displaying videos.

[0071] OLED uses three phosphor organic compounds, namely red, green and blue, which have self-emissive functions. It is a light-emitting display product that emits light by coupling electrons injected from the cathode and anode with positively charged particles in the organic material. Therefore, it does not need to use backlights (halo devices) to reduce color sensitivity.

[0072] LED (Light Emitting Diode) panels utilize a technology that uses an LED element as a pixel, enabling a reduction in the size of LED elements compared to the past, thus allowing for flexible display modules 150. Devices previously referred to as LED televisions used LEDs as the light source for a backlight unit that supplies light to the LCD; the LEDs themselves could not form an image.

[0073] The display module includes a display panel, a coupling magnet located on the back of the display panel, a first power supply unit, and a first signal module. The display panel may include multiple pixels R, G, and B. The multiple pixels R, G, and B may be formed in each region where multiple data lines intersect with multiple gate lines. The multiple pixels R, G, and B may be configured or arranged in a matrix.

[0074] For example, multiple pixels R, G, B may include red (hereinafter referred to as 'R') subpixels, green (hereinafter referred to as 'G') subpixels, and blue (hereinafter referred to as 'B') subpixels. Multiple pixels R, G, B may also include white (hereinafter referred to as 'W') subpixels.

[0075] In the display module 150, the side that displays the image can be referred to as the front or front surface. The side of the display module 150 where the image is not visible when it is displayed can be referred to as the rear or rear surface.

[0076] On the other hand, the display module 150 is composed of a touch screen and can be used as an input device in addition to being used as an output device.

[0077] The audio output unit 160 accepts the input signal after speech processing in the control unit 180 and outputs it as speech.

[0078] The interface unit 170 performs channel functions with various types of external devices connected to the display device 100. The interface unit may include not only wired methods for transmitting and receiving data via cable, but also wireless methods utilizing an antenna.

[0079] The interface section 170 may include at least one of the following: a wired / wireless headset port, an external charger port, a wired / wireless data port, a memory card port, a port for a connection device with an identification module, an audio I / O (Input / Output) port, a video I / O (Input / Output) port, and a headphone port.

[0080] As an example of a wireless method, the aforementioned broadcast receiver 110 may include not only broadcast signals, but also mobile communication signals, short-range communication signals, wireless Internet signals, etc.

[0081] The external device interface section 171 can transmit or receive data with connected external devices. For this purpose, the external device interface section 171 may include an A / V (audio / video) input / output section (not shown).

[0082] The external device interface unit 171 can be connected to external devices such as DVD (Digital Versatile Disk), Blu-ray, game devices, cameras, video recorders, computers (laptops), set-top boxes, etc. via wired / wireless means, and can also perform input / output operations with external devices.

[0083] Furthermore, the external device interface unit 171 can establish a communication network with various remote control devices 200, receive control signals related to the operation of the display device 100 from the remote control device 200, or transmit data related to the operation of the display device 100 to the remote control device 200.

[0084] The external device interface unit 171 may include a wireless communication unit (not shown) for short-range wireless communication with other electronic devices. Through this wireless communication unit (not shown), the external device interface unit 171 can exchange data with adjacent mobile terminals. Particularly in mirror mode, the external device interface unit 171 can receive device information, running application software information, application software images, etc., from the mobile terminal.

[0085] The network interface unit 172 can provide an interface for connecting the display device 100 to a wired / wireless network, including the Internet. For example, the network interface unit 172 can receive content or data provided by the Internet, a content provider, or a network operator via a network. Alternatively, the network interface unit 172 may include a communication module (not shown) for connecting to a wired / wireless network.

[0086] The external device interface section 171 and / or the network interface section 172 may include: a communication module for short-range communication, such as Wi-Fi (Wireless Fidelity), Bluetooth, Bluetooth Low Energy (BLE), Zigbee, NFC (Near Field Communication); and a communication module for cellular communication, such as LTE (long-term evolution), LTE-A (LTE Advance), CDMA (code division multiple access), WCDMA (wideband CDMA), UMTS (universal mobile telecommunications system), WiBro (Wireless Broadband).

[0087] The user input interface unit 173 can transmit signals input by the user to the control unit 180, or transmit signals from the control unit 180 to the user. For example, it can send / receive user input signals such as power on / off, channel selection, and screen settings from the remote control device 200, or transmit user input signals input from local keys (not shown) such as power button, channel button, volume button, and setting value to the control unit 180, or transmit user input signals input from a sensor unit (not shown) that senses user gestures to the control unit 180, or send signals from the control unit 180 to the sensor unit.

[0088] The control unit 180 may include at least one processor and may use the included processor to control the overall operation of the display device 100. Here, the processor may be a general processor, such as a CPU (central processing unit). Of course, the processor may also be a dedicated device or a processor based on other hardware, such as an ASIC (Application-Specific Integrated Circuit).

[0089] The control unit 180 can demultiplex the streams input through the tuner unit 111, demodulation unit 112, external device interface unit 171, or network interface unit 172, or process the demultiplexed signals to generate and output signals for outputting images or voice.

[0090] The image signal processed in the control unit 180 can be input to the display module 150 and displayed as an image corresponding to the image signal. Furthermore, the image signal processed in the control unit 180 can also be input to an external output device through the external device interface unit 171.

[0091] The voice signal processed in the control unit 180 can be output to the audio output unit 160. Furthermore, the voice signal processed in the control unit 180 can be input to an external output device via the external device interface unit 171. The control unit 180 may include a demultiplexing unit, an image processing unit, etc.

[0092] In addition, the control unit 180 can control the overall operation within the display device 100. For example, the control unit 180 can control the tuner unit 111 to select (Tuning) a channel that is equivalent to a user-selected channel or a stored channel for broadcasting.

[0093] Furthermore, the control unit 180 can control the display device 100 using user commands or internal programs input through the user input interface unit 173. On the other hand, the control unit 180 can control the display module 150 to display images. At this time, the images displayed on the display module 150 can be still images or videos, or 2D images or 3D images.

[0094] On the other hand, the control unit 180 can display specified 2D objects within the image displayed on the display module 150. For example, the object can be at least one of the following: a connected webpage (news website, e-magazine, etc.), an EPG (Electronic Program Guide), various menus, widgets, icons, still images, videos, and text.

[0095] On the other hand, the control unit 180 can modulate and / or demodulate the signal using amplitude shift keying (ASK). Here, amplitude shift keying (ASK) can refer to a method of modulating the signal by changing the amplitude of the carrier wave according to the data value, or a method of recovering the analog signal into digital data value according to the amplitude of the carrier wave.

[0096] For example, the control unit 180 can modulate the image signal using amplitude shift keying (ASK) and transmit it via a wireless communication module.

[0097] For example, the control unit 180 can demodulate and process the image signal received by the wireless communication module using amplitude shift keying (ASK).

[0098] Thus, the display device 100 can easily send and receive signals with other adjacent image display devices even without using inherent identifiers, such as MAC addresses (Media Access Control Addresses) or complex communication protocols, such as TCP / IP.

[0099] On the other hand, the display device 100 may also include a camera unit (not shown). The camera unit can capture images of the user. The camera unit can be implemented using a single camera, but is not limited to this; it can also be implemented using multiple cameras. Furthermore, the camera unit can be embedded in the display module 150 of the display device 100 or configured independently. The image information captured by the camera unit can be input to the control unit 180.

[0100] The control unit 180 can identify the user's position based on the images captured by the camera unit. For example, the control unit 180 can determine the distance (z-axis coordinate) between the user and the display device 100. In addition, the control unit 180 can also determine the x-axis coordinate and y-axis coordinate corresponding to the user's position within the display module 150.

[0101] The control unit 180 can sense the user's gestures based on images captured by the camera unit or signals sensed by the sensor unit, or a combination thereof.

[0102] The power supply unit 190 can supply power to the entire display device 100. In particular, it can supply power to the control unit 180, the display module 150 for displaying images, and the audio output unit 160 for outputting audio, which can be implemented as a system on chip (SOC).

[0103] Specifically, the power supply unit 190 may include: a converter (not shown) that converts AC power to DC power; and a DC / DC converter (not shown) that converts the level of the DC power.

[0104] On the other hand, the power supply unit 190's function is to obtain power from an external source and distribute that power to the various components. The power supply unit 190 may supply AC power by directly connecting to an external power source, or it may include a power supply unit 190 in the form of a rechargeable battery.

[0105] In the former case, using a wired cable makes movement difficult or limits the range of movement. In the latter case, although movement is free, it increases the weight and size of the battery. To charge, it needs to be directly connected to the power cable for a certain period of time, or coupled to the charging unit (not shown) that supplies power.

[0106] The charging unit can be connected to the display device via externally exposed terminals, or it can wirelessly approach and charge the built-in battery.

[0107] The remote control device 200 can transmit user input to the user input interface unit 173. For this purpose, the remote control device 200 can use Bluetooth, RF (Radio Frequency) communication, infrared (Infrared Radiation) communication, UWB (Ultra-wideband), ZigBee, etc. Furthermore, the remote control device 200 can receive image, voice, or data signals output from the user input interface unit 173 and display or output them via voice within the remote control device 200.

[0108] On the other hand, the aforementioned display device 100 may be a fixed or mobile digital broadcast receiver capable of receiving digital broadcasts.

[0109] on the other hand, Figure 1 The block diagram of the display device 100 shown is merely a block diagram for illustrating one embodiment of the present invention. The various components of the block diagram may be integrated, added, or omitted according to the specifications of the actual implemented display device 100.

[0110] That is, it can be configured to combine two or more constituent elements into one constituent element as needed, or to subdivide one constituent element into two or more constituent elements. Furthermore, the functions performed in each box are merely illustrative of embodiments of the present invention, and their specific actions or devices do not limit the scope of the invention.

[0111] Figure 2 This is a diagram illustrating a display device provided according to an embodiment of the present invention. Hereinafter, descriptions that are repeated above will be omitted.

[0112] Reference Figure 2 The display device 100 is a configuration in which the display module 150 is housed inside the housing 210. In this case, the housing 210 may include an upper housing 210a and a lower housing 210b, and is configured such that the upper housing 210a and the lower housing 210b can be opened and closed.

[0113] In one embodiment, an audio output unit 160 may be included in the upper housing 210a of the display device 100, and a motherboard, power board, power supply unit 190, battery, interface unit 170, sensing unit 120, and input unit 130 (including a local key) serving as a control unit 180 may be housed in the lower housing 210b. Here, the interface unit 170 may include a Wi-Fi module, Bluetooth module, and NFC module for communicating with external devices, and the sensing unit 120 may include an illuminance sensor and an IR sensor.

[0114] In one embodiment, the display module 150 may include a DC-DC board, a sensor, and an LVDS (low voltage differential signaling) conversion board.

[0115] Furthermore, in one embodiment, the display device 100 may also include four detachable legs 220a, 220b, 220c, and 220d. Here, the four legs 220a, 220b, 220c, and 220d may be attached to the lower housing 210b to separate the display device 100 from the ground.

[0116] Figure 2 The provided display device has a movable feature.

[0117] Figure 3 The present invention illustrates the wake word recognition or wake-up process for waking up a speech recognition service.

[0118] Figure 3 (a) through (d) show the sequential flow.

[0119] The user says "Hi, LG" to activate the voice recognition service for the display device 100. Figure 3 (a) This is the wake word used to activate the preset speech recognition service. This is just an example, and other wake words can also be set.

[0120] The display device 100 outputs the activation of voice command-based control (or voice recognition service) to the display screen. Figure 3 (b) thereby notifying the user to issue or should issue a voice command, or notifying the display device 100 that it is ready to recognize the user's voice command.

[0121] In response, users input the voice command by saying "search big bang theory". Figure 3 (c)).

[0122] The display device 100 recognizes the user's voice commands and performs a search based on them. Figure 3 (d)

[0123] By reference Figure 3 The description explains how the display device 100 provides voice recognition services. However, in order to access... Figure 3 In state (b), it is necessary to verify whether the wake word issued by the user is the preset wake word ("Hi, LG"), and the wake word issued by the user must be verified as the preset wake word.

[0124] This is because the display device 100 is for viewing images, but due to misrecognition, when the voice recognition service is enabled even though the user has not actually enabled it, it will operate in a way that obstructs the user's viewing environment.

[0125] Figure 4 The present invention illustrates the wake word recognition and verification process for wake-up voice recognition services.

[0126] Figure 4 (a) and (b) are respectively with Figure 3 Since (a) and (b) are the same, the explanation is omitted.

[0127] When a user speaks or uses voice input, the display device 100 executes the wake-word recognition (S101) and wake-word verification (S102) procedures. If the user's speech or voice is verified as a wake-word according to the wake-word verification (S102) procedure, then the process is initiated as follows: Figure 4 The speech recognition service shown in (b).

[0128] The wake word verification process will now be explained in more detail in connection with the update of the operating system software (hereinafter referred to as "software") of the display device.

[0129] Figure 5 A flowchart of the wake word recognition and verification process for wake-up speech recognition services provided by the present invention is shown.

[0130] Reference Figure 5 The diagram shows two display devices 1001 and 1002, but this is just an example. The process described later can also be applied using more than one display device and server 300.

[0131] Display devices 1001 and 1002 can be ordinary household televisions. Judging from the software version, display device 1002 is a newer model that was released slightly later than display device 1003.

[0132] Each display device is configured to use a pre-defined wake-word verification engine based on hardware and software version information. However, with the software version being updated to the latest version, the point of contention is how to use the wake-word verification engine.

[0133] like Figure 5 As shown, it is configured to continue using the wake word verification engine used before the software update, even after the software version of each display device is updated. However, this is just an example. The wake word verification engine to be used can also be determined according to the wake word verification engine information (S551, S552) provided by the server, such as using a version of the wake word verification engine that was improved from the wake word verification engine used before the software update, etc.

[0134] Return again Figure 5 Display devices 1001 and 1002 respectively receive voice input from the user to activate the voice recognition service, thereby performing wake-up word verification (S501, S502). To perform wake-up word verification, display devices 1001 and 1002 respectively use wake-up word verification engines (hereinafter referred to as "verification engines") A and B+C. "B+C" means using two verification engines, B and C.

[0135] Server 300 performs software updates for each display device 1001 and 1002 (software version changes from 2.00.13, 3.04.31 to 13.00.01) (S511, S521, S512, and S522). Software version 13.00.01 is the latest software version, assuming that the display devices designed or configured according to this software version use the wake word verification engine D.

[0136] Then, the long-distance voice recognition function is re-executed on each display device 1001, 1002 (S531, S532).

[0137] Display devices 1001 and 1002 can transmit a wake-word verification engine information request to server 300 (S541, S542). The wake-word verification engine information request is to request the wake-word verification engine to be used by display devices 1001 and 1002 from server 300. The wake-word verification engine information request may include software version information, hardware information, and server version information for display devices 1001 and 1002.

[0138] Therefore, server 300 can transmit the wake word verification engine information to be used by each display device 1001, 1002 to display devices 1001, 1002 (S551, S552).

[0139] The wake-word verification engine information to be used by each display device 1001 and 1002 depends on the hardware information of the display device. That is, the wake-word verification engine information to be used by the display device can be determined based on the hardware information of the display device. Speech recognition depends solely on the performance of the speech detector (i.e., a sensor such as a microphone) because, depending on the display device (i.e., the hardware information), the microphone position, the microphone-speaker configuration, and the distance between the microphone and speaker will affect the speech recognition results. Therefore, it is configured so that even if the software of the display device is updated to the latest version, the verification engine of the display device, at least based on the hardware information, can be used, rather than using the verification engine supported in the latest software. Furthermore, the software version of each display device 1001 and 1002 can also be used to determine the wake-word verification engine information to be used by the display devices 1001 and 1002.

[0140] The wake word verification engine information to be used by the display device can be preset. Alternatively, the wake word verification engine information to be used by the display device can also be determined dynamically by the server, at least based on the information included in the wake word verification engine information request.

[0141] Display devices 1001 and 1002 determine the verification engine according to the wake-word verification engine information received from server 300 (S561, S562). For example, when the received wake-word verification engine information indicates verification engine A, the display device determines verification engine A as the verification engine to be used. That is, although the latest software version (13.00.01) of the display device is set to require the display device to use verification engine D, it is set to use the verification engines A and B+C most suitable for display devices 1001 and 1002 according to the wake-word verification engine information requests and exchanges or transmissions with server 300. In short, the verification engines supported in the latest version of the software of display devices 1001 and 1002 may be different from the verification engines to be used by display devices 1001 and 1002.

[0142] The display devices 1001 and 1002 use the determined verification engine to perform wake word verification (S571, S572).

[0143] Figure 6 A flowchart of the wake word recognition and verification process for wake-up speech recognition services provided by the present invention is shown.

[0144] Reference Figure 6 The diagram shows two display devices 1003 and 1004, but this is just an example. The process described later can be applied using more than one display device and server 300.

[0145] Display device 1003 can be a regular indoor television set, and display device 1004 can be an outdoor television set.

[0146] As can be seen from the software version, the display device 1003 is a newer model that was released slightly later than the display device 1004.

[0147] Figure 6 and Figure 5 The differences lie in the hardware and software version information of each display device; the rest are the same. Figure 5 same.

[0148] On the other hand, combining Figure 5 It has been explained that the verification engine used by display devices 1003 and 1004 after the software update is the same as the verification engine used before the software update. This is just an example, using the verification engine given in the wake word verification engine information (S651, S652).

[0149] Figure 7 A flowchart of the wake word recognition and verification process for wake-up speech recognition services provided by the present invention is shown. Figure 7 The process is executed by the display device 100. The display device 100 includes: a sensor for detecting user voice; and a controller for verifying whether the user voice is an instruction to activate the voice recognition function. Furthermore, the display device 100 may also include a display 150.

[0150] The display device 100 can recognize the wake word used to wake up the voice recognition service by recognizing the user's voice (S710).

[0151] Then, the display device 100 runs the wake word verification engine and issues a ticket (S720).

[0152] Tickets are generated for each identified wake word. That is, the same ticket is used for each wake word; simply put, tickets can be set to have positive integer values. Furthermore, the number of tickets generated corresponds to the number of verification engines to be used.

[0153] Therefore, when assuming a wake word is identified, and as... Figure 5 as well as Figure 6 As explained in the description, when two verification engines are received from server 300 (2), two tickets with a value of "1" are issued. Furthermore, in order to manage the number of tickets with a value of "1" issued, a verification count value can be defined, which is set to the same value as the number of tickets with a value of one issued (i.e., the verification count value is 2 in the example above).

[0154] If wake words are input and recognized consecutively, tickets with different values ​​will be issued for each wake word. For example, a ticket with a value of "1" will be issued for the first wake word, a ticket with a value of "2" will be issued for the second wake word, and so on.

[0155] The display device 100 can run the wake word verification engine in parallel to perform wake word verification (S730-1, S730-2, ..., S730-N). Figure 7 The example shows the case of using N verification engines, where N is an integer greater than 1. For reference, when there is only one verification engine, they will not "run in parallel".

[0156] Each verification engine can use verification algorithms to output verification results for the identified wake words. The verification results obtained from the verification engines can be stored and managed as in the example in Table 1 below.

[0157] The verification result is determined by collecting the results of one or more verification engines for tickets with the same value, i.e., for a single wake word. The display device 100 can then confirm whether the verification results for all issued tickets have been generated or collected (S740). The aforementioned verification count value can be used here. Each time the verification result of each verification engine is generated and collected, the verification count value set when issuing the ticket is decreased by 1, thus the result of S740 can be determined by whether the verification count value becomes 0.

[0158] The display device 100 can collect the verification results determined in each verification engine to determine the verification result (S750). At this time, when multiple verification engines are used, the display device 100 can apply individual weights to the results of each verification engine to derive the verification result. The individual weights applied to the results of each verification engine can be the same, or they can be different, or at least some of them can be the same and the rest different.

[0159] The verification results are divided into pass and fail. The thresholds or benchmarks used to determine these results can be set and implemented in a variety of ways.

[0160] The display device 100 can transmit the verification result to a storage device such as a memory for managing the verification result (S760).

[0161] On the other hand, when the verification result is 'pass', the verification result can be passed to other services, applications, programs, or other services or related devices supported by the display device 100 within the display device 100.

[0162] The following table illustrates how the display device 100 stores and manages the verification results related to the wake word verification engine as a Q (Queue) data structure.

[0163] [Table 1]

[0164] Ticket represents the ticket value, VeriCnt represents the verification count value, and VeriRst represents the verification result exported from the verification engine.

[0165] Figure 8 A block diagram of the control unit 180 of the display device 100 for performing wake-up word recognition and verification for wake-up voice recognition services provided by the present invention is shown. The "control unit" may be referred to as a "controller". Furthermore, reference is made below... Figure 8 A brief explanation of the block diagram is provided, regarding the references above. Figures 3 to 7 Any parts of the invention described repeatedly will be omitted hereafter, referring to the corresponding description.

[0166] The controller 180 may include a voice recognizer 181, a verification executor 182, a verification result judge 183, and a verification result transmitter 184.

[0167] The voice recognizer 181 can recognize a user's voice. The voice recognizer 181 can be composed of a sensor such as a microphone.

[0168] Verification executor 182 can verify whether the recognized user voice is equivalent to the wake word used to activate the voice recognition service. Verification executor 182 can use the verification engine described above, and can use at least one verification engine.

[0169] The verification result judge 183 can collect verification results from various verification engines and thus export the verification result for a wake word.

[0170] The verification result transmitter 184 can store the verification result obtained from the verification result judge 183 to a storage device such as a memory, and transmit the verification result to other services, applications, programs, or other services or related devices supported by the display device 100 within the display device 100.

[0171] Figure 9 A block diagram of the system provided by the present invention for performing wake word recognition and verification for wake-up speech recognition services is shown.

[0172] Based on the above description, the case where the verification engine is embedded in the display device 100 or the controller 180 of the display device 100, or other components of the display device 100, has been explained.

[0173] However, the verification engine can also exist outside the display device 100.

[0174] As an example, such as Figure 9 As shown, the verification engine can be built into the wake word verification server 300. Figure 9 The wake word verification server 300 is shown as being related to Figures 5 to 6 The server 300 is the same, but it can also be other servers or devices.

[0175] When using the wakeword verification server 300's verification engine, there are two possible scenarios.

[0176] 1) Using only the wakeword to verify the server's 300 verification engine

[0177] Consider the following scenario: Display device 100 receives wake-word verification engine information from server 300 ( Figure 5 as well as Figure 6 The display device 100 can transmit the user's voice (s551, S552, S651, S652) to the server 300 based on the wake-word verification engine information.

[0178] The server 300 that receives the user's voice can run a wake word verification engine to perform wake word verification on the user's voice.

[0179] When the wake word verification server 300 has multiple verification engines, the display device 100 can control these multiple verification engines to perform verification in parallel.

[0180] Server 300 can obtain the verification results and retransmit them to display device 100. When display device 100 receives the verification results of all verification engines from server 300, it can determine the final verification result (i.e., whether it is a wake word) based on this result.

[0181] 2) Case where the verification engine of the wake word verification server 300 is used in conjunction with the verification engine of the display device 100.

[0182] Most of the embodiments are similar to those in 1) above. The case where the wake-word verification engine information received by the display device 100 includes not only the verification engine of the display device 100 but also the verification engine of the server 300. In this case, the display device 100 can also transmit the user's voice (speaking) to the server 300. Furthermore, the display device 100 can be controlled to allow its own verification engine and the verification engine of the wake-word verification server 300 to perform verification in parallel.

[0183] The server 300, upon receiving the user's voice, can run a wake-word verification engine to perform wake-word verification on the user's voice. The server 300 can obtain the verification result and retransmit it to the display device 100.

[0184] The display device 100 can determine the final verification result based on the verification result received from the server 300 and the verification result obtained by running its own verification engine.

[0185] On the other hand, in case 2), the authentication engine of the server 300 used may be different from the authentication engine built into the display device 100.

[0186] As shown in 1) or 2), the advantage of using the verification engine of server 300 is that the verification algorithm used for the verification engine can be applied more flexibly. Server 300 can collect verification results from the display device and collect verification results from those that have personally performed wake word verification, thereby optimizing or updating the verification algorithm through machine learning and other methods.

[0187] In this way, a system consisting of a display device 100 and a server 300 can be configured to perform wake word recognition and verification. (No reference provided) Figure 9 Explanation and reference Figures 3 to 7 The content of the explanation can also be applied to Figure 9 The system.

[0188] Furthermore, as another aspect of the present invention, the actions of the above-described proposals or inventions can also be provided as code that can be implemented, carried out, or executed by a "computer" (including the general concept of a system on chip (SoC) or (micro) processor, etc.), or a computer-readable storage medium storing or containing said code, or a computer program product, etc. The scope of the present invention can be extended to said code, or a computer-readable storage medium storing or containing said code, or a computer program product.

[0189] The preferred embodiments of the invention disclosed above have been described in detail in a manner that enables those skilled in the art to implement and practice the invention. While the invention has been described above with reference to preferred embodiments, those skilled in the art should understand that various modifications and variations can be made to the invention as set forth in the following claims. Therefore, the invention is not limited to the embodiments shown in this specification, but is intended to be given the broadest scope consistent with the principles and novel features disclosed herein.

Claims

1. A display device providing voice recognition functionality, the display device comprising: Sensors used to detect user voice; as well as The controller is used to verify whether the user's voice is a command to activate the voice recognition function. The controller drives multiple verification engines in parallel to verify the user's voice. It first outputs the results of the verification engine that completes the verification of the user's voice first. Based on the results of the multiple verification engines, it is determined whether the user's voice is an instruction used to activate the voice recognition function.

2. The display device according to claim 1, wherein, The controller transmits a request to the server regarding the verification engine, and in response to the request receives information about the plurality of verification engines.

3. The display device according to claim 1, wherein, The controller applies different weights to the results of the multiple verification engines to obtain the results.

4. The display device according to claim 1, wherein, The plurality of verification engines are determined based on the hardware or software version of the display device.

5. The display device according to claim 1, wherein, The controller transmits the user's voice to the server. Receive the verification result of the user's voice from the server's verification engine. The result received from the server is used to verify whether the user's voice is an instruction to activate the voice recognition function.

6. The display device according to claim 5, wherein, The server's verification engine is different from the verification engine embedded in the display device.

7. The display device according to claim 1, wherein, The plurality of verification engines are different from the verification engine used to perform wake word verification supported in the latest version of the operating system software of the display device.

8. A system comprising a display device and a wake-word verification server, wherein, The display device includes: a sensor for detecting user voice; and a controller for verifying whether the user voice is a command to activate the voice recognition function. The wake-word verification server is used to receive the user's voice from the display device, and includes a controller for verifying whether the user's voice is an instruction for waking up the voice recognition function. The controller of the display device controls multiple verification engines of the display device or the wake word verification server to drive them in parallel to verify the user's voice. Based on the results of the verification engine embedded in the display device and the results of the verification engine embedded in the wake word verification server, it is determined whether the user's voice is an instruction for waking up the voice recognition function.

9. A method for a speech recognition function, the method being performed by a display device, the display device comprising: Sensors used to detect user voice; The method includes a controller for verifying whether the user's voice is an instruction to activate the voice recognition function, and the method comprises the following steps: Recognize the user's voice; Multiple verification engines are driven in parallel to initiate the verification process of the user's voice. as well as First, the results of the verification engines that have completed the user voice verification are output, and based on the results of the multiple verification engines, it is determined whether the user voice is an instruction used to wake up the voice recognition function.

10. The method according to claim 9, wherein, The method includes the following steps: Send a request to the server regarding the verification engine; and In response to the request, information about the plurality of verification engines is received.

11. The method according to claim 9, wherein, The method includes the step of applying different weights to the results of the plurality of verification engines to obtain the results.

12. The method according to claim 9, wherein, The plurality of verification engines are determined based on the hardware and software versions of the display device.

13. The method according to claim 9, wherein, The method includes the following steps: Transmit the user's voice to the server; and Receive the verification result of the user's voice from the server's verification engine. The result received from the server is used to verify whether the user's voice is an instruction to activate the voice recognition function.

14. The method according to claim 13, wherein, The server's verification engine is different from the verification engine embedded in the display device.

15. The method according to claim 9, wherein, The plurality of verification engines are different from the verification engines used to perform wake word verification supported in the latest version of the operating system software of the display device.