Display device and operation method therefor

The AI device enhances voice recognition in digital TV services by anticipating and correcting errors in content names and search terms, thereby reducing search failures and improving user experience.

WO2025095151A1PCT designated stage expired Publication Date: 2025-05-08LG ELECTRONICS INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2023/016981
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-10-30
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

Conventional voice recognition systems in digital TV services often misinterpret content titles or misunderstand user voices, leading to search failures and inconvenient user experiences.

Method used

An AI device that anticipates and identifies potential errors in content names or search terms by analyzing voice data, utilizing an AI server to generate error names, and adjusting the number of stored error names based on popularity and search frequency.

Benefits of technology

Reduces search failures and improves user experience by accurately identifying and correcting voice recognition errors, ensuring that users can find desired content even with partial or misinterpreted voice commands.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2023016981_08052025_PF_FP_ABST
    Figure KR2023016981_08052025_PF_FP_ABST
Patent Text Reader

Abstract

An artificial intelligence device according to an embodiment of the present disclosure may comprise: a memory for storing a name and multiple error names matched to the name; a communication unit for communicating with an electronic device or a generative artificial intelligence (AI) server; and a processor for receiving, from the electronic device, voice data corresponding to a voice command uttered by a user, obtaining an analysis result on the basis of the received voice data, obtaining the name matched to the multiple error names when any one of the multiple stored error names is included in the obtained analysis result, and transmitting a search result for the obtained name to the electronic device.
Need to check novelty before this filing date? Find Prior Art

Description

Display device and method of operation thereof

[0001] The present disclosure relates to a display device and to proactively responding to mispronunciation or misrecognition of a speaker.

[0002] Digital TV services utilizing wired or wireless networks are becoming more widespread. Digital TV services can offer a variety of services not available with existing analog broadcasting services.

[0003] For example, IPTV (Internet Protocol Television), a type of digital TV service, and smart TV services offer interactivity, allowing users to actively choose the type of program they want to watch and when. Building on this interactivity, IPTV and smart TV services can also offer a variety of additional services, such as internet search, home shopping, and online games.

[0004] Recent TVs provide voice recognition services that recognize the voice spoken by the user and provide services.

[0005] However, in the past, in a content search environment through voice search on TV, there were many cases where the content title was spoken incorrectly or the content name spoken by the user was incorrectly recognized by the server.

[0006] These misfires or misrecognitions have a certain degree of regularity, but at the same time they are complex enough that they cannot be covered by a rule-base.

[0007] Users may experience inconvenience in the provision of voice recognition services due to mispronunciation or misrecognition of content.

[0008] The purpose of this disclosure is to proactively identify the number of cases in which a user's content name or search term is mispronounced or misrecognized, and to proactively respond to mispronunciations and misrecognitions.

[0009] The present disclosure aims to reduce search failures when a user speaks only some of the key words of a content name or substitutes some of the key words with synonyms.

[0010] An artificial intelligence device according to an embodiment of the present disclosure may include a memory that stores a name and a plurality of error names matching the name; a communication unit that communicates with an electronic device or a generation AI (Artificial Intelligence) server; and a processor that receives voice data corresponding to a voice command uttered by a user from the electronic device, obtains an analysis result based on the received voice data, and, if any one of the plurality of stored error names is included in the obtained analysis result, obtains the name matching the plurality of error names, and transmits a search result for the obtained name to the electronic device.

[0011] The above name may be either a content name or a search term.

[0012] The above processor can transmit a prompt requesting the generation of error names for the names of new content obtained from the new content database to the generation AI server, and receive error names for the names of the new content from the generation AI server.

[0013] The above processor can transmit a prompt requesting the generation of error names for the search words having a search frequency greater than a certain frequency to the generation AI server, and receive error names for the search words from the generation AI server.

[0014] The processor may transmit a command to the electronic device to cause the electronic device to output a pop-up window to confirm whether the acquired name matches the user's intention.

[0015] The above processor can adjust the number of stored error names according to the popularity ranking of the content corresponding to the above content name or the search ranking of the above search word.

[0016] The processor may reduce the number of error names as the popularity ranking or the search ranking decreases.

[0017] If the above name includes one or more keywords, each of the plurality of error names may be a combination of variant words of each keyword.

[0018] According to one embodiment of the present disclosure, a method of operating an artificial intelligence device may include: storing a name and a plurality of error names matching the name; receiving voice data corresponding to a voice command uttered by a user from an electronic device; obtaining an analysis result based on the received voice data; obtaining a name matching the plurality of error names when any one of the stored plurality of error names is included in the obtained analysis result; and transmitting a search result for the obtained name to the electronic device.

[0019] The above name may be either a content name or a search term.

[0020] The above method of operation may further include a step of transmitting a prompt requesting generation of an error name for a name of a new content acquired from a new content database to a generation AI server; and a step of receiving error names for the name of the new content from the generation AI server.

[0021] The above method of operation may further include a step of transmitting a prompt requesting generation of an error name for the search term having a search frequency of a certain frequency or higher to a generation AI server; and a step of receiving error names for the search term from the generation AI server.

[0022] The above operating method may further include a step of transmitting a command to the electronic device to cause the electronic device to output a pop-up window for confirming whether the obtained name matches the user's intention.

[0023] The above operating method may further include a step of adjusting the number of stored error names according to the popularity ranking of the content corresponding to the content name or the search ranking of the search word.

[0024] The above adjusting step may include a step of reducing the number of error names as the popularity ranking or the search ranking becomes lower.

[0025] According to an embodiment of the present disclosure, by proactively addressing the possibility of incorrect content search, search failures due to frequent recognition failures caused by errors in voice recognition according to STT can be reduced.

[0026] Accordingly, the user experience for voice search can be improved by reducing search failures when the user speaks only some of the key words of the content name or replaces some of the key words with synonyms.

[0027] Figure 1 is a block diagram illustrating the configuration of a display device according to one embodiment of the present invention.

[0028] Figure 2 is a block diagram of a remote control device according to one embodiment of the present invention.

[0029] Figure 3 shows an example of an actual configuration of a remote control device according to one embodiment of the present invention.

[0030] Figure 4 shows an example of utilizing a remote control device according to an embodiment of the present invention.

[0031] FIG. 5 illustrates an artificial intelligence (AI) server according to one embodiment of the present disclosure.

[0032] FIG. 6 is a diagram for explaining the configuration of an AI system according to one embodiment of the present disclosure.

[0033] FIG. 7 is a ladder diagram for explaining an operation method of an AI system according to an embodiment of the present disclosure.

[0034] FIG. 8 is a diagram illustrating a process of receiving a plurality of error names corresponding to content names from a generation AI server according to one embodiment of the present disclosure.

[0035] Figures 9a and 9b are diagrams illustrating the output results of LLM according to a prompt to generate an error name of a content name.

[0036] FIG. 10 is a diagram illustrating a table matching content names and multiple error names according to one embodiment of the present disclosure.

[0037] Figure 11 is a diagram illustrating a process for providing search results for desired content even when a user incorrectly pronounces the content name.

[0038] Figure 12 is a diagram illustrating an example of providing a pop-up window to confirm whether the content name intended by the user is correct when the user incorrectly pronounces the content name.

[0039] Hereinafter, embodiments related to the present invention will be described in more detail with reference to the drawings. The suffixes "module" and "part" used in the following description for components are assigned or used interchangeably solely for the convenience of writing the specification, and do not in themselves have distinct meanings or roles.

[0040] A display device according to an embodiment of the present invention is, for example, an intelligent display device that adds computer-assisted functionality to its broadcast reception function. While faithfully performing the broadcast reception function, it can also be equipped with Internet functionality and other features, providing a more user-friendly interface, such as a manual input device, touch screen, or space remote control. Furthermore, with support for wired or wireless Internet functionality, it can connect to the Internet and a computer, enabling functions such as email, web browsing, banking, or gaming. A standardized, general-purpose operating system can be used for these various functions.

[0041] Accordingly, the display device described in the present invention can perform various user-friendly functions, for example, by allowing various applications to be freely added or deleted on a general-purpose operating system kernel. More specifically, the display device can be a network TV, HBB TV, smart TV, LED TV, OLED TV, etc., and in some cases, it can also be applied to smartphones.

[0042] FIG. 1 is a block diagram illustrating the configuration of a display device according to one embodiment of the present invention.

[0043] Referring to FIG. 1, the display device (100) may include a broadcast receiving unit (130), an external device interface (135), a memory (140), a user input interface (150), a controller (170), a wireless communication interface (173), a display (180), a speaker (185), and a power supply circuit (190).

[0044] The broadcast receiving unit (130) may include a tuner (131), a demodulator (132), and a network interface (133).

[0045] The tuner (131) can select a specific broadcast channel according to a channel selection command. The tuner (131) can receive a broadcast signal for the selected specific broadcast channel.

[0046] The demodulator (132) can separate the received broadcast signal into a video signal, an audio signal, and a data signal related to the broadcast program, and can restore the separated video signal, audio signal, and data signal into a form that can be output.

[0047] The external device interface (135) can receive an application or a list of applications within an adjacent external device and transmit it to the controller (170) or memory (140).

[0048] The external device interface (135) can provide a connection path between the display device (100) and the external device. The external device interface (135) can receive one or more of images and audio output from an external device connected wirelessly or wiredly to the display device (100) and transmit them to the controller (170). The external device interface (135) can include a plurality of external input terminals. The plurality of external input terminals can include an RGB terminal, one or more HDMI (High Definition Multimedia Interface) terminals, and a component terminal.

[0049] A video signal of an external device input through an external device interface (135) can be output through a display (180). A voice signal of an external device input through an external device interface (135) can be output through a speaker (185).

[0050] An external device that can be connected to the external device interface (135) may be any one of a set-top box, a Blu-ray player, a DVD player, a game console, a sound bar, a smartphone, a PC, a USB memory, and a home theater, but this is only an example.

[0051] The network interface (133) may provide an interface for connecting the display device (100) to a wired / wireless network including the Internet. The network interface (133) may transmit or receive data to or from other users or other electronic devices via the connected network or another network linked to the connected network.

[0052] Additionally, some content data stored in the display device (100) can be transmitted to a selected user or electronic device among other users or other electronic devices pre-registered in the display device (100).

[0053] The network interface (133) can access a predetermined web page through a connected network or another network linked to the connected network. That is, by accessing a predetermined web page through a network, data can be transmitted or received with the corresponding server.

[0054] In addition, the network interface (133) can receive content or data provided by a content provider or network operator. That is, the network interface (133) can receive content such as movies, advertisements, games, VOD, broadcast signals, etc. and information related thereto provided from a content provider or network provider via a network.

[0055] Additionally, the network interface (133) can receive firmware update information and update files provided by the network operator, and can transmit data to the Internet or content provider or network operator.

[0056] The network interface (133) can select and receive a desired application from among applications open to the public via a network.

[0057] The memory (140) stores a program for each signal processing and control within the controller (170), and can store signal-processed image, voice, or data signals.

[0058] In addition, the memory (140) may perform a function for temporary storage of video, audio, or data signals input from an external device interface (135) or a network interface (133), and may store information about a specific image through a channel memory function.

[0059] The memory (140) can store an application or a list of applications input from an external device interface (135) or a network interface (133).

[0060] The display device (100) can play content files (video files, still image files, music files, document files, application files, etc.) stored in the memory (140) and provide them to the user.

[0061] The user input interface (150) can transmit a signal input by the user to the controller (170) or transmit a signal from the controller (170) to the user. For example, the user input interface (150) can receive and process control signals such as power on / off, channel selection, and screen setting from the remote control device (200) according to various communication methods such as Bluetooth, Ultra Wideband (WB), ZigBee, Radio Frequency (RF) communication, or infrared (IR) communication, or process control signals from the controller (170) to be transmitted to the remote control device (200).

[0062] In addition, the user input interface (150) can transmit control signals input from local keys (not shown) such as a power key, channel key, volume key, and setting value to the controller (170).

[0063] An image signal processed by the controller (170) can be input to the display (180) and displayed as an image corresponding to the image signal. In addition, an image signal processed by the controller (170) can be input to an external output device through an external device interface (135).

[0064] The voice signal processed by the controller (170) can be output as audio to the speaker (185). In addition, the voice signal processed by the controller (170) can be input to an external output device through the external device interface (135).

[0065] In addition, the controller (170) can control the overall operation within the display device (100).

[0066] In addition, the controller (170) can control the display device (100) by a user command or internal program input through the user input interface (150), and can connect to a network to enable the user to download a desired application or application list into the display device (100).

[0067] The controller (170) enables the user-selected channel information, etc. to be output through a display (180) or speaker (185) together with processed video or audio signals.

[0068] In addition, the controller (170) allows a video signal or audio signal from an external device, for example, a camera or camcorder, input through the external device interface (135) to be output through the display (180) or speaker (185) in accordance with an external device video playback command received through the user input interface (150).

[0069] Meanwhile, the controller (170) can control the display (180) to display an image, for example, a broadcast image input through a tuner (131), an external input image input through an external device interface (135), an image input through a network interface, or an image stored in a memory (140) can be controlled to be displayed on the display (180). In this case, the image displayed on the display (180) can be a still image or a moving image, and can be a 2D image or a 3D image.

[0070] In addition, the controller (170) can control the playback of content stored in the display device (100), received broadcast content, or external input content input from outside, and the content can be in various forms such as broadcast video, external input video, audio file, still image, connected web screen, and document file.

[0071] The wireless communication interface (173) can communicate with an external device through wired or wireless communication. The wireless communication interface (173) can perform short-range communication with the external device. To this end, the wireless communication interface (173) can support short-range communication using at least one of Bluetooth™, RFID (Radio Frequency Identification), Infrared Data Association (IrDA), UWB (Ultra Wideband), ZigBee, NFC (Near Field Communication), Wi-Fi (Wireless-Fidelity), Wi-Fi Direct, and Wireless USB (Wireless Universal Serial Bus) technologies. This wireless communication interface (173) can support wireless communication between the display device (100) and a wireless communication system, between the display device (100) and another display device (100), or between the display device (100) and a network in which the display device (100, or an external server) is located via a short-range wireless communication network (Wireless Area Network). The short-range wireless communication network can be a short-range wireless personal area network (Wireless Personal Area Network).

[0072] Here, the other display device (100) may be a wearable device (e.g., a smartwatch, smart glasses, a head-mounted display (HMD)) or a mobile terminal such as a smart phone that can exchange data with (or be linked to) the display device (100) according to the present invention. The wireless communication interface (173) may detect (or recognize) a wearable device capable of communication around the display device (100).

[0073] Furthermore, if the detected wearable device is a device certified to communicate with the display device (100) according to the present invention, the controller (170) can transmit at least a portion of the data processed in the display device (100) to the wearable device via the wireless communication interface (173). Accordingly, a user of the wearable device can utilize the data processed in the display device (100) via the wearable device.

[0074] The display (180) can generate a driving signal by converting a video signal, data signal, OSD signal processed by the controller (170) or a video signal, data signal, etc. received from an external device interface (135) into R, G, and B signals, respectively.

[0075] Meanwhile, since the display device (100) illustrated in FIG. 1 is merely an embodiment of the present invention, some of the illustrated components may be integrated, added, or omitted depending on the specifications of the display device (100) actually implemented.

[0076] That is, two or more components may be combined into a single component, or a single component may be subdivided into two or more components, as needed. Furthermore, the functions performed by each block are intended to illustrate embodiments of the present invention, and their specific operations or devices do not limit the scope of the present invention.

[0077] According to another embodiment of the present invention, the display device (100) may receive and play back an image through a network interface (133) or an external device interface (135) without having a tuner (131) and a demodulator (132), unlike that shown in FIG. 1.

[0078] For example, the display device (100) may be implemented separately as an image processing device, such as a set-top box, for receiving broadcast signals or contents according to various network services, and a content playback device for playing contents input from the image processing device.

[0079] In this case, the operating method of the display device according to the embodiment of the present invention described below may be performed by any one of the display device (100) described with reference to FIG. 1, as well as an image processing device such as the separated set-top box, or a content playback device having a display (180) and an audio output unit (185).

[0080] Next, a remote control device according to an embodiment of the present invention will be described with reference to FIGS. 2 and 3.

[0081] FIG. 2 is a block diagram of a remote control device according to an embodiment of the present invention, and FIG. 3 shows an example of an actual configuration of a remote control device (200) according to an embodiment of the present invention.

[0082] First, referring to FIG. 2, the remote control device (200) may include a fingerprint recognition device (210), a wireless communication circuit (220), a user input interface (230), a sensor (240), an output interface (250), a power supply circuit (260), a memory (270), a controller (280), and a microphone (290).

[0083] Referring to FIG. 2, the wireless communication circuit (220) transmits and receives signals with any one of the display devices according to the embodiments of the present invention described above.

[0084] The remote control device (200) may be equipped with an RF circuit (221) capable of transmitting and receiving signals with the display device (100) according to RF communication standards, and an IR circuit (223) capable of transmitting and receiving signals with the display device (100) according to IR communication standards. In addition, the remote control device (200) may be equipped with a Bluetooth circuit (225) capable of transmitting and receiving signals with the display device (100) according to Bluetooth communication standards. In addition, the remote control device (200) may be equipped with an NFC circuit (227) capable of transmitting and receiving signals with the display device (100) according to NFC (Near Field Communication) communication standards, and a WLAN circuit (229) capable of transmitting and receiving signals with the display device (100) according to WLAN (Wireless LAN) communication standards.

[0085] In addition, the remote control device (200) transmits a signal containing information about the movement of the remote control device (200) to the display device (100) through a wireless communication circuit (220).

[0086] Meanwhile, the remote control device (200) can receive a signal transmitted by the display device (100) through the RF circuit (221), and, if necessary, can transmit commands for power on / off, channel change, volume change, etc. to the display device (100) through the IR circuit (223).

[0087] The user input interface (230) may be configured as a keypad, buttons, a touchpad, or a touch screen. The user can input commands related to the display device (100) to the remote control device (200) by operating the user input interface (230). If the user input interface (230) includes a hard key button, the user can input commands related to the display device (100) to the remote control device (200) by pushing the hard key button. This will be described with reference to FIG. 3.

[0088] Referring to FIG. 3, the remote control device (200) may include a plurality of buttons. The plurality of buttons may include a fingerprint recognition button (212), a power button (231), a home button (232), a live button (233), an external input button (234), a volume control button (235), a voice recognition button (236), a channel change button (237), a confirmation button (238), and a back button (239).

[0089] The fingerprint recognition button (212) may be a button for recognizing a user's fingerprint. In one embodiment, the fingerprint recognition button (212) may be capable of a push operation, and may receive a push operation and a fingerprint recognition operation.

[0090] The power button (231) may be a button for turning the power of the display device (100) on / off.

[0091] The home button (232) may be a button for moving to the home screen of the display device (100).

[0092] The live button (233) may be a button for displaying a real-time broadcast program.

[0093] The external input button (234) may be a button for receiving an external input connected to the display device (100).

[0094] The volume control button (235) may be a button for adjusting the size of the volume output by the display device (100).

[0095] The voice recognition button (236) may be a button for receiving a user's voice and recognizing the received voice.

[0096] The channel change button (237) may be a button for receiving a broadcast signal of a specific broadcast channel.

[0097] The confirmation button (238) may be a button for selecting a specific function, and the back button (239) may be a button for returning to the previous screen.

[0098] Let's explain Figure 2 again.

[0099] When the user input interface (230) has a touch screen, the user can input commands related to the display device (100) using the remote control device (200) by touching the soft keys of the touch screen. In addition, the user input interface (230) may have various types of input means that can be operated by the user, such as a scroll key or a jog key, and this embodiment does not limit the scope of the present invention.

[0100] The sensor (240) may include a gyro sensor (241) or an acceleration sensor (243), and the gyro sensor (241) may sense information about the movement of the remote control device (200).

[0101] For example, the gyro sensor (241) can sense information about the operation of the remote control device (200) based on the x, y, and z axes, and the acceleration sensor (243) can sense information about the movement speed of the remote control device (200). Meanwhile, the remote control device (200) can further include a distance measuring sensor, so as to sense the distance to the display (180) of the display device (100).

[0102] The output interface (250) can output a video or audio signal corresponding to an operation of the user input interface (230) or a signal transmitted from the display device (100).

[0103] The user can recognize whether the output interface (250) is manipulating the user input interface (230) or controlling the display device (100).

[0104] For example, the output interface (250) may include an LED (251) that lights up when the user input interface (230) is operated or a signal is transmitted and received with the display device (100) via the wireless communication unit (225), a vibrator (253) that generates vibrations, a speaker (255) that outputs sound, or a display (257) that outputs images.

[0105] In addition, the power supply circuit (260) supplies power to the remote control device (200), and power waste can be reduced by stopping the power supply when the remote control device (200) does not move for a predetermined period of time.

[0106] The power supply circuit (260) can resume power supply when a predetermined key provided in the remote control device (200) is operated.

[0107] The memory (270) can store various types of programs, application data, etc. required for the control or operation of the remote control device (200).

[0108] When the remote control device (200) wirelessly transmits and receives signals through the display device (100) and the RF circuit (221), the remote control device (200) and the display device (100) transmit and receive signals through a predetermined frequency band.

[0109] The controller (280) of the remote control device (200) can store and reference information regarding the frequency band that can wirelessly transmit and receive signals with the display device (100) paired with the remote control device (200) in the memory (270).

[0110] The controller (280) controls all matters related to the control of the remote control device (200). The controller (280) can transmit a signal corresponding to a predetermined key operation of the user input interface (230) or a signal corresponding to the movement of the remote control device (200) sensed by the sensor (240) to the display device (100) via the wireless communication unit (225).

[0111] Additionally, the microphone (290) of the remote control device (200) can acquire voice.

[0112] A plurality of microphones (290) may be provided.

[0113] Next, Figure 4 is described.

[0114] Figure 4 shows an example of utilizing a remote control device according to an embodiment of the present invention.

[0115] Figure 4 (a) illustrates that a pointer (205) corresponding to a remote control device (200) is displayed on a display (180).

[0116] The user can move or rotate the remote control device (200) up and down, left and right. The pointer (205) displayed on the display (180) of the display device (100) corresponds to the movement of the remote control device (200). This remote control device (200) can be called a space remote control because, as shown in the drawing, the pointer (205) moves and is displayed according to the movement in 3D space.

[0117] Figure 4 (b) illustrates that when a user moves the remote control device (200) to the left, the pointer (205) displayed on the display (180) of the display device (100) also moves to the left correspondingly.

[0118] Information about the movement of the remote control device (200) detected through the sensor of the remote control device (200) is transmitted to the display device (100). The display device (100) can calculate the coordinates of the pointer (205) from the information about the movement of the remote control device (200). The display device (100) can display the pointer (205) to correspond to the calculated coordinates.

[0119] Figure 4 (c) illustrates a case where a user moves the remote control device (200) away from the display (180) while pressing a specific button within the remote control device (200). As a result, a selection area within the display (180) corresponding to the pointer (205) can be zoomed in and displayed in an enlarged manner.

[0120] Conversely, when the user moves the remote control device (200) closer to the display (180), the selection area within the display (180) corresponding to the pointer (205) may be zoomed out and displayed in a reduced size.

[0121] Meanwhile, when the remote control device (200) moves away from the display (180), the selection area may be zoomed out, and when the remote control device (200) moves closer to the display (180), the selection area may be zoomed in.

[0122] Additionally, when a specific button within the remote control device (200) is pressed, recognition of up, down, left, and right movements may be excluded. That is, when the remote control device (200) moves away from or toward the display (180), up, down, left, and right movements may not be recognized, and only forward and backward movements may be recognized. When a specific button within the remote control device (200) is not pressed, only the pointer (205) moves in accordance with the up, down, left, and right movements of the remote control device (200).

[0123] Meanwhile, the movement speed or movement direction of the pointer (205) can correspond to the movement speed or movement direction of the remote control device (200).

[0124] Meanwhile, the pointer in this specification refers to an object displayed on the display (180) in response to the operation of the remote control device (200). Accordingly, objects of various shapes other than the arrow shape illustrated in the drawing can be used as the pointer (205). For example, the pointer may be a concept including a point, a cursor, a prompt, a thick outline, etc. In addition, the pointer (205) may be displayed corresponding to one point on the horizontal or vertical axis on the display (180), or may be displayed corresponding to multiple points such as lines or surfaces.

[0125] FIG. 5 illustrates an artificial intelligence (AI) server according to one embodiment of the present disclosure.

[0126] Referring to FIG. 5, the AI ​​server (500) may refer to a device that trains an artificial neural network using a machine learning algorithm or uses a trained artificial neural network.

[0127] The AI ​​server (500) may be composed of multiple servers to perform distributed processing, and may be defined as a 5G network.

[0128] The AI ​​server (500) may be included as part of the configuration of the AI ​​device (100) and may perform at least part of the AI ​​processing together.

[0129] The AI ​​server (500) may include a communication unit (510), a memory (530), a learning processor (540), and a processor (560).

[0130] The communication unit (510) can transmit and receive data with an external device such as a display device (100).

[0131] The memory (530) may include a model storage unit (531).

[0132] The model storage unit (531) can store a model (or artificial neural network, 531a) that is being learned or has been learned through the learning processor (540).

[0133] A learning processor (540) can train an artificial neural network (531a) using learning data. The learning model can be used while mounted on the AI ​​server (500) of the artificial neural network, or can be mounted on an external device such as a display device (100).

[0134] The learning model may be implemented in hardware, software, or a combination of hardware and software. If part or all of the learning model is implemented in software, one or more instructions constituting the learning model may be stored in memory (530).

[0135] The processor (560) can use a learning model to infer a result value for new input data and generate a response or control command based on the inferred result value.

[0136] FIG. 6 is a diagram for explaining the configuration of an AI system according to one embodiment of the present disclosure.

[0137] Referring to FIG. 6, an artificial intelligence (AI) system (60) may include a display device (100), an AI server (500), and a generation AI server (600).

[0138] In one embodiment, the AI ​​server (500) and the generation AI server (600) may be configured as one server.

[0139] The AI ​​server (500) may also be referred to as an artificial intelligence device.

[0140] Components of the AI ​​system (60) can communicate with each other via the Internet.

[0141] Each generation AI server (600) may include the components illustrated in FIG. 5.

[0142] The AI ​​server (500) may be a natural language processing (NLP) server that obtains the results of intent analysis of a voice command through natural language processing.

[0143] The display device (100) may be referred to as an electronic device.

[0144] The AI ​​server (500) can obtain multiple error names corresponding to the content name.

[0145] The AI ​​server (500) can match the acquired multiple error names to content names and store them in the memory (530).

[0146] The display device (100) can obtain a voice command spoken by a user and transmit voice data corresponding to the obtained voice command to the AI ​​server (500) through a network interface (133).

[0147] The AI ​​server (500) can obtain analysis results of voice commands based on voice data.

[0148] The AI ​​server (500) can determine whether the acquired analysis results include a stored error name.

[0149] If the AI ​​server (500) does not include an error name or content name stored in the analysis result, it can transmit a non-recognition result indicating that the content name was not recognized to the display device (100) through the communication unit (510).

[0150] The display device (100) can display the unrecognized result on the display (180).

[0151] If the acquired analysis result includes a stored error name, the AI ​​server (500) can acquire a content name matching the error name from the memory (530).

[0152] The processor (560) of the AI ​​server (500) can obtain search results for the acquired content name and transmit the obtained search results to the display device (100) through the communication unit (510).

[0153] The display device (100) can output search results for content names received from the AI ​​server (500).

[0154] FIG. 7 is a ladder diagram for explaining an operation method of an AI system according to an embodiment of the present disclosure.

[0155] Below, the content name is used as an example, but it is not limited to the content name and can also be applied to words and sentences spoken by the user.

[0156] The processor (560) of the AI ​​server (500) can obtain multiple error names corresponding to the content name (S701).

[0157] The multiple error names may include one or more of a name similar to the content name, a name recognized based on the user's pronunciation error, and an abbreviation of the content name.

[0158] In one embodiment, the processor (560) may receive multiple error names from the generation AI server (600).

[0159] This is explained with reference to Fig. 8.

[0160] FIG. 8 is a diagram illustrating a process of receiving a plurality of error names corresponding to content names from a generation AI server according to one embodiment of the present disclosure.

[0161] The processor (560) can generate a prompt that commands the generation of an error name based on the content name through the communication unit (510) and transmit the generated prompt to the generation AI server (600).

[0162] The prompt may be a command to generate error names that are likely to be misrecognized / mispronounced for the content name.

[0163] The prompt may include an example of a content name and an error name.

[0164] The name included in the prompt may be either the name of a new content or a search term with a search frequency greater than a preset frequency.

[0165] The processor (560) may generate a prompt requesting the generation of error names for the names of new content from a new content database that stores information about the new content.

[0166] The processor (560) may generate a prompt requesting the generation of error names for search terms with a search frequency greater than a certain frequency from a content provider server or a web server.

[0167] The generation AI server (600) can generate multiple error names as a response result to a prompt using a large language model (LLM) (601).

[0168] The super-large language model (601) may be a model that learns a large amount of text data and outputs automatic translation results for prompts, question-answering results, etc.

[0169] The super-large language model (601) can generate multiple error names that can be uttered as content names, given a prompt that commands it to generate error names based on content names.

[0170] The processor (560) can transmit multiple prompts of different forms to the generating AI server (600).

[0171] The generation AI server (600) can generate multiple response results corresponding to each of multiple prompts using the LLM (601). Each of the multiple response results can include multiple error names.

[0172] Figures 9a and 9b are diagrams illustrating the output results of LLM according to a prompt to generate an error name of a content name.

[0173] Referring to Figure 9a, the execution screen of a chatbot application providing a generation AI service is shown.

[0174] The execution screen of the chatbot application can be displayed on the AI ​​server (500) or the display device (100).

[0175] That is, the prompt transmitted to the generation AI server (600) can be generated in the AI ​​server (500) or the display device (100).

[0176] The processor (560) of the AI ​​server (500)<your house> When a first prompt (901) is entered to generate error names with the name , the first prompt (901) can be transmitted to the generation AI server (600).

[0177] The LLM (601) of the generation AI server (600) responds to the first prompt (901).<your house> It can respond with five error names (902) that can be mispronounced.

[0178] The five error names (902) are<Yore howse> ,<Yer house> ,<Yor house> , <youse>and<Your hice> It could be.

[0179] In another embodiment, the processor (560) generates error names (902).<your house> It can be stored in memory (530) by matching the name.

[0180] Next, Figure 9b is described.

[0181] Referring to Figure 9b, another execution screen of a chatbot application providing a generation AI service is shown.

[0182] The processor (560) of the AI ​​server (500)<Explain the principle of shortening "let's go" to "lg"> When a second prompt (911) is entered to explain the principle of generating an abbreviation of the name, the second prompt (911) can be transmitted to the generation AI server (600).

[0183] The LLM (601) of the generation AI server (600) can output a response result (912) indicating the principle of generating an abbreviation of a name in response to the second prompt (911).

[0184] The processor (560) of the AI ​​server (500)<bad boys> If a command to generate an abbreviation for a name is entered in the third prompt (913), the third prompt (913) can be transmitted to the generation AI server (600).

[0185] The LLM (601) of the generation AI server (600) responds to the third prompt (913) with the abbreviation of the name <bb>The response result (914) including can be output.

[0186] In this way, the processor (560) can obtain multiple error names corresponding to the name through the LLM (601) of the generation AI server (600) based on the prompt.

[0187] The display device (100) can also obtain multiple error names corresponding to the name by executing a chatbot application that provides a generation AI service.

[0188] The acquired multiple error names can be stored in the memory (530) of the AI ​​server (500) or in a separate database.

[0189] Again, Figure 7 is explained.

[0190] In another embodiment, the processor (560) may obtain error names for content names from the display device (100).

[0191] The processor (560) may determine that the content name was mispronounced based on the analysis of the voice command uttered by the user.

[0192] Afterwards, the processor (560) can receive a text input for the content name from the display device (100).

[0193] When the processor (560) recognizes the content name through text input, it can obtain the error text of the mispronounced voice as the error name.

[0194] The processor (560) of the AI ​​server (500) can match the acquired multiple error names to content names and store them in the memory (530) (S703).

[0195] The processor (560) can match a plurality of error names corresponding to one acquired content name to the content name and store them in the memory (530) or database.

[0196] The database may be included in the AI ​​server (500) or may be a storage facility provided separately from the AI ​​server (500).

[0197] FIG. 10 is a diagram illustrating a table matching content names and multiple error names according to one embodiment of the present disclosure.

[0198] Referring to Fig. 10,<your house> The content name and multiple error names matching the content name<Yore howse> ,<Your hice> and <youse>A table (1000) showing the correspondence relationship between the livers is shown.

[0199] The table (1000) can be stored in the memory (530) of the AI ​​server (500) or in a content database (not shown).

[0200] Multiple tables can be stored in the form of a table (1000) in which multiple names and multiple error names corresponding to each name are matched.

[0201] The name can be either a content name or a search term.

[0202] In one embodiment, if the content name includes one or more keywords, variant words of each of the one or more keywords may be used to generate the error name.

[0203] LLM (601) can generate variant words for each keyword when the content name contains one or more keywords, and generate an error name by combining the generated variant words.

[0204] For example, assume that the content title consists of two words, including a first keyword and a second keyword.

[0205] LLM (601) can generate variant words of the first keyword and variant words of the second keyword.

[0206] LLM (601) can generate an error name by combining any one of the variant words of the first keyword and any one of the variant words of the second keyword.

[0207] A plurality of error names generated by a combination of variant words of LLM (601) can be transmitted to the AI ​​server (500).

[0208] In another embodiment, the processor (560) of the AI ​​server (500) may generate variant words of the first keyword and variant words of the second keyword. The processor (560) may generate an error name by combining any one of the variant words of the first keyword with any one of the variant words of the second keyword.

[0209] The processor (560) of the AI ​​server (500) may adjust the number of error names based on the popularity ranking of the content (in the case of search terms, the search ranking). For example, the processor (560) may increase the number of error names corresponding to the content name as the popularity ranking of the content increases, and may decrease the number of error names corresponding to the content name as the popularity ranking of the content decreases.

[0210] That is, the processor (560) can add or delete error names to adjust the capacity of the memory (530).

[0211] If the number of error names corresponding to content names with high popularity rankings is increased, recognition failure due to user mispronunciation can be significantly reduced.

[0212] When the number of error names corresponding to content names with low popularity rankings is reduced, there is an effect of reducing the capacity of the memory (530) or content database.

[0213] The AI ​​server (500) can receive the popularity ranking of content or the search ranking of search words from a content provider server or a search server.

[0214] Again, Figure 7 is explained.

[0215] The controller (170) of the display device (100) can obtain a voice command spoken by a user (S705) and transmit voice data corresponding to the obtained voice command to the AI ​​server (500) via a network interface (133) (S707).

[0216] In one embodiment, the AI ​​server (500) may be equipped with an STT (Speech To Text) engine and may convert voice data into text data using the STT engine.

[0217] In another embodiment, the display device (100) may transmit voice data corresponding to a voice command to an STT server (not shown). The STT server may convert the voice data into text data and transmit the converted text data to an AI server (500).

[0218] The processor (560) of the AI ​​server (500) can obtain the analysis result of the voice command based on the voice data (S709).

[0219] The processor (560) can obtain analysis results using text data of voice data.

[0220] The processor (560) can obtain analysis results from text data using a natural language processing engine.

[0221] The processor (560) can sequentially perform a morphological analysis step, a syntax analysis step, a speech act analysis step, and a dialogue processing step on text data to generate an analysis result.

[0222] The morphological analysis step is the step of classifying text data corresponding to the user's spoken voice into morphemes, which are the smallest units that have meaning, and determining which part of speech each classified morpheme has.

[0223] The syntactic analysis step is the step that uses the results of the morphological analysis step to divide text data into noun phrases, verb phrases, adjective phrases, etc., and determines what kind of relationship exists between each divided phrase.

[0224] Through the parsing step, the subject, object, and modifiers of the speech spoken by the user can be determined.

[0225] The speech act analysis step uses the results of the syntactic analysis step to analyze the intent of the user's speech. Specifically, the speech act analysis step determines the intent of the sentence, such as whether the user is asking a question, making a request, or simply expressing emotion.

[0226] The conversation processing stage uses the results of the speech act analysis stage to determine whether to respond to the user's utterance, respond, or ask a question for additional information.

[0227] After the conversation processing step, the processor (560) can generate an analysis result including one or more of the user's spoken intention, a response to the intention, a response, and an inquiry for additional information.

[0228] The processor (560) of the AI ​​server (500) can determine whether the acquired analysis result includes a stored error name (S711).

[0229] If the analysis result is a search intent for content corresponding to the content name, the processor (560) can determine whether the analysis result includes an error name through the memory (530) or the content database.

[0230] If the stored error name or content name is not included in the analysis result, the processor (560) of the AI ​​server (500) can transmit a non-recognition result indicating that the content name was not recognized to the display device (100) through the communication unit (510) (S712).

[0231] The display device (100) can display the unrecognized result on the display (180).

[0232] If the acquired analysis result includes a stored error name, the processor (560) of the AI ​​server (500) can acquire a content name matching the error name from the memory (530) (S713).

[0233] If the text data of a voice command spoken by a user for a content name corresponds to an error name, the processor (560) can recognize that the user has spoken a content name that matches the error name.

[0234] The processor (560) of the AI ​​server (500) can obtain search results for the acquired content name (S715) and transmit the obtained search results to the display device (100) via the communication unit (510) (S717).

[0235] The processor (560) can transmit a search request for a content name recognized as spoken by a user to a search server (not shown) and receive a search result for the content name from the search server.

[0236] The search server may be a web server or a content provider server. Search results may include one or more of the following: characters, plot, episode information, or access address.

[0237] The controller (170) of the display device (100) can output search results for content names received from the AI ​​server (500) (S719).

[0238] The controller (170) can display the search results on the display (180).

[0239] In this way, according to an embodiment of the present disclosure, by proactively addressing the possibility of content being incorrectly searched, search failures due to frequent recognition failures caused by errors in voice recognition according to STT can be reduced.

[0240] Accordingly, the user experience for voice search can be improved by reducing search failures when the user speaks only some of the key words of the content name or replaces some of the key words with synonyms.

[0241] Figure 11 is a diagram illustrating a process for providing search results for desired content even when a user incorrectly pronounces the content name.

[0242] In Figure 11, the content name intended by the user is<your house> and the user is<your hice> Assume that the utterance was uttered.

[0243] The display device (100) is used to<your hice> Voice data for voice commands can be transmitted to the AI ​​server (500).

[0244] The display device (100) can receive voice commands from a remote control device (200) or through a microphone provided therein.

[0245] The AI ​​server (500) can obtain analysis results by analyzing the intent of text data for voice data.

[0246] The analysis results may indicate a search intent for content corresponding to the content name. The search intent may include a first text corresponding to the content name and a first text indicating a search request for content corresponding to the content name.

[0247] The AI ​​server (500) can determine whether the first text included in the analysis result matches the error names stored in the content database (1100).

[0248] That is, the AI ​​server (500) can determine whether the first text included in the analysis result is stored in the table (1100) stored in the content DB (1100).

[0249] The AI ​​server (500) can extract a content name matching the first text when the first text included in the analysis result is stored in the table (1100).

[0250] The AI ​​server (500) can determine that the content name has been recognized if the first text included in the analysis result is stored in the table (1100).

[0251] The AI ​​server (500) can provide search results of content corresponding to the recognized content name to the display device (100).

[0252] In this way, according to an embodiment of the present disclosure, by proactively addressing the possibility of content being incorrectly searched, search failures due to frequent recognition failures caused by errors in voice recognition according to STT can be reduced.

[0253] Accordingly, the user experience for voice search can be improved by reducing search failures when the user speaks only some of the key words of the content name or replaces some of the key words with synonyms.

[0254] Figure 12 is a diagram illustrating an example of providing a pop-up window to confirm whether the content name intended by the user is correct when the user incorrectly pronounces the content name.

[0255] In Figure 12, the content name intended by the user is<your house> and the user is<your hice> Assume that the utterance was made.

[0256] The AI ​​server (500) recognizes the content of a voice command mispronounced by the user, as in the embodiment of FIG. 11.<your house> can be judged by

[0257] The AI ​​server (500) can transmit a control command to the display device (100) to check whether the content recognition result is correct.

[0258] The display device (100) can display a pop-up window (1200) on the display (180) to confirm whether the content name intended by the user is correct according to a control command received from the AI ​​server (500).

[0259] The pop-up window (1200) may include the recognized content name and text asking if the content name is correct.

[0260] When the display device (100) receives a confirmation input indicating that the content name is correct through a pop-up window (1200), it can transmit a confirmation command indicating that the confirmation input has been received to the AI ​​server (500).

[0261] The AI ​​server (500) can perform a search for a content name according to a received confirmation command and transmit the search result to the display device (100).

[0262] The display of the pop-up window (1200) of FIG. 12 can be performed between steps S713 and S715 of FIG. 7.

[0263] When the user's confirmation process is performed through a pop-up window (1200) as in the embodiment of Fig. 12, it can be more accurately determined whether the content name intended by the user is correct.

[0264] According to one embodiment of the present invention, the above-described method can be implemented as processor-readable code on a medium in which a program is recorded. Examples of processor-readable media include ROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage devices.< / youse> < / bb> < / youse>

Claims

1. In artificial intelligence devices, A memory for storing a name and multiple error names matching the name; A communication unit that communicates with an electronic device or a generated AI (Artificial Intelligence) server; and A processor that receives voice data corresponding to a voice command spoken by a user from the electronic device, obtains an analysis result based on the received voice data, and if any one of the plurality of stored error names is included in the obtained analysis result, obtains the name matching the plurality of error names, and transmits a search result for the obtained name to the electronic device. Artificial intelligence devices.

2. In paragraph 1, The above name is Either the content name or the search term Artificial intelligence devices.

3. In paragraph 2, The above processor Sending a prompt to the generation AI server requesting the generation of an error name for the name of a new content obtained from a new content database, Receiving error names for the names of the new content from the above-mentioned generation AI server Artificial intelligence devices.

4. In paragraph 2, The above processor A prompt is sent to the generation AI server requesting the generation of an error name for the search term whose search frequency is greater than a certain frequency, Receiving error names for the search term from the above-mentioned generation AI server Artificial intelligence devices.

5. In paragraph 1, The above processor A command is transmitted to the electronic device to cause the electronic device to output a pop-up window to confirm whether the acquired name matches the user's intention. Artificial intelligence devices.

6. In paragraph 2, The above processor The number of stored error names is adjusted according to the popularity ranking of the content corresponding to the above content name or the search ranking of the above search term. Artificial intelligence devices.

7. In paragraph 6, The above processor The lower the popularity ranking or search ranking, the lower the number of error names. Artificial intelligence devices.

8. In paragraph 1, If the above name contains one or more keywords, each of the plurality of error names is a combination of variant words of each keyword. Artificial intelligence devices.

9. In the method of operating an artificial intelligence device, A step of storing a name and a plurality of error names matching the name; A step of receiving voice data corresponding to a voice command spoken by a user from an electronic device; A step of obtaining analysis results based on received voice data; If the acquired analysis result includes any one of the plurality of stored error names, a step of acquiring the name matching the plurality of error names; and comprising a step of transmitting the search result for the acquired name to the electronic device; How artificial intelligence devices work.

10. In paragraph 9, The above name is Either the content name or the search term How artificial intelligence devices work.

11. In paragraph 10, A step of generating a prompt requesting the generation of an error name for the name of a new content acquired from a new content database and transmitting the prompt to the AI ​​server; and Further comprising a step of receiving error names for the names of the new content from the above generation AI server. How artificial intelligence devices work.

12. In paragraph 10, A step of generating a prompt requesting the generation of an error name for the search term having a search frequency greater than a certain frequency and transmitting the same to the AI ​​server; and Further comprising a step of receiving error names for the search term from the above generating AI server. How artificial intelligence devices work.

13. In paragraph 9, Further comprising a step of transmitting a command to the electronic device to cause the electronic device to output a pop-up window for confirming whether the acquired name matches the user's intention. How artificial intelligence devices work.

14. In paragraph 10, Further comprising a step of adjusting the number of stored error names according to the popularity ranking of the content corresponding to the content name or the search ranking of the search word. How artificial intelligence devices work.

15. In paragraph 14, The above adjustment steps are A step of reducing the number of error names as the popularity ranking or search ranking is lowered. How artificial intelligence devices work.

Citation Information

Patent Citations

  • Method and apparatus for training language model, method and apparatus for recognizing speech

    KR1020160069329A

  • Apparatus and method for speech recognition

    KR1020170063037A

  • Height-Adjustable Pillow Having User-Customized Variable Structure

    KR1020230140208A

  • Open flame grilling plate

    KR102551931B1

  • Speech recognition error diagnosis

    US20160253989A1