Artificial intelligence-based multimodal intelligent kiosk device capable of voice recognition, and method for driving device
The multimodal intelligent kiosk device with AI-based voice recognition optimizes user interaction by integrating touch and voice interfaces, using sensors and real-time text display to enhance usability and accuracy in order processing.
Patent Information
- Application Number
- PCT/KR2024/008733
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-24
- Filing Date
- 2024-06-24
- Publication Date
- 2026-01-02
AI Technical Summary
Existing voice-enabled kiosks face user confusion due to separate operation of touch and voice recognition interfaces, leading to frustration and inefficiency in using multiple modal functions.
A multimodal intelligent kiosk device with AI-based voice recognition that integrates touch screen and voice guidance, using sensors to detect user approach and real-time text display of voice recognition progress, optimizing service scenarios and algorithms to enhance usability.
The integration of touch and voice recognition through AI enhances user experience by reducing confusion and ensuring accurate order processing, allowing seamless interaction with the kiosk.
Smart Images

Figure KR2024008733_02012026_PF_FP_ABST
Abstract
Description
Multimodal intelligent kiosk device capable of artificial intelligence-based voice recognition and its operating method
[0001] The present invention relates to a multimodal intelligent kiosk device capable of artificial intelligence-based voice recognition and a method for operating the device, and more particularly, when a kiosk capable of voice recognition, etc., has multiple multimodal functions and operates each function individually, there is a concern that a user may be confused in using the kiosk, so the present invention relates to a multimodal intelligent kiosk device capable of artificial intelligence-based voice recognition and a method for operating the device, which prevents this and optimizes the usability of the actual user experience (UX) / user interface (UI).
[0002] Traditional kiosks have been primarily used in the public sector for guidance purposes, displaying menus on a screen and providing information about user choices. They are also widely used in the private sector, allowing customers to order and pay for food at restaurants and other establishments. Recently, with the rapid increase in contactless services that minimize human contact, demand for kiosks has grown even more. Furthermore, the growing aversion to interacting with kiosks has led to a significant increase in demand for voice-activated kiosks.
[0003] However, existing voice-enabled kiosks have the problem of users feeling frustrated due to the long time it takes for voice recognition, and in particular, if the touch user interface (UI) and the voice recognition UI operate separately, there is a concern that users may become confused when using the kiosk.
[0004] The problem to be solved by the present invention is that, when a kiosk capable of voice recognition, etc., has multiple multi-modal functions and operates each function individually, there is a concern that the user may be confused when using the kiosk, so the purpose of the present invention is to provide a multi-modal intelligent kiosk device capable of voice recognition based on artificial intelligence and a method of operating the device, which prevents this and optimizes the usability of the actual user experience (UX) / user interface (UI).
[0005] The tasks of the present invention are not limited to the tasks mentioned above, and other tasks not mentioned will be clearly understood by those skilled in the art from the description below.
[0006] In order to solve the above problem, a multimodal intelligent kiosk device capable of artificial intelligence-based voice recognition according to an embodiment of the present invention includes a user interface unit including a touch screen panel capable of touching a screen and a voice receiving unit for receiving voice, and a control unit for displaying the voice content of a voice received through the voice receiving unit as text on the screen of the screen panel to guide a user to recognize that voice recognition is being performed when using the kiosk, and for displaying the text in the form of a string on the screen according to the voice content and speed.
[0007] The above intelligent kiosk device further includes a sensor unit that detects whether a user using the kiosk is approaching, and the control unit, when it is determined that the user has first approached the kiosk based on the sensing data of the sensor unit, displays a standby screen showing in words that the kiosk is a kiosk for ordering by voice on the touch screen panel, and when it is determined that the user's approach is for ordering, starts screen guidance and voice guidance, and the control unit verifies in real time whether the user has properly ordered by expressing in text the names of the menus and options pronounced by the user, and the contents of the order, payment, and memos (e.g., expressing the results of STT (Speech to text) on the screen in a screen string), and can recognize the text expression when ordering again.
[0008] The above control unit can display an order page for ordering a menu and an option page for selecting options of the order menu on the screen of the touch screen panel when proceeding with screen guidance according to a preset service scenario, and can simultaneously proceed with preset voice guidance when proceeding with the screen guidance.
[0009] The above control unit performs different voice guidance when a required option is selected or not selected on the option page when there is no user selection for a specified period of time on the option page, and performs different voice guidance when there is an item in the shopping cart or when there is no item in the shopping cart when there is no user selection for a specified period of time on the payment step, and expresses the option name pronounced by the user in a string (below) so that it can be checked in real time whether the order has been made correctly.
[0010] A method for driving a multimodal intelligent kiosk device capable of artificial intelligence-based voice recognition according to an embodiment of the present invention for solving the above problem includes a step in which a control unit displays the voice content of a voice received through a voice receiving unit as text on a screen of a screen panel to guide a user to recognize that voice recognition is being performed when using a kiosk, and a step in which the control unit displays the text on the screen in the form of a string according to the voice content and speed.
[0011] The above driving method may further include a step of a sensor unit detecting whether a user using the kiosk is approaching, a step of the control unit displaying a standby screen showing in words that the kiosk is a voice ordering kiosk on the screen panel when it is determined that the user has approached the kiosk for the first time based on the sensing data of the sensor unit, a step of starting screen guidance and voice guidance when it is determined that the user's approach is for ordering, and a step of checking in real time whether the user has correctly ordered by expressing in text the names of the menus and options pronounced by the user, and the contents of the order, payment, and memo, and recognizing the text expression when ordering again.
[0012] The step of starting the above screen guidance and voice guidance is to display an order page for ordering a menu and an option page for selecting options of the order menu on the screen of the touch screen panel when the screen guidance is performed according to a preset service scenario, and to simultaneously perform preset voice guidance when the screen guidance is performed.
[0013] The step of starting the above screen guidance and voice guidance may include a step of performing different voice guidance by distinguishing between cases where a required option is selected and cases where a required option is not selected on the option page when there is no user selection for a specified period of time on the option page, a step of performing different voice guidance by distinguishing between cases where there are items in the shopping cart and cases where there are no items in the shopping cart when there is no user selection for a specified period of time on the payment step, and a step of expressing the option name pronounced by the user in a string to check in real time whether the order was made correctly.
[0014] According to an embodiment of the present invention, for example, by enabling recognition of screen touch and voice recognition through an artificial intelligence (AI) voice recognition kiosk, and optimizing service scenarios and algorithms so that users can naturally select and use menus at the kiosk, confusion in the use of the kiosk can be eliminated.
[0015] The effects of the present invention are not limited to the effects mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art from the description of the claims.
[0016] FIG. 1 is a drawing illustrating an example of a kiosk service system according to an embodiment of the present invention.
[0017] Figure 2 is a block diagram illustrating the detailed structure of the intelligent kiosk shown in Figure 1.
[0018] FIGS. 3 to 5 are drawings for explaining the integrated recognition optimization function of the intelligent kiosk device of FIG. 1.
[0019] FIG. 6 is a flowchart showing the driving process of the intelligent kiosk device of FIG. 1 according to an embodiment of the present invention.
[0020] The present invention is not limited to the embodiments described below, but can be implemented in various different forms. These embodiments are merely illustrative of the contents of the present invention and are provided to provide those skilled in the art with a detailed understanding of the scope of the invention. The present invention is defined solely by the scope of the claims. Like reference numerals refer to like elements throughout the specification.
[0021] Embodiments described herein will be described with reference to cross-sectional and / or plan views, which are ideal examples of the present invention. In the drawings, the illustrated regions are expressed for the effective explanation of the technical contents. Therefore, the regions illustrated in the drawings have a schematic nature, and the shapes of the regions illustrated in the drawings are intended to illustrate specific forms of the device regions and are not intended to limit the scope of the invention. Although terms such as first, second, and third are used to describe various components in various embodiments of the present specification, these components should not be limited by such terms. These terms are used only to distinguish one component from another. The embodiments described and illustrated herein also include complementary embodiments thereof.
[0022] The terminology used herein is for the purpose of describing embodiments only and is not intended to be limiting of the present invention. In this specification, the singular also includes the plural unless specifically stated otherwise. As used herein, the terms "comprises" and / or "comprising" do not exclude the presence or addition of one or more other components, steps, operations, and / or elements to the mentioned components, steps, operations, and / or elements.
[0023] Unless otherwise defined, all terms (including technical and scientific terms) used herein may be used in their common sense to those of ordinary skill in the art to which the present invention pertains. Furthermore, terms defined in commonly used dictionaries are not to be interpreted ideally or excessively unless explicitly and specifically defined otherwise.
[0024] Hereinafter, with reference to the drawings, the concept of the present invention and embodiments thereof will be described in detail.
[0025] FIG. 1 is a drawing illustrating an example of a kiosk service system according to an embodiment of the present invention.
[0026] As illustrated in FIG. 1, a kiosk service system (90) according to an embodiment of the present invention includes part or all of an intelligent kiosk (device) (100), a communication network (110), and a kiosk integrated service device (120).
[0027] Here, “including some or all” means that the communication network (110) or the kiosk integrated service device (120) of FIG. 1 is omitted so that the intelligent kiosk (100) can operate independently in a stand-alone form, or the intelligent kiosk (100) and the kiosk integrated service device (120) perform direct (e.g., P2P, etc.) communication, or some or all of the components constituting the kiosk integrated service device (120) can be integrated into a network device (e.g., a wireless switching device, a gateway, etc.) constituting the communication network (110), and in order to help a sufficient understanding of the invention, it is described as including all.
[0028] The intelligent kiosk (100) can, for example, be in a stand-alone form, and then execute a program for performing operations according to an embodiment of the present invention, and can process data such as service screens in a streaming form by linking with the kiosk integrated service device (120). For example, the intelligent kiosk (100) can analyze the voice signal and convert it into text through a text conversion module installed inside when converting a voice signal into text, but such an operation can also be requested from the kiosk integrated service device (120). This is because the kiosk integrated service device (120) is a more expensive device and can thus possess high-performance hardware and software resources. Of course, the intelligent kiosk (100) according to the embodiment of the present invention can also be equipped with only the minimum hardware and software resources for performing operations according to the embodiment of the present invention, and all data can be processed by the kiosk integrated service device (120) in a cloud form.
[0029] The intelligent kiosk (100) can be installed in general stores such as government offices or restaurants to help users select desired documents, order menus, and make payments. Of course, the intelligent kiosk (100) can also be used for electronic voting, etc., so it can be installed and used in various places. It can also be used on golf courses, etc. Above all, the intelligent kiosk (100) according to an embodiment of the present invention can help users effectively use the voice recognition function by applying an artificial intelligence (AI) program to guide users to recognize that voice recognition is being used when using the kiosk through voice recognition. In addition, when an AI robot provides guidance through TTS (Text to Speech) or when a customer speaks an order, the voice content can be displayed on the screen as text to emphasize the feeling of an actual conversation. For example, when a user speaks "Eat and Go" or "Take it home" according to the voice guidance, the words "Eat and Go" or "Take it home" can be displayed as text next to the voice microphone icon on the screen. This is also well illustrated in (d) of FIG. 4.
[0030] An intelligent kiosk (100) according to an embodiment of the present invention may be a voice recognition kiosk using intelligent AI. Although it primarily uses voice recognition, it may have three or more multi-modal functions, including a UI for touching the screen, a UI for voice recognition, and a UI for detecting a person's approach. For example, in the case where each function is operated individually, the usability of the UI and UX actually used is optimized to prevent confusion among users when using the kiosk. Of course, the key to the success of an AI voice recognition kiosk is to design and optimize service scenarios and algorithms so that users can use the kiosk naturally without difficulty. The optimized scenarios and algorithms are installed and operated.
[0031] As will be discussed in detail later, the intelligent kiosk (100) has a user interface operation through a panel that is a combination of a touch screen, that is, a video panel and a sensing panel, so that a user command can be input by the user's touch on the screen, and in the case of voice recognition, a microphone that receives voice may be installed to receive the user's spoken voice or voice content, and operate based on the voice command or voice content. The spoken voice may be a command to wake up a kiosk device or an AI program to use a service. Of course, the embodiment of the present invention will not be particularly limited to such spoken commands. Above all, in the embodiment of the present invention, the presence or absence of a user or the movement of the user may be detected through a sensor unit such as a detection sensor (e.g., a proximity sensor such as an infrared sensor or a camera, etc.) installed in the intelligent kiosk (100), and screen guidance and voice guidance may be started based on this. In addition, the intelligent kiosk (100) can check in real time whether the order was made correctly by expressing the menu and option names pronounced by the user, as well as the order and payment, memo contents, etc., in a screen string (STT: Speech to text result displayed on the screen), and can also perform an action to recognize the text expression when ordering again.
[0032] The intelligent kiosk (100) according to the embodiment of the present invention may not operate with a single service scenario or algorithm set when optimizing the service scenario and algorithm, that is, when the program is designed to have multi-modal functions such as a touch UI, a voice recognition UI, and a user approach detection UI. In other words, this is because the intelligent kiosk (100) used in stores such as restaurants or public offices may have various different service scenarios. Therefore, when the intelligent kiosk (100) is installed in a specific location such as a public office or a restaurant, the initial setting is to provide screen guidance and voice guidance with a service scenario optimized for that location. However, such screen guidance and voice guidance may be changed at any time based on the analysis results of collected data related to the surrounding environment in which the intelligent kiosk (100) is installed or the usage patterns of users.
[0033] For example, a representative example could be a behavior that distinguishes whether a user approaches the intelligent kiosk (100) to use it or simply passes by. For example, in crowded places like restaurants, a voice-activated kiosk function might feel a bit awkward. For example, it would be perfectly possible to automatically switch between on-screen and voice guidance when orders are continuously generated over a certain period of time, using proximity sensors or the like. It would also be perfectly possible to check the number of people waiting using a camera. For example, voice recognition could replace on-screen guidance in the form of text. In this way, the intelligent kiosk (100) can analyze the surrounding environment through its built-in artificial intelligence (AI) program and optimize service scenarios based on the analysis results. Of course, these actions could also be temporarily activated based on time information provided by users. It would also be perfectly possible to simplify the menu, leaving only frequently used menus in a specific store, and automatically reconfigure the screen layout, or menu composition. In addition, for menus selected a preset number of times, titles such as Best or Popular Menu can be automatically assigned, i.e., displayed on one side of the order menu. Of course, such service operations can also be performed at the request of the kiosk integrated service device (120) of Fig. 1. Of course, in the case of Best or Popular Menu, voice guidance can also be provided when a user selects a menu.
[0034] Looking into it in more detail, the intelligent kiosk (100) according to an embodiment of the present invention detects a user's approach and distinguishes whether the user is approaching to use the kiosk or is simply passing by. Then, when it is determined that the user is approaching to use the kiosk, it provides guidance through voice and screen. When the proximity sensor first operates, it only displays a standby screen. That is, only phrases such as "This is a kiosk that allows you to order by voice" are displayed on the screen. This is well illustrated in Fig. 3(b) and Fig. 4(b). Then, if the proximity sensor continues to detect a user for a certain period of time, for example, 0.2 to 0.8 seconds (after a period of time longer than the time it takes for a person to pass by the kiosk), it determines that the user is approaching the kiosk to order rather than passing by, and can initiate on-screen and voice guidance. This operation can be set as the default operation. In addition, the intelligent kiosk (100) displays the menu and option names pronounced by the user, as well as the order and payment and memo contents, in characters on the screen string to check in real time whether the order has been made correctly, and can recognize the expressed characters when ordering again.
[0035] The communication network (110) can be configured in various forms. The communication network (110) can include both wired and wireless communication networks. For example, a wired or wireless Internet network can be used or linked as the communication network (110). Here, the wired network includes an Internet network such as a cable network or a public switched telephone network (PSTN), and the wireless communication network includes a CDMA, WCDMA, GSM, EPC (Evolved Packet Core), LTE (Long Term Evolution), Wibro network, etc. Of course, the communication network (110) according to the embodiment of the present invention is not limited thereto, and can be used as an access network of a next-generation mobile communication system to be implemented in the future, for example, a cloud computing network in a cloud computing environment, a 5G network, a 6G network, etc. For example, if the communication network (110) is a wired communication network, an access point within the communication network can connect to a telephone exchange, etc., but if it is a wireless communication network, data can be processed by connecting to an SGSN or GGSN (Gateway GPRS Support Node) operated by a communication company, or data can be processed by connecting to various relays such as a BTS (Base Transceiver Station), NodeB, or e-NodeB.
[0036] The communication network (110) may include an access point. The access point may include a small base station such as a femto or pico base station, which is widely installed inside a building. Here, the femto or pico base station may be classified according to the maximum number of intelligent kiosks (100) of FIG. 1 that can be connected to the small base station. Of course, the communication network (110) may include a short-range communication module for performing short-range communication such as Zigbee and Wi-Fi with the intelligent kiosk (100). The access point may use TCP / IP or RTSP (Real-Time Streaming Protocol) for wireless communication. Here, short-range communication may be performed in various standards such as Bluetooth, Zigbee, infrared (IrDA), radio frequency (RF) such as ultra-high frequency (UHF) and very high frequency (VHF), and ultra-wideband (UWB) in addition to Wi-Fi. Accordingly, the access point can extract the location of the data packet, designate the best communication path for the extracted location, and forward the data packet along the designated communication path to the next device, such as the kiosk integrated service device (120). The access point can share multiple lines in a typical network environment, and may include, for example, a router, a repeater, and a repeater.
[0037] The kiosk integrated service device (120) may include, for example, a cloud server, and may be configured to include a DB (120a) linked to the server. Here, the server and the DB (120a) may be connected and operated via a dedicated network such as an intranet, and since they communicate via the internal network, separate data processing such as modulation / demodulation, encoding / decoding, etc. may not be performed, unlike when using an external network such as a communication network (110). However, operations such as encryption or decryption may be performed to safely protect personal information, etc. The kiosk integrated service device (120) may operate as an integrated platform to provide a service according to an embodiment of the present invention. However, since various types of kiosks are actually released and used on the market, it may operate as a monitoring device that integrates and monitors only kiosks of a specific product.
[0038] The kiosk integrated service device (120) may also perform an operation to update a program according to an embodiment of the present invention installed within the intelligent kiosk (100). For example, when a function is added, an update operation is performed to automatically add the function. The kiosk integrated service device (120) stores information in the database (120a) regarding which programs are installed and which functions are executed by each intelligent kiosk (100), and thus can select a specific target and perform an update operation based on this information.
[0039] Above all, the kiosk integrated service device (120) according to an embodiment of the present invention can collect operation data of intelligent kiosks (100) installed and operating in various locations across the country or data related to usage patterns of how users use the kiosks, for example, by generating big data and analyzing it through an artificial intelligence program to predict kiosk failures, or can perform optimization operations by analyzing user patterns of users of intelligent kiosks (100) installed in a specific area and performing corresponding screen guidance and voice guidance operations. For example, the intelligent kiosk (100) of a store analyzed as a good restaurant can reset its internal program to perform corresponding operations. In addition, it may be possible to control the intelligent kiosk (100) to add titles such as Best or Popular Menu for products or popular menus that are popular nationwide.
[0040] In addition to the above, the intelligent kiosk (100), communication network (110), and kiosk integrated service device (120) of FIG. 1 can perform various operations, and other detailed information will be covered in the following, so we will replace it with those contents.
[0041] FIG. 2 is a block diagram illustrating a detailed structure of the intelligent kiosk shown in FIG. 1, and FIGS. 3 to 5 are drawings for explaining the integrated recognition optimization function of the intelligent kiosk device of FIG. 1.
[0042] As illustrated in FIG. 2, the intelligent kiosk (device) (100) of FIG. 1 according to an embodiment of the present invention includes part or all of a sensor unit (200), a user interface unit (210), a control unit (220), a usability optimization unit (230), and a storage unit (240), and may further include a communication interface unit for communicating with the kiosk integrated service device (120) of FIG. 1.
[0043] Here, “including some or all” means that some components, such as the storage unit (240), may be omitted to configure the intelligent kiosk (100) of FIG. 1, or some components, such as the usability optimization unit (230), may be integrated into other components, such as the control unit (220), etc. In order to help a sufficient understanding of the invention, it is described as including all.
[0044] The sensor unit (200) includes a proximity sensor such as an infrared sensor or a distance sensor. Of course, the sensor unit (200) may include various types of sensors other than the proximity sensor to detect whether a user is approaching, and furthermore, it may be replaced with a camera sensor module or configured to include a camera module. The sensor unit (200) is used to distinguish whether a user is approaching to use the intelligent kiosk (100) or is just passing by. In practice, the sensor unit (200) can generate sensing data and provide it to the control unit (220), and the data is provided to the usability optimization unit (230) to perform an operation of providing guidance through voice and screen when it is determined through data analysis that a user is approaching to use the intelligent kiosk (100).
[0045] In the case of the infrared sensor constituting the sensor unit (200), it may be configured to include a light emitting unit (or light emitting sensor) and a light receiving unit (or light receiving sensor), and infrared rays may be emitted through the light emitting unit, and light emitted by, for example, a user, may be received through the light receiving unit. Accordingly, data sensed through the light receiving sensor of the light receiving unit may be generated and provided to the control unit (220). In addition, in the case of the distance sensor, a laser sensor, etc. may be used. A laser may be emitted through the laser sensor, and a laser reflected by the user may be received to measure the presence or distance of the user, and the proximity of the user may be determined through the measured distance. Of course, it can be considered that the corresponding operation is performed by the usability optimization unit (230).
[0046] The user interface unit (210) may include a touch screen and a voice receiving unit. The touch screen may include an image display panel that displays an image, and a sensing element or a sensing panel formed inside or outside the image display panel. Of course, since the configuration of such a touch screen is already known, further description will be omitted. The voice receiving unit may include a voice receiving module such as a microphone and may receive a user's voice. The user interface unit (210) may further include a voice output unit. The voice output unit may include a speaker, and voice guidance for using the kiosk may be provided through the voice output unit. The screen displayed on the touch screen is well illustrated in FIGS. 3 to 5 .
[0047] The control unit (220) includes a processor such as a CPU, an MPU, a GPU, and may further include a flash memory such as RAM. The processor and the memory may be configured as a single chip in the form of an IC chip, etc. The control unit (220) performs the overall control operation of the sensor unit (200), the user interface unit (210), the usability optimization unit (230), and the storage unit (240) of FIG. 2. When sensing data is provided through the sensor unit (200), the control unit (220) may temporarily store the data in the storage unit (240) and then retrieve it to request data analysis from the usability optimization unit (230). In addition, when the usability optimization unit (230) determines that a user has approached to use the intelligent kiosk (100) based on the data analysis results, the control unit (220) may execute an internal program of the usability optimization unit (230) to display the data on the touch screen constituting the user interface unit (210). Of course, the control unit (220) can be involved in this process and operate. The control unit (220) can display a designated screen on the touch screen according to the request of the usability optimization unit (230).
[0048] In addition, the control unit (220) can perform various operations in conjunction with the usability optimization unit (230). This is to perform usability optimization operations. Usability optimization may mean that in the case of a voice recognition kiosk in the past, for example, when the on-screen touch UI, voice recognition UI, and human approach UI were operated individually when it had a multi-modal function, which caused a lot of confusion for users when using the kiosk. In the embodiment of the present invention, in order to eliminate such confusion, the service scenario and algorithm are optimized and the kiosk service is provided to users accordingly. Of course, the control unit (220) may execute a program installed in the usability optimization unit (230) to execute a program for the optimization.
[0049] The usability optimization unit (230) optimizes the usability of UX and UI used by actual users for multi-modal functions such as touch UI on a touch screen, voice recognition UI that recognizes voice, and UI that detects human approach, and loads the optimized program and executes the program at the request of the control unit (220). Of course, data according to the execution of the program may be provided by generating a service screen by a graphic (GUI) program.
[0050] The usability optimization unit (230) according to an embodiment of the present invention can be optimized to perform operations as shown in FIG. 3 with respect to a UI that detects a person's approach. The usability optimization unit (230) analyzes the sensing data provided by the sensor unit (200) and, based on the analysis results, determines whether the user is approaching the intelligent kiosk (100) to use it or is simply passing by. When it is determined that the user is approaching the kiosk to use it, it can provide guidance on how to use the kiosk using voice and screen. As shown in FIG. 3, when the proximity sensor is first activated, only a standby screen may be displayed. FIG. 3 (a) and (b) show the standby screen. That is, only a phrase such as "This is a kiosk that accepts voice orders" may be displayed. In addition, when the proximity sensor continues to detect a user for a period longer than the time required for a person to pass by the kiosk, the usability optimization unit determines that the user is approaching the kiosk to order rather than passing by, and initiates on-screen and voice guidance. The time it takes for a person to pass through a kiosk can be used to distinguish between ordinary people and those with special needs, such as infants or the elderly. While the average person typically passes within 0.2 to 0.8 seconds (or the first range), the second range (e.g., 0.8 to 1.2 seconds) can be further considered to determine whether the person is infants or the elderly. For example, infants may exhibit different patterns, so the time can be set with this in mind. More precisely, on-screen and voice guidance can be provided based on the final judgment. For example, if the person is identified as an elderly person within the second range, the speed of the on-screen and voice guidance can be automatically adjusted. While the camera can be used temporarily during this process to clearly determine infants or the elderly, proximity sensor sensing data can also be used to determine the individual using the above method.
[0051] In addition, the usability optimization unit (230) can perform an operation to help the user use the voice recognition function well by guiding the user to recognize that voice recognition is performed when using the voice recognition AI kiosk. For this purpose, when the AI robot provides guidance with TTS and when the customer speaks an order, the voice content can be displayed on the screen as text to emphasize the feeling of an actual conversation. For example, an operation to display string characters according to the speaking content and speed can be performed. As shown in Fig. 4 (c) and (d), please say [Eat and go] or [Package] in the blue box, and [Eat and go] is implemented according to the user's speaking content and speed. For this purpose, the usability optimization unit (230) can analyze the user's speaking content or speed based on the recognition result when the user's voice is recognized, and can display string characters in particular by considering the speed, etc. If the speed is considered, the characters can be displayed by dividing them into three stages, such as [Please say / Eat and go] or / [Package]. Furthermore, the spoken content can be recognized as sentences and contextually understood through a natural language recognition module. For example, artificial intelligence models such as the Large Language Model (LLM) can be used. Figure 4 (d) shows the spoken content of "Eat and Go" displayed as a string of text on one side of the microphone icon.
[0052] Furthermore, the usability optimization unit (230) can optimize the service scenario and algorithm so that the recognition of touch and voice recognition on the screen (the user can select and use touch or voice) naturally, which is a key element of the success of the AI voice recognition kiosk. This is well shown in Fig. 5. For example, if there is no action for 10 seconds on the order page after the screen changes from (a) to (b) of Fig. 5, the sound "ding-dong" can be heard along with the voice "Please tell me the order menu again." Also, in the screen (c) of Fig. 5, that is, if there is no action for 10 seconds on the option page, the voice "Please tell me the option again" can be output along with the sound "ding-dong."
[0053] The usability optimization unit (230) can output “Please say or select the required option” after 10 seconds in the screen (d) of Fig. 5 if (condition 1) the required option is not selected, but if all required options are selected (condition 2), then “Please select the additional option” or “Please say order” after 10 seconds. In the screen (e) of Fig. 5, if there is no action for 10 seconds and there are items in the shopping cart (condition 1), “Please say the menu to add to dingdong or “Please pay” can be output, and if there are no items in the shopping cart (condition 2) (no option selected, cancel button), “Please say the order menu” can be output again with the dingdong sound. In addition, the usability optimization unit (230) can perform an action to check in real time whether the order was made correctly by expressing the option name pronounced by the user in the string below.
[0054] Of course, in order to prevent confusion among users when using a voice recognition kiosk with multimodal functions in the embodiment of the present invention, the usability optimization unit (230) optimizes the service scenario and algorithm so that users can naturally select and recognize screen touch and voice recognition. However, this can be changed in part by comprehensively considering the surrounding environment of the kiosk or the service usage patterns of users. Of course, the actions such as whether there are items in the shopping cart or not in the screen (e) of Fig. 5 can be fixed as the default. However, when the kiosk is installed in an area with a large elderly population, actions such as repeatedly outputting “ding dong” “ding dong” or repeatedly outputting “Please insert cash / card” can be automatically set and performed. In reality, such automatic settings or actions can be automatically set based on the analysis results of sensing data acquired by the proximity sensor of the sensor unit (200). Currently, when the kiosk user is determined to be an elderly person, screen guidance and voice guidance actions are performed accordingly. If a regular person places an order behind an elderly person, the system can assess factors like speech rate (or tone of voice) to determine whether the person is a regular person and provide on-screen and voice guidance using pre-defined basic functions. Of course, it's also possible to use AI programs to distinguish between elderly people based on factors like speech rate or tone of voice, using learning data to identify them.
[0055] The storage unit (240) can temporarily store various types of data processed under the control of the control unit (220). The storage unit (240) can store sensing data provided by the sensor unit (200) and then provide it to the usability optimization unit (230) under the control of the control unit (220).
[0056] In addition to the above, the sensor unit (200), user interface unit (210), control unit (220), usability optimization unit (230), and storage unit (240) of FIG. 2 can perform various operations, and other detailed information has been sufficiently explained above, so it will be replaced with those contents.
[0057] Meanwhile, the sensor unit (200), user interface unit (210), control unit (220), usability optimization unit (230), and storage unit (240) of FIG. 2 according to an embodiment of the present invention are configured as physically separate hardware modules, but each module may store software for performing the above operations therein and execute the same. However, the software is a collection of software modules, and each module may be formed of hardware, so there will be no particular limitation on the configuration such as software or hardware. For example, the storage unit (240) may be hardware such as storage or memory. However, it is also possible to store information (repository) in software, so there will be no particular limitation on the above.
[0058] In addition, as another embodiment of the present invention, the control unit (220) may include a CPU and a memory, and may be formed as a single chip. The CPU may include a control circuit, an operation unit (ALU), a command interpretation unit, and a registry, and the memory may include a RAM. The control circuit may perform a control operation, the operation unit may perform an operation of binary bit information, and the command interpretation unit may perform an operation of converting a high-level language into machine language and vice versa, including an interpreter or a compiler, and the registry may be involved in software data storage. According to the above configuration, for example, at the initial operation of the intelligent kiosk (100) of FIG. 1, the program stored in the usability optimization unit (230) may be copied and loaded into the memory, i.e., RAM, and then executed, thereby rapidly increasing the data operation processing speed. In the case of a deep learning model, it may be loaded into the GPU memory instead of the RAM and executed by accelerating the execution speed using the GPU.
[0059] FIG. 6 is a flowchart showing the driving process of the intelligent kiosk of FIG. 1 according to an embodiment of the present invention.
[0060] For convenience of explanation, referring to FIG. 6 together with FIG. 1 and FIG. 2, the intelligent kiosk (device) (100) of FIG. 1 according to an embodiment of the present invention displays the voice content of the voice received through the voice receiving unit as text on the screen of the screen panel to guide the user to recognize that voice recognition is in progress when using the kiosk (S600). In FIG. 4 (d), when there is guidance for eating and taking out and when the user selects eating and taking out through the voice output unit, the voice content thereof can be recognized and displayed on the screen.
[0061] The intelligent kiosk (100) includes a voice receiving unit such as a microphone, through which the voice is received, and the received voice signal is analyzed and converted into text using a text conversion module, etc., and the text can be displayed on the screen through natural language processing or context recognition, etc. of the text. Of course, natural language processing or context recognition, etc. can be processed at any time using an artificial intelligence program such as LLM. In the embodiment of the present invention, when the AI robot provides guidance through TTS and when the customer speaks an order, the voice content is displayed on the screen as text, thereby emphasizing the feeling of an actual conversation.
[0062] In addition, the intelligent kiosk (100) can display characters in a string form on the screen panel according to the speech content (e.g., "eat and go," "package," etc.) and speed (S610). For example, the analysis results of the speech signal show that the speed is different if the user speaks five characters per second and ten characters per second. The latter case is faster. Therefore, in the latter case, the speed of the characters displayed on the screen can be slowed down compared to the former case, and the kiosk can be operated according to the user's speech speed. The speech content can mean things like "eat and go" or "package." Of course, with regard to the speech speed, it is entirely possible to classify the speech signal based on the analysis of the speech signal through an artificial intelligence program, or in other words, based on the learning results of the learning data.
[0063] In addition, the intelligent kiosk (100) can perform various operations. Since the service scenario and algorithm are optimized so that recognition of screen touch and voice recognition can be performed naturally, as a service scenario, an order page and an option page for selecting options from a menu selected from the order page are displayed on the screen, and a preset voice corresponding to the case where a required option is not selected and when all required options are selected, for example, when there are no additional selections for a specified time period of 10 seconds, can be transmitted. In addition, even on the screen where payment is made by inserting cash / card after option selection is completed on the option page, a preset voice can be transmitted by determining whether there are items in the shopping cart or no items in the shopping cart. In the former case, a voice is transmitted saying, "Please tell me the additional menu or say 'Pay'" along with a beep sound, while in the latter case, a voice is transmitted saying, "Please tell me the order menu again" along with a beep sound. For example, if the kiosk customer is determined to be elderly, the screen can be enlarged to display the text, and it is also possible to operate the kiosk by repeating the voice guidance pattern. Of course, it is also possible to set a separate layout for the elderly (e.g., enlarged text, simple layout structure, etc.) and then display the corresponding screen. The intelligent kiosk (100) can check in real time whether the user has ordered correctly by expressing the option name pronounced by the user in the string below.
[0064] In addition to the above, the intelligent kiosk (100) of FIG. 1 and FIG. 1 can perform various operations, and other detailed information has been sufficiently explained above, so it will be replaced with that information.
[0065] Meanwhile, even though all components constituting the embodiments of the present invention have been described as being combined or operating in combination, the present invention is not necessarily limited to such embodiments. That is, within the scope of the purpose of the present invention, all of the components may be selectively combined and operated one or more times. In addition, although all of the components may be implemented as individual independent hardware, some or all of the components may be selectively combined and implemented as a computer program having program modules that perform some or all of the functions of the combined hardware in one or more pieces. The codes and code segments constituting the computer program can be easily inferred by those skilled in the art of the present invention. Such a computer program may be stored in a non-transitory computer-readable storage medium and read and executed by a computer, thereby implementing the embodiments of the present invention.
[0066] Here, the non-transitory readable storage medium refers to a medium that permanently stores data and can be read by a device, rather than a medium that stores data for a short period of time, such as a register, cache, or memory. Specifically, the above-described programs may be stored and provided on a non-transitory readable storage medium, such as a CD, DVD, hard disk, Blu-ray disc, USB, memory card, or ROM.
[0067] Although the preferred embodiments of the present invention have been illustrated and described above, the present invention is not limited to the specific embodiments described above, and various modifications may be made by those skilled in the art without departing from the spirit or scope of the present invention as claimed in the claims. Furthermore, such modifications should not be understood individually from the technical idea or prospect of the present invention.
Claims
1. A user interface unit including a touch screen panel capable of touching the screen and a voice receiving unit for receiving voice; and A control unit that displays the voice content of the voice received through the voice receiving unit as text on the screen of the screen panel to guide the user to recognize that voice recognition is being performed when using the kiosk, and displays the text in the form of a string on the screen according to the voice content and speed; A multimodal intelligent kiosk device capable of artificial intelligence-based voice recognition.
2. In paragraph 1, It further includes a sensor unit that detects whether a user using the above kiosk is approaching; The control unit, when it is determined that the user has first approached the kiosk based on the sensing data of the sensor unit, displays a standby screen on the touch screen panel indicating that the kiosk is a kiosk for ordering, and when it is determined that the user's approach is for ordering, it starts screen guidance and voice guidance. The above control unit is a multimodal intelligent kiosk device capable of artificial intelligence-based voice recognition, which displays the menu and option names pronounced by the user, and the order, payment, and memo contents on the screen string to check in real time whether the order has been made correctly, and recognizes the above-mentioned text expressions when ordering again.
3. In paragraph 1, The above control unit is a multimodal intelligent kiosk device capable of artificial intelligence-based voice recognition, which displays an order page for ordering a menu and an option page for selecting options from the order menu on the screen of the touch screen panel when proceeding with screen guidance according to a preset service scenario, and simultaneously provides preset voice guidance when proceeding with the screen guidance.
4. In paragraph 3, The above control unit performs different voice guidance when a required option is selected or not selected on the option page when there is no user selection for a specified period of time on the option page, and performs different voice guidance when there is an item in the shopping cart and when there is no item in the shopping cart when there is no user selection for a specified period of time on the payment stage. The above control unit is a multimodal intelligent kiosk device capable of artificial intelligence-based voice recognition that expresses the option name pronounced by the user in a string and checks in real time whether the order has been made correctly.
5. A step for the control unit to display the voice content of the voice received through the voice receiving unit as text on the screen of the screen panel to guide the user to recognize that voice recognition is being performed when using the kiosk; and The step of the above control unit displaying the characters in the form of a string on the screen according to the voice content and speed; A method for operating a multimodal intelligent kiosk device capable of artificial intelligence-based voice recognition.
6. In paragraph 5, A step of the sensor unit detecting whether a user using the kiosk is approaching; A step for the control unit to display a standby screen on the screen panel indicating that the kiosk is a kiosk that allows verbal ordering when it is determined that the user has first approached the kiosk based on the sensing data of the sensor unit; and A step of starting screen guidance and voice guidance when the user's access is judged to be an access for ordering; A method for operating a multimodal intelligent kiosk device capable of artificial intelligence-based voice recognition, which further includes:
7. In paragraph 5, The steps to start the above screen guidance and voice guidance are: When proceeding with screen guidance according to a preset service scenario, an order page for ordering a menu and an option page for selecting options from the order menu are displayed on the screen of the touch screen panel, and a step for proceeding with preset voice guidance at the same time when proceeding with the screen guidance; and A step of displaying the results of STT on the screen in real time to check whether the order was correctly made by displaying the menu and option names pronounced by the user, as well as the order and payment and memo contents in the screen string, and recognizing the above character expressions when ordering again; A method for operating a multimodal intelligent kiosk device capable of artificial intelligence-based voice recognition, including:
8. In paragraph 7, The steps to start the above screen guidance and voice guidance are: A step of performing different voice guidance by distinguishing between cases where a required option is selected and cases where a required option is not selected on the above option page when there is no user selection for a specified period of time on the above option page; and A step of performing voice guidance differently depending on whether there are items in the shopping cart or not when there are no items in the shopping cart when there is no user selection for a specified period of time during the payment step; and A step to check in real time whether the user has ordered correctly by expressing the option name pronounced by the user in a string; A method for operating a multimodal intelligent kiosk device capable of artificial intelligence-based voice recognition, including:
Citation Information
Patent Citations
Water hammer swivelling ground drill and ground drilling method using the same
KR1020210087130A
System for providing order service for non-face-to-face and noncontact based on speech recognition
KR102216783B1
Kiosk apparatus capable of recognizing voice
KR102306092B1
Kiosk system for providing help service for kiosk use
KR102547308B1
Character recognition method and apparatus for automatically recognizing marked text
KR102848864B1