Electronic device and method for processing user voice

By converting user speech into text and classifying it into different types of intent information, and generating matching information to perform tasks, the problem of low task execution efficiency in voice assistant functions is solved, and more efficient user intent understanding and operation execution are achieved.

CN122029601APending Publication Date: 2026-05-12SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SAMSUNG ELECTRONICS CO LTD
Filing Date
2024-10-10
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing voice assistant electronic devices struggle to efficiently categorize and execute tasks corresponding to user intent when processing user speech, resulting in low task execution efficiency.

Method used

By converting user speech into text, segmenting it into multiple text fragments, classifying it into different types of intent information, generating matching information to perform corresponding tasks, and utilizing processors and memory to realize speech processing and control operations.

Benefits of technology

It improves the efficiency and accuracy of voice assistant functions, enabling them to better understand user intent and perform corresponding operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122029601A_ABST
    Figure CN122029601A_ABST
Patent Text Reader

Abstract

A method according to an embodiment may include an operation of converting a first speech of a user into text. The method may include an operation of dividing text into a plurality of text segments including a first text segment and a second text segment. The method may include classifying a first text segment mapped to intent information for performing a task as a first type of operation. The method may include classifying a second text segment that is not mapped to intent information for performing the task as a second type of operation. The method may include performing an operation of a first task including a device control operation for a target device based on first intent information corresponding to a first text segment. The method may include an operation of generating pairing information by pairing the first intent information and the second text segment. When text corresponding to the second text segment is recognized in the second speech of the user, the pairing information may be used to perform the first task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of this disclosure relate to an electronic device and method for processing user voice. Background Technology

[0002] Electronic devices, including those offering voice assistants based on user speech, are widely distributed. These devices can use artificial intelligence (AI) servers to recognize user speech and determine its meaning and intent. The AI ​​servers can interpret user speech to infer the user's intent and perform tasks based on that inferred intent. Furthermore, AI servers can perform tasks based on the user's intent expressed through natural language interactions between the user and the AI ​​server.

[0003] Electronic devices that include voice assistant functionality can perform operations in chronological order to classify domains used for processing user speech and to perform operations (e.g., applications) on tasks corresponding to the user speech within the classified domains (e.g., capsules).

[0004] The above information is presented as relevant technical information to aid in understanding this disclosure. There is no argument or determination as to whether any of the foregoing content can be used as prior art in connection with this disclosure. Summary of the Invention

[0005] Technical solution

[0006] Embodiments of this disclosure may provide a method comprising converting a user's first utterance into text. The method may include segmenting the text into multiple text segments, including a first text segment and a second text segment. The method may include classifying the first text segment, mapped to intent information for performing a task, into a first type. The method may include classifying a second text segment, not mapped to intent information for performing a task, into a second type. The method may include performing a first task, including device control operations, on a target device based on the first intent information corresponding to the first text segment. The method may include generating pairing information by pairing the first intent information with the second text segment. When text corresponding to the second text segment is identified in the user's second utterance, the pairing information can be used to perform the first task.

[0007] Another embodiment of this disclosure provides an electronic device including a processor. The electronic device may include a memory storing instructions. When executed by the processor alone or in conjunction with other processors, the instructions may cause the electronic device to convert a user's first speech into text. When executed by the processor alone or in conjunction with other processors, the instructions may cause the electronic device to segment the text into multiple text segments, including a first text segment and a second text segment. When executed by the processor alone or in conjunction with other processors, the instructions may cause the electronic device to classify the first text segment, mapped to intent information for performing a task, into a first type. When executed by the processor alone or in conjunction with other processors, the instructions may cause the electronic device to classify second text segments, not mapped to intent information for performing a task, into a second type. When executed by the processor alone or in conjunction with other processors, the instructions may cause the electronic device to perform a first task, including device control operations, on a target device based on the first intent information corresponding to the first text segment. When executed by the processor alone or in conjunction with other processors, the instructions may cause the electronic device to generate pairing information by pairing the first intent information with the second text segment. When text corresponding to the second text segment is identified in the user's second speech, the pairing information can be used to perform the first task.

[0008] Another embodiment of this disclosure may provide a method for converting a user's speech into text. The method may include segmenting the text into text segments that include at least one of a first text segment and a second text segment. The method may include classifying the first text segment, mapped to intent information for performing a task, into a first type. The method may include classifying a second text segment, not mapped to intent information for performing a task, into a second type. The method may include identifying a speech chain corresponding to the second text segment. The method may include performing a first task, including device control operations, on a target device based on the first intent information included in the speech chain. The speech chain may be a pairing of first intent information identified in speech preceding the first text segment with a second type of text segment obtained in speech preceding the first text segment.

[0009] Another embodiment of this disclosure provides an electronic device including a processor. The electronic device may include a memory storing instructions. When executed by the processor alone or in conjunction with other processors, the instructions may cause the electronic device to convert a user's speech into text. When executed by the processor alone or in conjunction with other processors, the instructions may cause the electronic device to segment the text into text segments including at least one of a first text segment and a second text segment. When executed by the processor alone or in conjunction with other processors, the instructions may cause the electronic device to classify a first text segment mapped to intent information for performing a task into a first type. When executed by the processor alone or in conjunction with other processors, the instructions may cause the electronic device to classify a second text segment not mapped to intent information for performing a task into a second type. When executed by the processor alone or in conjunction with other processors, the instructions may cause the electronic device to identify a speech chain corresponding to the second text segment. When executed by the processor alone or in conjunction with other processors, the instructions may cause the electronic device to perform a first task, including device control operations, on a target device based on the first intent information included in the speech chain. The speech chain may be a pairing of first intent information identified in a speech preceding the speech and a second type of text segment obtained in the speech preceding the speech. Attached Figure Description

[0010] Figure 1 This is a block diagram illustrating an electronic device in a network environment according to an embodiment.

[0011] Figure 2 This is a block diagram illustrating an integrated intelligent system according to an embodiment.

[0012] Figure 3 This is a diagram illustrating the form in which information about the relationship between concepts and actions according to an embodiment is stored in a database.

[0013] Figure 4 This is a diagram illustrating the screen of an electronic device that processes voice input received via a smart application according to an embodiment.

[0014] Figure 5 This is a diagram illustrating the operation of an electronic device for processing user speech according to an embodiment.

[0015] Figure 6 This is a schematic block diagram of an electronic device according to an embodiment.

[0016] Figures 7a to 10c This is a diagram illustrating the operation of an electronic device for processing user speech according to an embodiment.

[0017] Figure 11 This is a flowchart illustrating an operation method of an electronic device according to an embodiment. Detailed Implementation

[0018] In the following description, embodiments will be described in detail with reference to the accompanying drawings. When describing embodiments with reference to the accompanying drawings, the same reference numerals refer to the same elements, and repeated descriptions related to them are omitted.

[0019] Figure 1 This is a block diagram illustrating an electronic device 101 in a network environment 100 according to various embodiments.

[0020] refer to Figure 1 In network environment 100, electronic device 101 can communicate with electronic device 102 via a first network 198 (e.g., a short-range wireless communication network), or with at least one of electronic device 104 or server 108 via a second network 199 (e.g., a long-range wireless communication network). According to an embodiment, electronic device 101 can communicate with electronic device 104 via server 108. According to an embodiment, electronic device 101 may include a processor 120, memory 130, input module 150, sound output module 155, display module 160, audio module 170, sensor module 176, interface 177, connection terminal 178, haptic module 179, camera module 180, power management module 188, battery 189, communication module 190, user identification module (SIM) 196, or antenna module 197. In some embodiments, at least one component (e.g., connection terminal 178) may be omitted from electronic device 101, or one or more other components may be added to electronic device 101. In some embodiments, some of the components (e.g., sensor module 176, camera module 180, or antenna module 197) may be implemented as a single component (e.g., display module 160).

[0021] Processor 120 can execute, for example, software (e.g., program 140) to control at least one other component (e.g., hardware or software component) of electronic device 101 coupled to processor 120, and can perform various data processing or calculations. According to an embodiment, as at least part of data processing or calculation, processor 120 can store commands or data received from another component (e.g., sensor module 176 or communication module 190) in volatile memory 132, process the commands or data stored in volatile memory 132, and store the result data in non-volatile memory 134.

[0022] According to embodiments, processor 120 may be implemented as a circuit (e.g., a processing circuit), such as a system-on-a-chip (SoC) or an integrated circuit (IC). Processor 120 may include one or more processors. For example, processor 120 may include a combination of one or more processors, such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor unit (MPU), an application processor (AP), and a communication processor (CP).

[0023] According to embodiments, processor 120 may include a main processor 121 (e.g., a CPU or AP) or an auxiliary processor 123 (e.g., a GPU, neural processing unit (NPU), image signal processor (ISP), sensor central processor, or CP) that is operationally independent of or combined with the main processor 121. For example, when electronic device 101 includes a main processor 121 and an auxiliary processor 123, the auxiliary processor 123 may be adapted to consume less power than the main processor 121 or to be specifically adapted for a given function. The auxiliary processor 123 may be implemented separately from the main processor 121 or as part of the main processor 121.

[0024] When the main processor 121 is inactive (e.g., in sleep mode), the auxiliary processor 123 may control at least some of the functions or states associated with at least one component of the electronic device 101 (other than the main processor 121) (e.g., display module 160, sensor module 176, or communication module 190), or when the main processor 121 is active (e.g., running an application), the auxiliary processor 123 may work with the main processor 121 to control at least some of the functions or states associated with at least one component of the electronic device 101 (e.g., display module 160, sensor module 176, or communication module 190). According to embodiments, the auxiliary processor 123 (e.g., ISP or CP) may be implemented as part of another component (e.g., camera module 180 or communication module 190) functionally associated with the auxiliary processor 123. According to embodiments, the auxiliary processor 123 (e.g., neural processing unit) may include hardware architectures specified for processing artificial intelligence models. Artificial intelligence models can be generated through machine learning. This learning can be performed, for example, by an electronic device 101 performing artificial intelligence or via a separate server (e.g., server 108). The learning algorithm can include, but is not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. The artificial intelligence model can include multiple layers of artificial neural networks. The artificial neural network can be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of these, but is not limited thereto. Additionally or alternatively, the artificial intelligence model can include software structures in addition to hardware structures.

[0025] The memory 130 may store various data used by at least one component of the electronic device 101 (e.g., processor 120 or sensor module 176). The various data may include, for example, software (e.g., program 140) and input or output data for commands associated therewith.

[0026] According to an embodiment, memory 130 may include one or more memories. Instructions stored in memory 130 may be stored in a single memory. Instructions stored in memory 130 may be distributed and stored in multiple memories. Instructions stored in memory 130, when executed individually or jointly by processor 120, can enable electronic device 101 (e.g., Figure 2 Electronic device 201 or Figure 5 Electronic device 501) performs and / or controls reference Figures 5 to 11 The described user speech processing method. Instructions stored in memory 130, when executed individually or jointly by multiple processors, can enable electronic device 101 (e.g., Figure 2 Electronic device 201 or Figure 5 Electronic device 501) performs and / or controls reference Figures 5 to 11 The described user speech processing method. According to an embodiment, memory 130 may include volatile memory 132 or non-volatile memory 134.

[0027] Program 140 may be stored as software in memory 130 and may include, for example, an operating system (OS) 142, middleware 144, or application 146.

[0028] Input module 150 can receive commands or data from outside electronic device 101 (e.g., a user) to be used by another component of electronic device 101 (e.g., processor 120). Input module 150 may include, for example, a microphone, mouse, keyboard, keypad (e.g., button), or digital pen (e.g., stylus).

[0029] The audio output module 155 can output audio signals to the outside of the electronic device 101. The audio output module 155 may include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as playing multimedia or playing records. The receiver can be used to receive incoming calls. According to an embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.

[0030] Display module 160 can visually provide information to the outside of electronic device 101 (e.g., to a user). Display module 160 may include, for example, a display, a holographic device, or a projector, and control circuitry for controlling a respective one of the display, holographic device, and projector. According to an embodiment, display module 160 may include a touch sensor adapted to detect touch or a pressure sensor adapted to measure the intensity of the force caused by touch.

[0031] Audio module 170 can convert sound into electrical signals and vice versa. According to an embodiment, audio module 170 can obtain sound via input module 150 or output sound via an external electronic device (e.g., electronic device 102) (e.g., a speaker or headphones) that is directly (e.g., wired) or wirelessly connected to electronic device 101.

[0032] Sensor module 176 can detect the operating state of electronic device 101 (e.g., power or temperature) or the environmental state outside electronic device 101 (e.g., user state), and then generate an electrical signal or data value corresponding to the detected state. According to embodiments, sensor module 176 may include, for example, a gesture sensor, gyroscope sensor, atmospheric pressure sensor, magnetic sensor, accelerometer, grip sensor, proximity sensor, color sensor, infrared (IR) sensor, biometric sensor, temperature sensor, humidity sensor, or illuminance sensor.

[0033] Interface 177 may support one or more specific protocols used to enable electronic device 101 to connect directly (e.g., wired) or wirelessly to external electronic devices (e.g., electronic device 102). According to embodiments, interface 177 may include, for example, a High Definition Multimedia Interface (HDMI), a Universal Serial Bus (USB) interface, a Secure Digital Card (SD) interface, or an audio interface.

[0034] Connection 178 may include a connector, through which electronic device 101 may be physically connected to an external electronic device (e.g., electronic device 102). According to embodiments, connection 178 may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0035] The haptic module 179 can convert electrical signals into mechanical stimuli (e.g., vibration or motion) or electrical stimuli that can be recognized by a user through his touch or kinesthesia. According to embodiments, the haptic module 179 may include, for example, a motor, a piezoelectric element, or an electrical stimulator.

[0036] Camera module 180 can capture still or moving images. According to an embodiment, camera module 180 may include one or more lenses, an image sensor, an ISP, or a flash.

[0037] The power management module 188 manages the power supply to the electronic device 101. According to an embodiment, the power management module 188 may be implemented as at least part of, for example, a power management integrated circuit (PMIC).

[0038] Battery 189 can power at least one component of electronic device 101. According to an embodiment, battery 189 may include, for example, a non-rechargeable primary battery, a rechargeable accumulator, or a fuel cell.

[0039] Communication module 190 can support the establishment of a direct (e.g., wired) or wireless communication channel between electronic device 101 and external electronic devices (e.g., electronic device 102, electronic device 104, or server 108), and perform communication via the established communication channel. Communication module 190 may include one or more CPs that can operate independently of processor 120 (e.g., AP) and support direct (e.g., wired) or wireless communication. According to embodiments, communication module 190 may include wireless communication module 192 (e.g., cellular communication module, short-range wireless communication module, or Global Navigation Satellite System (GNSS) communication module) or wired communication module 194 (e.g., local area network (LAN) communication module or power line communication (PLC) module). One of these communication modules can communicate with an external electronic device via a first network 198 (e.g., a short-range communication network such as Bluetooth, Wi-Fi Direct, or Infrared Data Association (IrDA)) or a second network 199 (e.g., a long-range communication network such as a traditional cellular network, 5G network, next-generation communication network, the Internet, or a computer network (e.g., a LAN or a wide area network (WAN))). These various types of communication modules can be implemented as a single component (e.g., a single chip) or as multiple components separate from each other (e.g., multiple chips). The wireless communication module 192 can use user information (e.g., an International Mobile Subscriber Identity (IMSI)) stored in the SIM 196 to identify and authenticate the electronic device 101 in the communication network (such as the first network 198 or the second network 199).

[0040] Wireless communication module 192 can support 5G networks and next-generation communication technologies, such as New Radio (NR) access technologies, following 4G networks. NR access technologies can support enhanced mobile broadband (eMBB), massive machine-type communications (mMTC), or ultra-reliable and low-latency communications (URLLC). Wireless communication module 192 can support high-frequency bands (e.g., millimeter-wave bands) to achieve, for example, high data transmission rates. Wireless communication module 192 can support various technologies used to ensure performance in high-frequency bands, such as beamforming, massive MIMO, full-dimensional MIMO (FD-MIMO), array antennas, analog beamforming, or massive antennas. Wireless communication module 192 can support various requirements specified in electronic device 101, external electronic devices (e.g., electronic device 104), or network systems (e.g., second network 199). According to an embodiment, the wireless communication module 192 may support peak data rates (e.g., 20 Gbps or higher) for implementing eMBB, loss coverage (e.g., 164 dB or lower) for implementing mMTC, or U-plane delay (e.g., 0.5 ms or less for each of the downlink (DL) and uplink (UL), or 1 ms or less for round trip) for implementing URLLC.

[0041] Antenna module 197 can transmit or receive signals or power to or from the outside of electronic device 101 (e.g., external electronic device). According to an embodiment, antenna module 197 may include an antenna comprising a radiating element formed of a conductive material or conductive pattern formed in or on a substrate (e.g., a printed circuit board (PCB)). According to an embodiment, antenna module 197 may include multiple antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication scheme used in a communication network (such as a first network 198 or a second network 199) can be selected from the multiple antennas, for example by communication module 190 (e.g., wireless communication module 192). Signals or power can then be transmitted or received between communication module 190 and the external electronic device via the selected at least one antenna. According to an embodiment, another component besides the radiating element (e.g., a radio frequency integrated circuit (RFIC)) may be additionally incorporated into antenna module 197.

[0042] According to various embodiments, antenna module 197 can form a millimeter-wave antenna module. According to embodiments, the millimeter-wave antenna module may include: a PCB, an RFIC, and multiple antennas (e.g., an array antenna), wherein the RFIC is disposed on or adjacent to a first surface (e.g., a bottom surface) of the printed circuit board and is capable of supporting a specified high-frequency band (e.g., a millimeter-wave band), and the multiple antennas are disposed on or adjacent to a second surface (e.g., a top surface or a side surface) of the printed circuit board and are capable of transmitting or receiving signals in the specified high-frequency band.

[0043] At least some of the aforementioned components may be coupled to each other and transmit signals (e.g., commands or data) between them via an inter-peripheral communication scheme (e.g., bus, general purpose input and output (GPIO), serial peripheral interface (SPI), or mobile industrial processor interface (MIPI)).

[0044] According to an embodiment, commands or data can be sent or received between electronic device 101 and external electronic device 104 via server 108 connected to a second network 199. Each of electronic devices 102 or 104 can be a device of the same or different type as electronic device 101. According to an embodiment, all or some operations to be performed at electronic device 101 can be performed at one or more of external electronic devices 102, 104, or 108. For example, if electronic device 101 is required to automatically perform a function or service, or in response to a request from a user or another device, electronic device 101 may request one or more external electronic devices to perform at least a portion of the function or service, or in addition to performing the function or service, to perform at least a portion of the function or service. Upon receiving the request, one or more external electronic devices may perform at least a requested portion of the function or service, or perform additional functions or services related to the request, and transmit the result of the performance to electronic device 101. Electronic device 101 may provide the result, with or without further processing, as at least part of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technologies can be used, for example. Electronic device 101 can use, for example, distributed computing or mobile edge computing to provide ultra-low latency services. In another embodiment, external electronic device 104 may include an Internet of Things (IoT) device. Server 108 may be an intelligent server using machine learning and / or neural networks. According to embodiments, external electronic device 104 or server 108 may be included in a second network 199. Electronic device 101 can be applied to intelligent services based on 5G communication technology or IoT-related technologies (e.g., smart homes, smart cities, smart cars, or healthcare).

[0045] refer to Figure 2 The integrated intelligent system 20 according to the embodiment may include an electronic device 201 (e.g., Figure 1 Electronic device 101), intelligent server 200 (e.g., Figure 1 Server 108) and server 300 (for example, Figure 1 Server 108).

[0046] Electronic device 201 may be an internet-connected terminal device (or electronic device) and may be, for example, a mobile phone, smartphone, personal digital assistant (PDA), laptop computer, TV, white goods, wearable device, head-mounted display (HMD), or smart speaker.

[0047] According to the illustrated embodiment, electronic device 201 may include communication interface 202 (e.g., Figure 1 Interface 177), microphone 206 (e.g., Figure 1 Input module 150), speaker 205 (e.g., Figure 1 (e.g., sound output module 155), display module 204) Figure 1 The display module 160), and the memory 207 (e.g., Figure 1 The memory 130) or processor 203 (e.g., Figure 1 (Processor 120). The components listed above can be operatively connected or electrically connected to each other.

[0048] Communication interface 202 can be connected to an external device and is configured to send data to and receive data from the external device. Microphone 206 can receive sound (e.g., user speech) and convert the sound into an electrical signal. Speaker 205 can output the electrical signal as sound (e.g., speech).

[0049] Display module 204 can be configured to display images or videos. Display module 204 can also display the graphical user interface (GUI) of a running application (app). In this embodiment, display module 204 can receive touch input via a touch sensor. For example, display module 204 can receive text input via a touch sensor in the keyboard area of ​​the screen displayed on display module 204.

[0050] The memory 207 of the embodiment can store the client module 209, the software development kit (SDK) 208, and multiple applications 211. The client module 209 and the SDK 208 can be configured to perform a framework (or solution program) for general functions. In addition, the client module 209 or the SDK 208 can be configured to handle user input (e.g., voice input, text input, or touch input).

[0051] The plurality of applications 211 stored in memory 207 in the embodiment may be programs for performing specified functions. The plurality of applications 211 may include a first application 211_1, a second application 211_2, etc. According to the embodiment, each of the plurality of applications 211 may include multiple actions for a specified function. For example, the applications may include an alarm clock application, a messaging application, and / or a scheduling application. The plurality of applications 211 may be executed by processor 203 to sequentially perform at least a portion of the plurality of actions.

[0052] The processor 203 can control the overall operation of the electronic device 201. For example, the processor 203 can be electrically connected to the communication interface 202, microphone 206, speaker 205 and display module 204 to perform specified actions.

[0053] According to embodiments, processor 203 may be implemented as a circuit (e.g., a processing circuit), such as a SoC or IC. Processor 203 may include one or more processors. For example, processor 203 may include a combination of one or more processors, such as a CPU, GPU, MPU, AP, and CP.

[0054] The processor 203 in this embodiment can also perform specified functions by executing programs stored in memory 207. For example, the processor 203 can execute at least one of client module 209 or SDK 208 to perform the following operations for processing user input. The processor 203 can control the actions of multiple applications 211 through, for example, SDK 208. The following operations, which are operations of client module 209 or SDK 208, can be performed by the processor 203.

[0055] According to an embodiment, memory 207 may include one or more memories. Instructions stored in memory 207 may be stored in a single memory. Instructions stored in memory 207 may be distributed and stored in multiple memories. Instructions stored in memory 207, when executed individually or jointly by processor 203, can enable electronic device 201 (e.g., Figure 1 Electronic device 101 or Figure 5 Electronic device 501) performs and / or controls reference Figures 5 to 11The described user speech processing method. Instructions stored in memory 207, when executed individually or jointly by multiple processors, can enable electronic device 201 (e.g., Figure 1 Electronic device 101 or Figure 5 Electronic device 501) performs and / or controls reference Figures 5 to 11 The described user speech processing method.

[0056] The client module 209 in this embodiment can receive user input. For example, the client module 209 can receive voice signals corresponding to user speech sensed by the microphone 206. Alternatively, the client module 209 can receive touch input sensed by the display module 204. Alternatively, the client module 209 can receive text input sensed by a keyboard or on-screen keyboard. Furthermore, the client module 209 can receive various types of user input sensed by an input module included in or connected to the electronic device 201. The client module 209 can send the received user input to the intelligent server 200. The client module 209 can also send the status information of the electronic device 201 along with the received user input to the intelligent server 200. The status information may be, for example, application execution status information.

[0057] The client module 209 in this embodiment can receive results corresponding to the received user input. For example, when the intelligent server 200 determines a result corresponding to the received user input, the client module 209 can receive the result. The client module 209 can display the received result on the display module 204. Additionally, the client module 209 can output the received result in audio form via the speaker 205.

[0058] The client module 209 of this embodiment can receive a plan corresponding to the received user input. The client module 209 can display the results of executing multiple actions of the application according to the plan on the display module 204. For example, the client module 209 can sequentially display the results of executing multiple actions on the display module 204 and output the results in audio form via the speaker 205. In another example, the electronic device 201 can display only a portion of the results of executing multiple actions (e.g., the result of the last action) on the display module 204 and output that portion of the result in audio form via the speaker 205.

[0059] According to an embodiment, the client module 209 can receive a request from the intelligent server 200 for obtaining information needed to calculate the result corresponding to the user input. According to an embodiment, the client module 209 can respond to this request by sending the necessary information to the intelligent server 200.

[0060] The client module 209 in this embodiment can send information about the results of performing multiple actions according to a plan to the intelligent server 200. The intelligent server 200 can use the information about the results to confirm that the received user input has been processed correctly.

[0061] Client module 209 may include a speech recognition module. According to an embodiment, client module 209 can recognize voice input for limited functions via the speech recognition module. For example, client module 209 can execute a smart application for processing voice input to perform organic operations upon specified input (e.g., wake-up!).

[0062] The intelligent server 200 can receive information about user voice input from the electronic device 201 via a communication network. According to an embodiment, the intelligent server 200 can convert the data about the received voice input into text (e.g., text data). According to an embodiment, the intelligent server 200 can generate a task plan corresponding to the user's voice input based on the text.

[0063] According to an embodiment, the plan can be generated by an artificial intelligence system. The artificial intelligence system can be a rule-based system or a neural network-based system (e.g., a feedforward neural network (FNN) or an RNN). Alternatively, the artificial intelligence system can be a combination of these or other artificial intelligence systems. According to an embodiment, the plan can be selected from a set of predefined plans, or it can be generated in real time in response to a user request. For example, the artificial intelligence system can select at least one plan from predefined plans.

[0064] The intelligent server 200 can send the results of the generated plan to the electronic device 201, or send the generated plan to the electronic device 201. According to an embodiment, the electronic device 201 can display the results of the plan on the display module 204. According to an embodiment, the electronic device 201 can also display the results of actions performed according to the plan on the display module 204.

[0065] The intelligent server 200 in the embodiment may include a front-end 215, a natural language platform 220, a capsule database (DB) 230, an execution engine 240, a terminal UI 250, a management platform 260, a big data platform 270, or an analysis platform 280.

[0066] Front-end 215 can receive user input from electronic device 201. Front-end 215 can send a response corresponding to the user input.

[0067] According to an embodiment, the natural language platform 220 may include an automatic speech recognition (ASR) module 221, a natural language understanding (NLU) module 223, a planner module 225, a natural language generator (NLG) module 227, or a text-to-speech (TTS) module 229.

[0068] ASR module 221 can convert data about voice input received from electronic device 201 into text (e.g., text data). NLU module 223 can use the text of the voice input to discern the user's intent. For example, NLU module 223 can discern the user's intent by performing syntactic or semantic analysis on the user input in the form of text data. According to an embodiment, NLU module 223 can use linguistic features (e.g., grammatical elements) of morphemes or phrases to discern the meaning of words extracted from the user input, and can determine the user's intent by matching the discerned meaning of the words with the intent. In other words, NLU module 223 can obtain intent information corresponding to the user's utterance. The intent information can indicate the user's intent determined by analyzing the text. The intent information can include information indicating the user's intent to perform an action or function using the device. The intent information can also be referred to as target information. Slots can be detailed information about the intent information. Slots can be parameters necessary for an action based on the user's intent. Slots can be variable information required for the action.

[0069] Planner module 225 can generate a plan using parameters (e.g., slots) and an intent determined by NLU module 223. According to an embodiment, planner module 225 can determine multiple domains required for a task based on the determined intent. Planner module 225 can determine multiple actions included in each of the multiple domains determined based on the intent. According to an embodiment, planner module 225 can determine the parameters required for the determined multiple actions, or the result values ​​output by executing the multiple actions. Parameters and result values ​​can be defined as concepts of a specified form (or category). Therefore, the plan can include multiple actions and multiple concepts determined by the user's intent. Planner module 225 can determine the relationships between the multiple actions and the multiple concepts step-by-step (or hierarchically). For example, planner module 225 can determine the execution order of the multiple actions determined based on the user's intent based on multiple concepts. In other words, planner module 225 can determine the execution order of the multiple actions based on the parameters required to execute the multiple actions and the results output by executing the multiple actions. Therefore, planner module 225 can generate a plan that includes connection information (e.g., ontology) regarding the connections between the multiple actions and the multiple concepts. The planner module 225 can use information stored in capsule DB 230 to generate plans, which stores a set of relationships between concepts and actions.

[0070] NLG module 227 can convert specified information into text form. The information converted into text form can be in the form of natural language speech. TTS module 229 can convert information in text form into information in speech form.

[0071] According to an embodiment, some or all of the functions of the natural language platform 220 may also be implemented in the electronic device 201.

[0072] Capsule DB 230 can store information about relationships between multiple concepts and actions corresponding to multiple domains. According to embodiments, capsules may include multiple action objects (or action information) and concept objects (or concept information) included in the plan. According to embodiments, capsule DB 230 may store multiple capsules in the form of a Concept Action Network (CAN). According to embodiments, multiple capsules may be stored in a function registry included in capsule DB 230.

[0073] Capsule DB 230 may include a strategy registry storing strategy information required to determine a plan corresponding to voice input. The strategy information may include reference information for determining a plan when multiple plans exist corresponding to user input. According to an embodiment, Capsule DB 230 may include a follow-up registry storing information about follow-up actions for suggesting follow-up actions to the user in specified situations. Follow-up actions may include, for example, follow-up utterances. According to an embodiment, Capsule DB 230 may include a layout registry storing layout information output by electronic device 201. According to an embodiment, Capsule DB 230 may include a vocabulary registry storing vocabulary information included in the capsule information. According to an embodiment, Capsule DB 230 may include a dialogue registry storing information about dialogue (or interaction) with the user. Capsule DB 230 can update stored objects through developer tools. Developer tools may include, for example, a function editor for updating action objects or concept objects. Developer tools may include a vocabulary editor for updating vocabulary. Developer tools may include a strategy editor for generating and registering strategies for determining plans. Developer tools may include a dialogue editor for generating dialogue with the user. Developer tools may include a follow-up editor capable of activating subsequent goals and editing follow-up statements that provide prompts. The follow-up goal may be determined based on the currently set goal, user preferences, or environmental conditions. In this embodiment, capsule DB 230 may also be implemented in electronic device 201.

[0074] The execution engine 240 can use the generated plan to calculate the results. The terminal user interface 250 can send the calculation results to the electronic device 201. Therefore, the electronic device 201 can receive the results and provide them to the user. The management platform 260 can manage the information used by the intelligent server 200. The big data platform 270 can collect user data. The analysis platform 280 can manage the quality of service (QoS) of the intelligent server 200. For example, the analysis platform 280 can manage the components and processing speed (or efficiency) of the intelligent server 200.

[0075] Service server 300 can provide specified services (e.g., food orders or hotel reservations) to electronic device 201. According to an embodiment, service server 300 can be a server operated by a librarian. Services of service server 300 (such as CP service A 301 and CP service B 302) can interact with the front end 215 of intelligent server 200. Service server 300 can provide intelligent server 200 with information to be used to generate plans corresponding to received user input. The provided information can be stored in capsule DB 230. Furthermore, service server 300 can provide intelligent server 200 with results information based on the plans.

[0076] In the aforementioned integrated intelligent system 20, the electronic device 201 can provide various intelligent services to the user in response to user input. User input may include, for example, input via physical buttons, touch input, or voice input.

[0077] In this embodiment, the electronic device 201 can provide voice recognition services through a smart application (or voice recognition application) stored therein. For example, the electronic device 201 can recognize user speech or voice input received through a microphone and provide the user with services corresponding to the recognized voice input.

[0078] In this embodiment, the electronic device 201 can perform a specified action based on the received voice input, either alone or in conjunction with a smart server and / or a service server. For example, the electronic device 201 can execute an application corresponding to the received voice input and perform the specified action through the executed application.

[0079] In this embodiment, when the electronic device 201 provides services together with the intelligent server 200 and / or the service server 300, the electronic device 201 can use the microphone 206 to detect user speech and generate a signal (or voice data) corresponding to the detected user speech. The electronic device 201 can use the communication interface 202 to send the voice data to the intelligent server 200.

[0080] In response to voice input received from electronic device 201, intelligent server 200 can generate a plan for a task corresponding to the voice input or the result of performing actions according to the plan. The plan may include, for example, multiple actions for a task corresponding to the user's voice input, and multiple concepts associated with the multiple actions. A concept may be defined as parameters for inputs used to perform the multiple actions or as the result value output by performing the multiple actions. The plan may include connection information regarding the connections between the multiple actions and the multiple concepts.

[0081] Electronic device 201 can receive responses using communication interface 202. Electronic device 201 can output voice signals generated internally to the outside using speaker 205, or it can output images generated internally to the outside using display module 204.

[0082] Figure 3 This is a diagram illustrating the form in which information about the relationship between concepts and actions according to an embodiment is stored in a database.

[0083] Intelligent servers (e.g., Figure 2 Capsule DB (e.g., intelligent server 200) Figure 2 The capsule DB 230 can store capsules in CAN 400 format. The capsule DB can store actions for processing tasks corresponding to voice input from the user, as well as the parameters required for those actions, in CAN format.

[0084] The capsule database can store multiple capsules (capsule A 401 and capsule B 404) corresponding to multiple domains respectively. According to an embodiment, a capsule (e.g., capsule A 401) may correspond to a domain (e.g., location (geo)). Additionally, a capsule may correspond to at least one service provider (e.g., CP 1 402 or CP 2 403) for a function associated with the domain. According to an embodiment, a capsule may include at least one action 410 and at least one concept 420 to perform a specified function. CAN 400 may store other information, such as CP 3 406, and capsule B 404 may correspond to another service provider (e.g., CP 4 405).

[0085] Natural language platforms (e.g., Figure 2 The natural language platform 220 can use capsules stored in the capsule DB to generate a plan for a task corresponding to the received speech input. For example, the planner module of the natural language platform (e.g., Figure 2The planner module 225 can use capsules stored in the capsule DB to generate plans. For example, plans 407 can be generated using actions 4011 and 4013 and concepts 4012 and 4014 of capsule A 401 and actions 4041 and concepts 4042 of capsule B 404.

[0086] Figure 4 This is a diagram illustrating the screen of an electronic device that processes voice input received via a smart application according to an embodiment.

[0087] Electronic device 201 can execute intelligent applications via an intelligent server (e.g., Figure 2 The intelligent server 200 processes user input.

[0088] According to an embodiment, on screen 310, when a specified voice input (e.g., wake-up!) is recognized or input is received via a hardware key (e.g., a dedicated hardware key), electronic device 201 can execute a smart application for processing voice input. Electronic device 201 can execute the smart application, for example, while executing a scheduling application. According to an embodiment, electronic device 201 can be displayed on display module 204 (e.g., ...). Figure 1 Display module 160 and Figure 2 The display module 204 displays an object (e.g., an icon) 311 corresponding to the smart application. According to an embodiment, the electronic device 201 can receive voice input via user speech. For example, the electronic device 201 can receive voice input such as "Tell me my schedule for this week!". According to an embodiment, the electronic device 201 can display a user interface (UI) 313 (e.g., an input window) of the smart application, wherein the text (e.g., text data) of the received voice input is displayed on the display module 204.

[0089] According to an embodiment, on screen 320, electronic device 201 can display the result corresponding to the received voice input on display module 204. For example, electronic device 201 can receive a schedule corresponding to the received user input and display "this week's schedule" on display module 204 according to the schedule.

[0090] Figure 5 This is a diagram illustrating the operation of an electronic device for processing user speech according to an embodiment.

[0091] refer to Figure 5 The electronic device 501 may include a reference Figure 1 The described electronic device 101 and reference Figure 2 The described electronic device 201 includes at least some components. The intelligent server 601 may include references. Figure 2At least some components of the intelligent server 200 are described. References to electronic device 501 and intelligent server 601 are omitted. Figures 1 to 4 The provided description is duplicated.

[0092] According to an embodiment, electronic device 501 (e.g., Figure 1 Electronic device 101 or Figure 2 The electronic device 201 can be connected to the intelligent server 601 (e.g., via a LAN, WAN, value-added network (VAN), mobile radio communication network, satellite communication network, or any combination thereof) Figure 2 The intelligent server 200. Electronic device 501 and intelligent server 601 can communicate with each other via wired or wireless communication methods (e.g., Wireless LAN (WiFi), Bluetooth, Bluetooth Low Energy, ZigBee, WiFi Direct (WFD), Ultra Wideband (UWB), Infrared Data Association (IrDA), and Near Field Communication (NFC)). Electronic device 501 can communicate with peripheral devices around electronic device 501 (e.g., Figure 1 The electronic device 102 or electronic device 104 performs communication.

[0093] According to an embodiment, the electronic device 501 may be implemented as at least one of the following: a smartphone, a tablet PC, a mobile phone, a speaker (e.g., an AI speaker), a video phone, an e-book reader, a desktop PC, a laptop PC, a netbook computer, a workstation, a server, a personal digital assistant (PDA), a portable multimedia player (PMP), an MP3 player, a mobile medical device, a camera, or a wearable device.

[0094] According to an embodiment, electronic device 501 can acquire a speech signal corresponding to a user's speech and can send the speech signal to intelligent server 601. Intelligent server 601 can acquire text (e.g., text data) corresponding to the user's speech based on the speech signal. The text can be acquired by performing ASR on the speech signal to convert the speech portion into computer-readable text data. Intelligent server 601 can use the text to analyze the user's speech. Intelligent server 601 can use the analysis results (e.g., intent information, domains, and / or capsules) to perform necessary functions, or can provide responses (e.g., questions and answers) to be provided to the user to the device (e.g., electronic device 501). Intelligent server 601 can be implemented as software. Part or all of intelligent server 601 can be implemented in electronic device 501. In other words, artificial intelligence for processing speech without communicating with intelligent server 601 can be installed on electronic device 501. (See references...) Figures 2 to 4 At least some components of the described natural language platform 220 can be implemented in the electronic device 501.

[0095] According to an embodiment, the electronic device 501 can perform tasks corresponding to user input (e.g., user speech) (e.g., a unit specified by the manufacturer of the electronic device 501 equipped with a voice assistant) (e.g., device control operation on a target device).

[0096] According to an embodiment, firstly, the ASR module (e.g., included in the electronic device 501) Figure 2 The ASR module 221 can convert user speech into text (e.g., text data). The electronic device 501 can also use the NLU module (e.g., ...) Figure 2 The NLU module 223) determines the domain and / or intent information corresponding to the user's utterance based on the text.

[0097] According to embodiments, a domain can be a category (or service) associated with an action (or function) that a user expects to perform using the device. Domains can be categorized based on the services provided therefrom. For example, a music playback domain could support music playback services (e.g., music playback services included in the Melon and Spotify apps). A communication domain could support communication services (e.g., communication services included in messaging, chat, and email apps). Multiple user utterances can be processed separately based on their respective domains. Tasks corresponding to user utterances can be processed within capsules (e.g., applications). One capsule can correspond to one domain. A capsule can include at least one action and at least one concept for a predetermined function. Capsules can process tasks corresponding to user utterances based on intent information. Intent information can be determined within the capsule or in the NLU module.

[0098] According to an embodiment, intent information may be information indicating a user's intent determined through analysis of text (e.g., text data). Intent information may include information indicating the user's intention to perform an action (or function) using the terminal. Intent information can be used to perform a task. Intent information may be predefined by the manufacturer of the electronic device 501 equipped with a voice assistant. Intent information may also be referred to as target information.

[0099] According to an embodiment, the slot can be detailed information about intent. The slot can be parameters necessary for an action based on the user's intent.

[0100] According to embodiments, electronic device 501 can perform tasks corresponding to a user's voice input (e.g., device control operations) based on domains, intent information, and slots. For example, when the text converted from the user's voice input is "What time is it in San Francisco?", the domain could be the "Date and Time Domain", the intent information could correspond to "Date / Time Information Provision", the slot could be "San Francisco", and the capsule (e.g., an app) could provide the user with the time in San Francisco. For example, when the text converted from the user's voice input is "What's the weather like here?", the domain could be the "Weather Domain", the intent information could correspond to "Weather Information Provision", the slot could be "Current Location", and the capsule could provide the weather for the current location. For example, when the text converted from the user's voice input is "Set the oven temperature to 300 degrees", the domain could be the "Device Control Domain", the intent information could correspond to "Oven Control", the slot could be "300 degrees", and the capsule could attempt to set the oven temperature to 300 degrees.

[0101] Traditional voice assistants can perform actions based on predefined intent information (e.g., intent). When the predefined intent information matches the user input, the voice assistant can perform the action corresponding to the matched intent information. Voice assistants can perform accurate actions when they receive appropriate and concise utterances. However, voice assistants may not be able to accurately handle sentences or phrases with forms different from the trained utterances (e.g., by treating sentences or phrases as exceptions or by misidentifying user input and performing incorrect actions). As voice assistant capabilities advance, the frequency with which users input large amounts of text at once (e.g., as utterances) is increasing. Due to relatively limited resource capacity, conventional voice assistants may struggle to accurately handle large amounts of input.

[0102] In the case of dialogue engines based on Large Language Models (LLMs), accurate results can be obtained when detailed utterances are provided. LLM-based dialogue engines, using probabilistic models based on artificial neural networks, can be trained on large-scale language corpora and can efficiently handle diverse user inputs. Due to their relatively large resource capacity, LLM-based dialogue engines can handle a large amount of input. However, as LLM-based dialogue engines become more resource-intensive, their response time to user input increases.

[0103] According to an embodiment, electronic device 501 can process various user inputs (e.g., speech). Electronic device 501 can use the user's speech pattern. Electronic device 501 can respond to speech that does not match predefined intent information based on the user's speech pattern. Electronic device 501 can use user input that indicates a user situation.

[0104] According to an embodiment, electronic device 501 can appropriately process a large amount of user input (e.g., speech). Electronic device 501 can segment the user input. Electronic device 501 can respond to multi-intent scenarios (e.g., multi-intent speech or multi-target speech) by segmenting a large amount of user input.

[0105] Electronic device 501 can receive input (e.g., utterance) from a user (e.g., “It’s really hot today…”). Electronic device 501 can convert the user’s utterance (e.g., “It’s really hot today…”) into text (e.g., text data) (e.g., “It’s really hot today…”). Electronic device 501 can segment the text into text fragments based on the probability of matching intent information. Since the text (e.g., “It’s really hot today…”) does not include any parts that may match intent information, the text may not be segmented. However, for the sake of terminology consistency, the text (e.g., “It’s really hot today…”) may also be referred to as a text fragment (e.g., “It’s really hot today…”) below. Electronic device 501 can analyze the text fragment (e.g., “It’s really hot today…”) and classify the text fragment (e.g., “It’s really hot today…”) into a second type of text fragment that is not mapped to intent information (e.g., intent information for performing a task). Electronic device 501 can recognize a discourse chain (e.g., text fragment "It's really hot today," intent information "Turn on the air conditioner," and intent information "Set the air conditioner temperature to 22 degrees") that matches a second type of text fragment (e.g., "It's really hot today..."). Electronic device 501 can perform a task corresponding to the discourse chain (e.g., text fragment "It's really hot today," intent information "Turn on the air conditioner," and intent information "Set the air conditioner temperature to 22 degrees"). Electronic device 501 can provide a response to the user (e.g., "Do you want to turn on the air conditioner and set it to fanless mode?") to perform the task.

[0106] According to the embodiment, electronic device 501 may not require user-predefined triggers (e.g., the trigger "very hot" corresponding to the intent information "air conditioner on" and "set air conditioner temperature to 22 degrees"). Electronic device 501 may not require explicit trigger definitions.

[0107] According to an embodiment, electronic device 501 may disregard the association between repeatedly used intent information (e.g., "turn on the air conditioner" and "set the air conditioner temperature to 22 degrees"). Electronic device 501 can extract utterances that do not match the intent information from the user's utterance patterns and use the extracted utterances.

[0108] According to an embodiment, electronic device 501 may not require a probabilistic model (such as LLM). Electronic device 501 can minimize failure problems (e.g., hallucinations). Electronic device 501 can provide a user-predictable response as an intermediate method between direct input and complete inference.

[0109] Figure 6 This is a schematic block diagram of an electronic device according to an embodiment.

[0110] refer to Figure 6 The electronic device 501 may include a reference Figure 1 The described electronic device 101 and reference Figure 2 At least some components of the described electronic device 201. As described above, components for use in non-interacting with a smart server (e.g., Figure 2 Intelligent Server 200 or Figure 5 Artificial intelligence on a device that processes speech in the context of communication with an intelligent server (601). (See above for reference.) Figures 2 to 4 At least some of the functions of the described natural language platform 220 can be implemented in electronic device 501. References to electronic device 501 are omitted. Figures 1 to 4 The provided description is duplicated.

[0111] According to an embodiment, electronic device 501 may include wireless communication circuit 510 (e.g., Figure 1 The wireless communication module 192). Electronic device 501 may include processor 520 (e.g., wireless communication module 192). Figure 1 Processor 120 or Figure 2 The processor 203). Electronic device 501 may include memory 530 (e.g., processor 203). Figure 1 memory 130 or Figure 2 The memory 501 is 507. The processor 520 (e.g., an AP) can execute one or more instructions by accessing the memory 530. The processor 520 can enable the electronic device 501 to provide a response to the user. The memory 530 can store various types of data used by at least one component of the electronic device 501 (e.g., the processor 520).

[0112] According to an embodiment, processor 520 may be implemented as a circuit (e.g., a processing circuit), such as a SoC or IC. Processor 520 may include one or more processors. For example, processor 520 may include a combination of one or more processors, such as a CPU, GPU, MPU, AP, and CP.

[0113] According to an embodiment, memory 530 may include one or more memories. Instructions stored in memory 530 may be stored in a single memory. Instructions stored in memory 530 may be distributed and stored in multiple memories. Instructions stored in memory 530, when executed individually or jointly by processor 520, can enable electronic device 501 (e.g., Figure 1 Electronic device 101 or Figure 2 Electronic device 201) performs and / or controls reference Figures 5 to 11 The described user speech processing method. Instructions stored in memory 530, when executed individually or jointly by multiple processors, can enable electronic device 501 (e.g., Figure 1 Electronic device 101 or Figure 2 Electronic device 201) performs and / or controls reference Figures 5 to 11 The described user speech processing method.

[0114] According to an embodiment, electronic device 501 can receive input (e.g., speech) from a user. Electronic device 501 may be based on ASR module 521 (e.g., Figure 2 The automatic speech recognition module 221 converts the user's speech into text (e.g., text data).

[0115] According to an embodiment, electronic device 501 can segment text based on text segmentation module 522. Electronic device 501 can segment compound or complex sentences into at least one of simple sentences, independent clauses, or subordinate clauses. Electronic device 501 can segment text based on intent information. (References already provided) Figure 5 The intent information has been described in detail, therefore its repetitive description will be omitted. Electronic device 501 can segment the text based on the probability of a match with the intent information. Electronic device 501 can segment the text to obtain text fragments.

[0116] According to an embodiment, electronic device 501 may use classifier 523 (e.g., Figure 2The classifier 523 uses an NLU (Network Logic Unit) to analyze text segments. The classifier 523 can identify the user's intent based on each text segment. The classifier 523 can perform syntactic analysis, semantic analysis, or statistical classification (e.g., pattern recognition or machine learning) on ​​each text segment to identify the user's intent. In other words, the classifier 523 can obtain intent information corresponding to each text segment. Text segments may include text segments from which intent information is not obtained (e.g., derived). Text segments from which intent information is not derived (or do not match intent information) can be classified as a second type. Text segments from which intent information is derived (or match intent information) can be classified as a first type. The classifier 523 can classify text segments into first-type text segments that match intent information (e.g., first text segments) and second-type text segments that do not match intent information (e.g., second text segments). Second-type text segments may include text indicating the user's situation. Second-type text segments may include text classified as unprocessable commands or as descriptive sentences without intent.

[0117] According to an embodiment, electronic device 501 can use speech chain management module 524 to manage speech chains and / or pairing information. Pairing information can be a pairing of first intent information (e.g., first intent information corresponding to a first type of text fragment) and a second type of text fragment. When the number of pairings between the first intent information and the second type of text fragment exceeds a threshold, speech chain management module 524 can determine the pairing information as a speech chain. When text corresponding to a second type of text fragment (e.g., a second type of text fragment included in the speech chain) is identified from the user's subsequent speech, electronic device 501 can perform a task based on the first intent information included in the speech chain. The speech chain management operation will be described in detail below.

[0118] Figures 7a to 10c This is a diagram illustrating the operation of an electronic device for processing user speech according to an embodiment.

[0119] refer to Figure 7a According to an embodiment, electronic device 501 can generate speech chains (e.g., scenario 701) and can use speech chains (e.g., scenario 702).

[0120] According to an embodiment, in scenario 701, electronic device 501 may receive input (e.g., speech) from a user (e.g., “It’s really hot today. Turn on the air conditioner and set it to windless mode.”). Electronic device 501 may convert the user’s speech (e.g., “It’s really hot today. Turn on the air conditioner and set it to windless mode.”) into text (e.g., text data) (e.g., “It’s really hot today. Turn on the air conditioner and set it to windless mode.”). Electronic device 501 may segment the text based on the probability of matching intent information. The text (e.g., “It’s really hot today. Turn on the air conditioner and set it to windless mode”) may be segmented into text fragments (e.g., “It’s really hot today,” “Turn on the air conditioner,” and “Set it to windless mode”). Electronic device 501 may determine that the text (e.g., “It’s really hot today. Turn on the air conditioner and set it to windless mode.”) is a complex sentence, and may segment the complex sentence into clauses to obtain three text fragments (e.g., “It’s really hot today,” “Turn on the air conditioner,” and “Set it to windless mode”). Electronic device 501 may perform intent determination based on each segmented text fragment. Electronic device 501 can analyze text fragments (e.g., “It’s really hot today,” “Turn on the air conditioner,” and “Set it to windless mode”) to classify the text fragments (e.g., “Turn on the air conditioner” and “Set it to windless mode”) into a first type of text fragment that matches intent information, and classify the text fragments (e.g., “It’s really hot today”) into a second type of text fragment that does not match intent information. Electronic device 501 can generate discourse chains (e.g., text fragment “It’s really hot today,” intent information “Turn on the air conditioner,” and intent information “Set it to windless mode”) by pairing first intent information (e.g., intent information “Air conditioner is on” and intent information “Set it to windless mode”) corresponding to the first type of text fragments (e.g., “Turn on the air conditioner” and “Set it to windless mode”) with second type of text fragments (e.g., “It’s really hot today”) (e.g., second text fragment). Electronic device 501 can perform a task corresponding to the first intent information (e.g., intent information "air conditioner on" and intent information "set it to windless mode") (e.g., device control operation on the air conditioner as the target device) (e.g., turning on the air conditioner and setting it to windless mode). Electronic device 501 can provide a response to the user after performing the task (e.g., "the air conditioner has been turned on and set to windless mode").When the electronic device 501 identifies text corresponding to a second type of text fragment (e.g., “It’s really hot today”) from the user’s subsequent utterance (e.g., a second utterance), the electronic device 501 can perform a task (e.g., a task corresponding to the first intent information (e.g., intent information “It’s really hot today”, intent information “Turn on the air conditioner”, and intent information “Set it to windless mode”) based on the utterance chain (e.g., text fragment “It’s really hot today”, intent information “Turn on the air conditioner”, and intent information “Set it to windless mode”).

[0121] According to an embodiment, in scenario 702, electronic device 501 may receive input (e.g., utterance) from a user (e.g., “It’s really hot today…”). Electronic device 501 may convert the user’s utterance (e.g., “It’s really hot today…”) into text (e.g., text data) (e.g., “It’s really hot today…”). Electronic device 501 may segment the text based on the probability of matching intent information. Since the text (e.g., “It’s really hot today…”) does not include any parts that may match the intent information, the text may not be segmented. Electronic device 501 may segment text that includes compound sentences or complex sentences. Since the text (e.g., “It’s really hot today…”) is a simple sentence, the text may not be segmented. However, for consistency of terminology, the text (e.g., “It’s really hot today…”) may also be referred to as a text fragment (e.g., “It’s really hot today…”) hereinafter. Electronic device 501 may analyze the text fragment (e.g., “It’s really hot today…”) and classify the text fragment (e.g., “It’s really hot today…”) into a second type of text fragment that does not match the intent information (e.g., intent information for performing a task). Electronic device 501 can recognize a utterance chain (e.g., text fragment "It's really hot today...", intent information "Turn on the air conditioner" and intent information "Set it to windless mode") that matches a second type of text fragment (e.g., "It's really hot today..."). Electronic device 501 can perform a task corresponding to the utterance chain (e.g., text fragment "It's really hot today", intent information "Turn on the air conditioner" and intent information "Set it to windless mode") (e.g., device control operation on the air conditioner as the target device) (e.g., turning on the air conditioner and setting it to windless mode). Electronic device 501 can provide a response to the user before performing the task (e.g., "Would you like to turn on the air conditioner and set it to windless mode?"). It should be noted that for the utterance chain that serves as a trigger operation for the task, the number of times the utterance chain is obtained (e.g., scenario 701) (e.g., the number of pairings between the first intent information and the second type of text fragment) may need to exceed a threshold.

[0122] refer to Figure 7bAccording to an embodiment, in scenario 703, electronic device 501 can receive input (e.g., utterance) from a user (e.g., “It’s really hot today”). Electronic device 501 can convert the user’s utterance (e.g., “It’s really hot today”) into text (e.g., text data) (e.g., “It’s really hot today”). Electronic device 501 can segment the text based on the probability of matching intent information. Since the text (e.g., “It’s really hot today”) does not include any part that may match the intent information, the text may not be segmented. Electronic device 501 can segment text that includes compound sentences or complex sentences. Since the text (e.g., “It’s really hot today”) is a simple sentence, the text may not be segmented. However, for consistency of terminology, the text (e.g., “It’s really hot today”) may also be referred to as a text fragment (e.g., “It’s really hot today”) below. Electronic device 501 can analyze the text fragment (e.g., “It’s really hot today”) and classify the text fragment (e.g., “It’s really hot today”) into a second type of text fragment that does not match the intent information. Electronic device 501 can identify discourse chains (e.g., text fragment "It's really hot today", intent information "turn on the air conditioner", and intent information "set it to windless mode") that match a second type of text fragment (e.g., "It's really hot today"). Discourse chain recognition can be performed based on determining the similarity between text fragments included in the discourse chain (e.g., "It's really hot today") and second-type text fragments (e.g., "It's really hot today") classified by electronic device 501. The determination of similarity between text fragments can be performed based on matching scores between sentences (e.g., scores based on edit distance (e.g., Levenshtein distance) used to determine similarity between strings, or scores based on term frequency-inverse document frequency (TF-IDF) used to assess word importance). Text fragments with matching scores between sentences greater than or equal to a threshold can be determined as similar. Electronic device 501 can perform a task corresponding to the discourse chain (e.g., text fragment "It's really hot today", intent information "turn on the air conditioner", and intent information "set it to windless mode") (e.g., turning on the air conditioner and setting it to windless mode). Electronic device 501 can provide a response to the user before performing a task (e.g., “Do you want to turn on the air conditioner and set it to windless mode?”).

[0123] refer to Figure 8a and Figure 8b According to the embodiments, discourse chains can be generated based on multi-turn discourse.

[0124] According to an embodiment, in scenario 801, electronic device 501 can receive input (e.g., speech) from a user (e.g., “I’m thirsty. Set up the water purifier.”). Electronic device 501 can convert the user’s speech (e.g., “I’m thirsty. Set up the water purifier.”) into text (e.g., text data) (e.g., “I’m thirsty. Set up the water purifier.”). Electronic device 501 can segment the text based on the probability of matching intent information. The text (e.g., “I’m thirsty. Set up the water purifier”) can be segmented into text fragments (e.g., “I’m thirsty” and “Set up the water purifier”). Electronic device 501 can analyze the text fragments (e.g., “I’m thirsty” and “Set up the water purifier”) to classify the text fragments (e.g., “Set up the water purifier”) into a first type of text fragment that matches the intent information, and classify the text fragments (e.g., “I’m thirsty”) into a second type of text fragment that does not match the intent information. The first intent information (e.g., intent information “Set up the water purifier”) corresponding to the first type of text fragment (e.g., “Set up the water purifier”) may require additional parameters (e.g., slot). Electronic device 501 can provide the user with a response for obtaining additional parameters (e.g., "Which mode do you want: room temperature water, cold water, or hot water?").

[0125] According to an embodiment, in scenario 802, electronic device 501 can receive input from a user (e.g., "set it to cold water mode"). Based on the input (e.g., "set it to cold water mode"), electronic device 501 can determine first intent information (and / or slot) (e.g., "set it to cold water mode"). Electronic device 501 can generate a discourse chain (e.g., text fragment "I'm thirsty" and intent information "set it to cold water mode") by pairing the first intent information (e.g., intent information "set it to cold water mode") corresponding to a first type of text fragment (e.g., "set it to cold water mode") with a second type of text fragment (e.g., "I'm thirsty"). Electronic device 501 can perform a task corresponding to the first intent information (e.g., intent information "set it to cold water mode") (e.g., device control operation of a water purifier as the target device) (e.g., setting the water purifier to cold water mode). Electronic device 501 can provide a response to the user after performing the task (e.g., "the water purifier has been set to cold water mode"). Reference Figure 8a In multi-turn scenarios that maintain the conversational context, electronic device 501 can include a second round of input in the discourse chain (e.g., the utterance “set it to cold water mode.”).

[0126] refer to Figure 8bAccording to an embodiment, in scenario 803, electronic device 501 can receive input from a user (e.g., “Turn on the air conditioner. And I’m really thirsty. Set up the water purifier.”). Based on the input (e.g., “Turn on the air conditioner. And I’m really thirsty. Set up the water purifier.”), electronic device 501 can obtain first intent information (e.g., “Air conditioner on” and “Set up the water purifier”). See above reference. Figure 8a The intent information “set up the water purifier” may require additional parameters (e.g., tank location). Electronic device 501 can use other first intent information (e.g., “air conditioner on”) to obtain additional parameters for the intent information (e.g., “set up the water purifier”). Other first intent information (e.g., “air conditioner on”) may indicate that the user feels hot. Based on other first intent information (e.g., “air conditioner on”), electronic device 501 can determine the first intent information (and / or tank location) (e.g., “set it to cold water mode”). Electronic device 501 can generate a discourse chain (e.g., text fragment “I’m thirsty”, intent information “set it to cold water mode”, and intent information “air conditioner on”) by pairing first intent information corresponding to first type text fragments (e.g., “turn on the air conditioner” and “set up the water purifier”) with second type text fragments (e.g., “I’m thirsty”). Electronic device 501 can perform tasks corresponding to the first intent information (e.g., intent information "air conditioner on" and intent information "set it to cold water mode") (e.g., device control operations on the water purifier and air conditioner as target devices) (e.g., turning on the air conditioner and setting the water purifier to cold water mode). Electronic device 501 can provide a response to the user after performing the task (e.g., "Yes, the air conditioner has been turned on and the water purifier has been set to cold water mode.").

[0127] refer to Figure 9 According to an embodiment, the electronic device 501 can segment text, and therefore can even extract and use second-type text fragments from inverted sentences.

[0128] According to an embodiment, in scenario 901, electronic device 501 can receive input (e.g., speech) from a user (e.g., “Charging the robot vacuum. I’m too lazy to move. Turn off the living room light.”). Electronic device 501 can convert the user’s speech (e.g., “Charging the robot vacuum. I’m too lazy to move. Turn off the living room light.”) into text (e.g., text data) (e.g., “Charging the robot vacuum. I’m too lazy to move. Turn off the living room light.”). Electronic device 501 can segment the text based on the probability of matching intent information. The text (e.g., “Charging the robot vacuum. I’m too lazy to move. Turn off the living room light.”) can be segmented into text fragments (e.g., “Charging the robot vacuum,” “I’m too lazy to move,” and “Turn off the living room light”). Electronic device 501 can analyze text fragments (e.g., “charge the robot vacuum cleaner”, “I’m too lazy to move”, and “turn off the living room light”) to classify them into a first type of text fragment that matches intent information, and a second type of text fragment that does not match intent information. Electronic device 501 can generate a discourse chain (e.g., text fragment “I’m too lazy to move”, intent information “start charging the robot vacuum cleaner”, and intent information “turn off the living room light”) by pairing first intent information (e.g., intent information “start charging the robot vacuum cleaner” and intent information “turn off the living room light”) corresponding to the first type of text fragments (e.g., “I’m too lazy to move”) with the second type of text fragments (e.g., “I’m too lazy to move”). Electronic device 501 can perform tasks corresponding to the first intent information (e.g., intent information "start charging the robot vacuum cleaner" and intent information "turn off the living room light") (e.g., device control operations on the robot vacuum cleaner and the living room light as target devices) (e.g., start charging the robot vacuum cleaner and turn off the living room light). Electronic device 501 can provide a response to the user after performing the task (e.g., "The robot vacuum cleaner is charging and the living room light is off").

[0129] According to an embodiment, in scenario 902, electronic device 501 can receive input (e.g., speech) from a user (e.g., “I’m too lazy to move…”). Electronic device 501 can convert the user’s speech (e.g., “I’m too lazy to move…”) into text (e.g., text data) (e.g., “I’m too lazy to move…”). Electronic device 501 can segment the text based on the probability of matching intent information. Since the text (e.g., “I’m too lazy to move…”) does not include any part that can match the intent information, the text may not be segmented. However, for consistency of terminology, the text (e.g., “I’m too lazy to move…”) may also be referred to as a text fragment (e.g., “I’m too lazy to move…”) below. Electronic device 501 can analyze the text fragment (e.g., “I’m too lazy to move…”) and classify the text fragment (e.g., “I’m too lazy to move…”) into a second type of text fragment that does not match the intent information. Electronic device 501 can recognize a discourse chain (e.g., text fragment "I'm too lazy to move...", intent information "start charging the robot vacuum cleaner", and intent information "turn off the living room light") that matches a second type of text fragment (e.g., "I'm too lazy to move..."). Electronic device 501 can identify the discourse chain based on determining the similarity between the second type of text fragment (e.g., "I'm too lazy to move...") included in the discourse chain (e.g., text fragment "I'm too lazy to move...", intent information "start charging the robot vacuum cleaner", and intent information "turn off the living room light") and the second type of text fragment (e.g., "I'm too lazy to move...") identified by electronic device 501. Electronic device 501 can perform a task corresponding to the discourse chain (e.g., text fragment "I'm too lazy to move", intent information "start charging the robot vacuum cleaner", and intent information "turn off the living room light") (e.g., start charging the robot vacuum cleaner and turn off the living room light). Electronic device 501 can provide a response to the user before performing the task (e.g., "Would you like to charge the robot vacuum cleaner and turn off the living room light?").

[0130] refer to Figures 10a to 10c According to an embodiment, the discourse chain can be adaptively updated based on the user's subsequent utterances.

[0131] refer to Figure 10aAccording to an embodiment, in scenario 1001, electronic device 501 may receive input (e.g., speech) from a user (e.g., “I’m worried about the electricity bill. Turn up the air conditioning and turn off the TV.”). Electronic device 501 may convert the user’s speech (e.g., “I’m worried about the electricity bill. Turn up the air conditioning and turn off the TV.”) into text (e.g., text data) (e.g., “I’m worried about the electricity bill. Turn up the air conditioning and turn off the TV.”). Electronic device 501 may segment the text based on the probability of matching intent information. The text (e.g., “I’m worried about the electricity bill. Turn up the air conditioning and turn off the TV”) may be segmented into text fragments (e.g., “I’m worried about the electricity bill,” “Turn up the air conditioning,” and “Turn off the TV”). Electronic device 501 may analyze the text fragments (e.g., “I’m worried about the electricity bill,” “Turn up the air conditioning,” and “Turn off the TV”) to classify the text fragments (e.g., “Turn up the air conditioning” and “Turn off the TV”) into a first type of text fragment that matches the intent information, and classify the text fragments (e.g., “I’m worried about the electricity bill”) into a second type of text fragment that does not match the intent information. Electronic device 501 can generate a discourse chain (e.g., text fragment "I'm worried about the electricity bill," intent information "raise the air conditioner temperature," and intent information "turn off the TV") by pairing first intent information (e.g., intent information "raise the air conditioner temperature" and intent information "turn off the TV") corresponding to a first type of text fragment (e.g., "raise the air conditioner temperature" and "turn off the TV") with a second type of text fragment (e.g., "I'm worried about the electricity bill"). Electronic device 501 can perform a task corresponding to the first intent information (e.g., intent information "raise the air conditioner temperature" and intent information "turn off the TV") (e.g., device control operation on the air conditioner and TV as target devices). Electronic device 501 can provide a response to the user after performing the task (e.g., "The air conditioner has been turned off, and the TV has been turned off").

[0132] refer to Figure 10b According to an embodiment, in scenario 1002, electronic device 501 can receive input (e.g., speech) from a user (e.g., “I’m worried about the electricity bill…”). Text (e.g., text data) corresponding to the input (e.g., “I’m worried about the electricity bill…”) can match a speech chain (e.g., the text fragment “I’m worried about the electricity bill,” the intent information “Raise the air conditioner temperature,” and the intent information “Turn off the TV”). Electronic device 501 can provide a response to the user (e.g., “Do you want to turn off the air conditioner and the TV?”) before performing the task corresponding to the speech chain (e.g., the text fragment 'I’m worried about the electricity bill,’ the intent information 'Raise the air conditioner temperature,’ and the intent information 'Turn off the TV’).

[0133] According to an embodiment, in scenario 1003, electronic device 501 can receive subsequent input (e.g., subsequent utterances) from the user (e.g., "No need to turn off TV anymore."). Electronic device 501 can update (e.g., remove the intent information "TV off" from the utterance chain) the utterance chain (e.g., the text fragment "I'm worried about the electricity bill," the intent information "Raise the air conditioner temperature," and the intent information "TV off"). Electronic device 501 can perform a task corresponding to the updated utterance chain (e.g., the text fragment "I'm worried about the electricity bill" and the intent information "Raise the air conditioner temperature"). Electronic device 501 can provide a response to the user after performing the task (e.g., "Yes, the air conditioner temperature has been raised").

[0134] refer to Figure 10c According to an embodiment, in scenario 1004, electronic device 501 can receive input (e.g., utterances) from a user (e.g., “I’m worried about the electricity bill…”). Text (e.g., text data) corresponding to the input (e.g., utterances) (e.g., “I’m worried about the electricity bill…”) can be matched with an updated utterance chain (e.g., the text fragment “I’m worried about the electricity bill” and the intent information “Raise the air conditioning temperature”). Electronic device 501 can provide a response to the user (e.g., “Do you want to raise the air conditioning temperature?”) before performing the task corresponding to the updated utterance chain (e.g., the text fragment 'I’m worried about the electricity bill’ and the intent information 'Raise the air conditioning temperature’).

[0135] Figure 11 This is a flowchart illustrating an operation method of an electronic device according to an embodiment.

[0136] Operations 1110 to 1160 can be executed sequentially, but are not required to. For example, the order of each of operations 1110 to 1160 can be changed, and at least two operations can be executed in parallel.

[0137] According to the embodiments, operations 1110 to 1160 can be understood as being performed by an electronic device (e.g., Figure 6 The processor of the electronic device 501 (e.g., Figure 6 The processor 520 executes the commands.

[0138] In operation 1110, the electronic device according to the embodiment (e.g., Figure 5 The electronic device 501 can convert the user's speech into text.

[0139] In operation 1120, the electronic device according to the embodiment can segment text into multiple text segments including a first text segment and a second text segment.

[0140] In operation 1130, the electronic device according to the embodiment can classify a first text fragment mapped to intent information for performing a task into a first type.

[0141] In operation 1140, the electronic device according to the embodiment can classify a second text fragment that is not mapped to intent information for performing a task into a second type. Operations 1130 and 1140 are not limited to a specific order and can be executed in parallel.

[0142] In operation 1150, the electronic device according to the embodiment can perform a first task, including device control operations, on the target device based on first intent information corresponding to the first text fragment.

[0143] In operation 1160, the electronic device according to the embodiment can generate pairing information by matching first intent information with a second text fragment. When text corresponding to the second text fragment is identified in the user's second utterance, the pairing information can be used to perform the first task.

[0144] According to an embodiment, a method may include converting a user's first utterance into text. The method may include segmenting the text into multiple text segments, including a first text segment and a second text segment. The method may include classifying the first text segment, which is mapped to intent information for performing a task, into a first type. The method may include classifying a second text segment, which is not mapped to intent information for performing a task, into a second type. The method may include performing a first task, including device control operations, on a target device based on the first intent information corresponding to the first text segment. The method may include generating pairing information by pairing the first intent information with the second text segment. When text corresponding to the second text segment is identified in the user's second utterance, the pairing information can be used to perform the first task.

[0145] According to an embodiment, the method may further include: when the number of pairings between the first intent information and the second text segment exceeds a threshold, determining the pairing information as a discourse chain.

[0146] According to an embodiment, when text corresponding to a second text fragment is identified in a user's second utterance, the utterance chain can be used to perform a first task based on the first intent information included in the utterance chain.

[0147] According to an embodiment, the method may further include adaptively updating intent information included in the discourse chain based on a user's third utterance. The third utterance may include utterances in which text corresponding to a second text fragment and intent information different from the first intent information are identified. The third utterance may include a modification request for the first intent information.

[0148] According to an embodiment, the first intent information may include multiple intent information items. The first task may include multiple tasks corresponding to the multiple intent information items.

[0149] According to an embodiment, the second text fragment may include text indicating the user's situation.

[0150] According to an embodiment, the text corresponding to the second text fragment in the second discourse can be identified based on the similarity between the second text fragment and the second type of text fragment identified in the second discourse.

[0151] According to an embodiment, a method may include converting a user's utterance into text. The method may include segmenting the text into text segments that include at least one of a first text segment and a second text segment. The method may include classifying the first text segment, which is mapped to intent information for performing a task, into a first type. The method may include classifying a second text segment, which is not mapped to intent information for performing a task, into a second type. The method may include identifying a utterance chain corresponding to the second text segment. The method may include performing a first task, including device control operations, on a target device based on the first intent information included in the utterance chain. The utterance chain may be a pairing of first intent information identified in utterances preceding the utterance with second-type text segments obtained in utterances preceding the utterance.

[0152] According to an embodiment, discourse chain recognition can be performed based on determining the similarity between a second text fragment and a second type of text fragment included in the discourse chain.

[0153] According to an embodiment, the method may further include adaptively updating intent information included in the discourse chain based on subsequent utterances of the user's utterance.

[0154] According to an embodiment, subsequent utterances may include utterances in which text corresponding to a second type of text fragment and intent information different from the first intent information are identified. Subsequent utterances may include a request to modify the first intent information.

[0155] According to an embodiment, the first intent information may include multiple intent information items. The first task may include multiple tasks corresponding to the multiple intent information items.

[0156] According to an embodiment, the second text fragment may include text indicating the user's situation.

[0157] According to an embodiment, electronic devices (e.g., Figure 1 Electronic device 101 Figure 2 Electronic device 201 or Figure 5 The electronic device 501 may include a processor (e.g., Figure 1 Processor 120 Figure 2Processor 203 or Figure 5 The processor 520) and the memory for storing instructions (e.g., Figure 1 Memory 130, Figure 2 memory 207 or Figure 5 (Memory 530). When executed by the processor alone or in combination, the instructions can cause the electronic device to convert a user's first speech into text. When executed by the processor alone or in combination, the instructions can cause the electronic device to segment the text into multiple text segments, including a first text segment and a second text segment. When executed by the processor alone or in combination, the instructions can cause the electronic device to classify the first text segment mapped to intent information for performing a task into a first type. When executed by the processor alone or in combination, the instructions can cause the electronic device to classify a second text segment not mapped to intent information for performing a task into a second type. When executed by the processor alone or in combination, the instructions can cause the electronic device to perform a first task, including device control operations, on a target device based on the first intent information corresponding to the first text segment. When executed by the processor alone or in combination, the instructions can cause the electronic device to generate pairing information by pairing the first intent information with the second text segment. When text corresponding to the second text segment is identified in the user's second speech, the pairing information can be used to perform the first task.

[0158] According to an embodiment, when the instructions are executed by the processor alone or together, the electronic device can determine the pairing information as a speech chain when the number of pairings between the first intent information and the second text fragment exceeds a threshold.

[0159] According to an embodiment, when text corresponding to a second text fragment is identified in a user's second utterance, the utterance chain can be used to perform a first task based on the first intent information included in the utterance chain.

[0160] According to an embodiment, when executed individually or jointly by the processors (120; 203; 520), the instructions cause the electronic device (101; 201; 501) to adaptively update intent information included in the utterance chain based on a user's third utterance. The third utterance may include utterances in which text corresponding to a second text fragment and intent information different from the first intent information are identified. The third utterance may include a request to modify the first intent information.

[0161] According to an embodiment, the first intent information may include multiple intent information items. The task may include multiple tasks corresponding to the multiple intent information items.

[0162] According to an embodiment, the second text fragment may include text indicating the user's situation.

[0163] According to an embodiment, the text corresponding to the second text fragment in the second discourse can be identified based on the similarity between the second text fragment and the second type of text fragment identified in the second discourse.

[0164] The electronic device according to various embodiments can be one of a variety of types of electronic devices. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a home appliance. According to embodiments of this disclosure, the electronic device is not limited to those described above.

[0165] It should be understood that the various embodiments of this disclosure and the terminology used therein are not intended to limit the technical features set forth herein to the specific embodiments, but rather to include various changes, equivalents, or substitutions to the respective embodiments. Regarding the description of the drawings, similar reference numerals may be used to refer to similar or related elements. It should be understood that, unless the relevant context clearly indicates otherwise, the singular form of the noun corresponding to an item may include one or more things. As used herein, each of “A or B,” “at least one of A and B,” “at least one of A or B,” “A, B, or C,” “at least one of A, B, and C,” and “at least one of A, B, or C” may include any one of the items listed together in a corresponding phrase, or all possible combinations thereof. Terms such as “first” and “second” or “first” and “second” may be used simply to distinguish the corresponding component from other components and do not limit the component in other respects (e.g., importance or order). It will be understood that, whether the terms “operably” or “communically” are used or not, if an element (e.g., a first element) is referred to as “combined with another element (e.g., a second element),” “combined to another element (e.g., a second element),” “connected to another element (e.g., a second element),” or “connected to another element (e.g., a second element)”, it means that the element can be directly (e.g., via a wire), wirelessly connected to another element, or connected to another element via a third element.

[0166] As used in conjunction with various embodiments of this disclosure, the term "module" may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with other terms such as logic, logic block, component, or circuit. A module may be a single integrated component adapted to perform one or more functions, or its smallest unit or portion. For example, a module may be implemented as an application-specific integrated circuit (ASIC).

[0167] The various embodiments described herein can be implemented as software (e.g., a program) including one or more instructions stored in a storage medium (e.g., internal or external memory) readable by a machine (e.g., an electronic device). For example, a processor of a machine (e.g., an electronic device) can invoke and execute at least one of the one or more instructions stored in the storage medium. This allows the machine to operate to perform at least one function according to the invoked at least one instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. The term "non-transitory" simply means that the storage medium is a tangible device and does not include signals (e.g., electromagnetic waves), but this term does not distinguish between cases where data is stored semi-permanently in the storage medium and cases where data is temporarily stored in the storage medium.

[0168] According to embodiments, methods according to various embodiments of this disclosure may be included and provided in a computer program product. The computer program product can be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., an optical disc read-only memory (CD-ROM)) or via an app store (e.g., the Play Store). TM The computer program product may be distributed online (e.g., downloaded or uploaded) or directly between two user devices (e.g., smartphones). If distributed online, at least a portion of the computer program product may be temporarily generated or at least temporarily stored in a machine-readable storage medium, such as the memory of a manufacturer's server, an app store's server, or a relay server.

[0169] According to various embodiments, each of the above-described components (e.g., a module or program) may include a single entity or multiple entities, and some of the multiple entities may be arranged separately in different components. According to various embodiments, one or more of the above-described components may be omitted, or one or more other components may be added. Alternatively or additionally, multiple components (e.g., modules or programs) may be integrated into a single component. In this case, according to various embodiments, the integrated component may still perform one or more functions of each of the multiple components in the same or similar manner as the corresponding components in the multiple components before integration. According to various embodiments, operations performed by a module, program, or other component may be performed sequentially, in parallel, repeatedly, or heuristically, or one or more operations may be performed in a different order or omitted, or one or more other operations may be added.

[0170] Explanation of reference numerals in the attached figures

[0171] 501: Electronic Devices

[0172] 601: Intelligent Server

Claims

1. A method comprising: Convert the user's initial words into text; The text is divided into multiple text segments, including a first text segment and a second text segment; The first text fragment that will be mapped to intent information used to perform the task will be classified as the first type; The second text fragment that was not mapped to intent information used to perform the task is classified as the second type; Based on first intent information corresponding to the first text fragment, a first task including device control operations is performed on the target device. as well as Pairing information is generated by matching first intent information with second text fragments. In this context, when the text corresponding to the second text fragment is identified in the user's second utterance, the pairing information is used to perform the first task.

2. The method according to claim 1, further comprising: When the number of pairings between the first intent information and the second text fragment exceeds a threshold, the pairing information is identified as a discourse chain.

3. The method according to any one of claims 1 and 2, wherein, When the text corresponding to the second text fragment is identified in the user's second discourse, the discourse chain is used to perform the first task based on the first intent information included in the discourse chain.

4. The method according to any one of claims 1 to 3, further comprising: Adaptively update intent information included in the discourse chain based on the user's third utterance. The third discourse includes: The utterances in which the text corresponding to the second text fragment and the intent information different from the first intent information are identified; or A request to modify the initial intent information.

5. The method according to any one of claims 1 to 4, wherein, The first intent information includes multiple intent messages, and The first task includes multiple tasks corresponding to the multiple intent information.

6. The method according to any one of claims 1 to 5, wherein, The second text fragment includes text indicating the user's situation.

7. The method according to any one of claims 1 to 6, wherein, The identification of the text corresponding to the second text fragment in the second discourse is performed based on determining the similarity between the second text fragment and the second type of text fragment identified in the second discourse.

8. An electronic device (101; 201; 501), comprising: Processors (120; 203; 520); and Memory (130; 207; 530), store instructions, The instructions, when executed individually or jointly by the processors (120; 203; 520), cause the electronic devices (101; 201; 501) to: Convert the user's initial words into text; The text is divided into multiple text segments, including a first text segment and a second text segment; The first text fragment that will be mapped to intent information used to perform the task will be classified as the first type; The second text fragment that was not mapped to intent information used to perform the task is classified as the second type; Based on first intent information corresponding to the first text fragment, a first task, including device control operations, is performed on the target device; and Pairing information is generated by matching first intent information with second text fragments. In this context, when the text corresponding to the second text fragment is identified in the user's second utterance, the pairing information is used to perform the first task.

9. The electronic device (101; 201; 501) according to claim 8, wherein, When the instruction is executed alone or together by the processor (120; 203; 520), it causes the electronic device (101; 201; 501) to identify the pairing information as a speech chain when the number of pairings between the first intent information and the second text segment exceeds a threshold.

10. The electronic device (101; 201; 501) according to any one of claims 8 and 9, wherein, When the text corresponding to the second text fragment is identified in the user's second discourse, the discourse chain is used to perform the first task based on the first intent information included in the discourse chain.

11. The electronic device (101; 201; 501) according to any one of claims 8 to 10, wherein, When executed alone or in combination by the processors (120; 203; 520), the instructions cause the electronic devices (101; 201; 501) to: Adaptively update intent information included in the discourse chain based on the user's third utterance. The third discourse includes: The utterances in which the text corresponding to the second text fragment and the intent information different from the first intent information are identified; or A request to modify the initial intent information.

12. The electronic device (101; 201; 501) according to any one of claims 8 to 11, wherein, The first intent information includes multiple intent messages, and The first task includes multiple tasks corresponding to the multiple intent information.

13. The electronic device (101; 201; 501) according to any one of claims 8 to 12, wherein, The second text fragment includes text indicating the user's situation.

14. The electronic device (101; 201; 501) according to any one of claims 8 to 13, wherein, The identification of the text corresponding to the second text fragment in the second discourse is performed based on determining the similarity between the second text fragment and the second type of text fragment identified in the second discourse.