Method, server, and storage medium for processing voice command

The method analyzes voice commands to determine the intent of content playback and selects the appropriate output device based on content attributes and playback history, addressing the inefficiencies in existing technologies for processing voice commands for content playback.

WO2025121815A1PCT designated stage expired Publication Date: 2025-06-12SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/019481
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-20
Filing Date
2024-12-02
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Existing technologies lack an efficient method to process voice commands for playing content, particularly in determining the appropriate output device based on content attributes and playback history.

Method used

A method that analyzes voice commands to identify the intent of content playback, verifies attribute information of the content, and selects an output device capable of playing the content based on its playback history and attributes.

Benefits of technology

Enables seamless content playback by accurately identifying the intended output device, ensuring that the content is played on the most suitable device based on its attributes and playback history.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024019481_12062025_PF_FP_ABST
    Figure KR2024019481_12062025_PF_FP_ABST
Patent Text Reader

Abstract

According to one embodiment, a method for processing a voice command may be provided. The method may comprise an operation of identifying that an analysis result of the voice command indicates an intention for content playback. The method may comprise an operation of identifying at least one piece of attribute information of content associated with the voice command. The method may comprise an operation of identifying at least one output device capable of playing the content. The method may comprise an operation of identifying a first output device for playing the content among the at least one output device on the basis of at least a part of the at least one piece of attribute information of the content and playback history of at least a part of the at least one output device. The method may comprise an operation of performing at least one operation for playing the content by the first output device. Various other embodiments are possible.
Need to check novelty before this filing date? Find Prior Art

Description

Method, server, and storage medium for processing voice commands

[0001] The present disclosure relates to a method, a server, and a storage medium for processing a voice command to play content.

[0002] An AI-based voice command processing agent can perform a function associated with the intent of a voice command based on the text corresponding to the user's voice command. For example, the voice command processing agent can identify the keyword "Banana Papa" based on the text corresponding to the voice command, "Play Banana Papa," and determine that the intent is "play content (e.g., video or music)." For example, the voice agent can identify an output device capable of performing the identified intent of "playing content." The voice agent can output data that causes the identified output device to play content associated with the keyword "Banana Papa." The output device receiving the data can play content associated with the keyword "Banana Papa." In this way, content can be played based on the voice command.

[0003] The above information may be provided as background information to aid in understanding this document. None of the above is claimed to be prior art related to this document or can be used to determine prior art.

[0004] According to one embodiment, a method for processing a voice command may be provided. The method may include an operation of confirming that an analysis result of the voice command is intended to play content. The method may include an operation of confirming at least one attribute information of content associated with the voice command. The method may include an operation of confirming at least one output device capable of playing the content. The method may include an operation of confirming a first output device for playing the content among the at least one output device based on at least a portion of the at least one attribute information of the content and a playback history of at least a portion of the at least one output device. The method may include an operation of performing at least one operation for playing the content by the first output device.

[0005] According to one embodiment, a storage medium storing at least one computer-readable instruction may be provided. The at least one instruction, when executed by one or more processors including processing circuitry of the electronic device, may cause the electronic device to perform at least one operation. The at least one operation may include an operation of determining that an analysis result of a voice command is intended to play content. The at least one operation may include an operation of determining, based on determining that the analysis result of the voice command is intended to play content, at least one piece of attribute information of content associated with the voice command. The at least one operation may include an operation of determining at least one output device capable of playing the content. The at least one operation may include an operation of determining, based on at least a portion of the at least one piece of attribute information of the content and a playback history of at least a portion of the at least one output device, a first output device for playing the content from among the at least one output device. The at least one operation may include an operation of performing at least one operation for reproduction of the content by the first output device.

[0006] According to one embodiment, an electronic device may include a memory storing at least one instruction. The electronic device may include one or more processors, each processor including processing circuitry. The at least one instruction, when executed by the one or more processors, may cause the electronic device to perform at least one operation. The at least one operation may include an operation of confirming that an analysis result of a voice command is intended to play content. The at least one operation may include an operation of confirming at least one attribute information of content associated with the voice command, based on confirmation that the analysis result of the voice command is intended to play content. The at least one operation may include an operation of confirming at least one output device capable of playing the content. The at least one operation may include an operation of identifying a first output device for playing back the content among the at least one output device based on at least a portion of at least one attribute information of the content and a playback history of at least a portion of the at least one output device. The at least one operation may include an operation of performing at least one operation for playing back the content by the first output device.

[0007] According to one embodiment, a method for processing a voice command may be provided. The method may include an operation of confirming that an analysis result of a voice command provided by an electronic device reproducing first content is intended to change a playback device of the first content. The method may include an operation of confirming at least one attribute information of content associated with the voice command based on the confirmation that the analysis result of the voice command is intended to change a playback device of the first content. The method may include an operation of confirming at least one output device capable of reproducing the first content. The method may include an operation of confirming a first output device for reproducing the first content among the at least one output device based on at least a portion of the at least one attribute information of the first content and a playback history of at least a portion of the at least one output device. The method may include an operation of performing at least one operation for stopping playback of the first content by the electronic device and reproducing the first content by the first output device.

[0008] According to one embodiment, a storage medium storing at least one computer-readable instruction may be provided. The at least one instruction, when executed by one or more processors including processing circuitry of the electronic device, may cause the electronic device to perform at least one operation. The at least one operation may include an operation of determining that an analysis result of a voice command provided by the electronic device reproducing first content is intended to change a playback device of the first content. The at least one operation may include an operation of determining at least one attribute information of content associated with the voice command based on determining that the analysis result of the voice command is intended to change a playback device of the first content. The at least one operation may include an operation of determining at least one output device capable of reproducing the first content. The at least one operation may include an operation of identifying a first output device for playing back the first content among the at least one output device based on at least a portion of at least one attribute information of the first content and a playback history of at least a portion of the at least one output device. The at least one operation may include an operation of stopping playback of the first content by the electronic device and performing at least one operation for playing back the first content by the first output device.

[0009] According to one embodiment, an electronic device may include a memory storing at least one instruction. The electronic device may include one or more processors, each processor including processing circuitry. The at least one instruction, when executed by the one or more processors, may cause the electronic device to perform at least one operation. The at least one operation may include an operation of determining that an analysis result of a voice command provided by an electronic device reproducing a first content is intended to change a playback device of the first content. The at least one operation may include an operation of determining at least one attribute information of content associated with the voice command based on determining that the analysis result of the voice command is intended to change a playback device of the first content. The at least one operation may include an operation of determining at least one output device capable of reproducing the first content. The at least one operation may include an operation of identifying a first output device for playing back the first content among the at least one output device based on at least a portion of at least one attribute information of the first content and a playback history of at least a portion of the at least one output device. The at least one operation may include an operation of stopping playback of the first content by the electronic device and performing at least one operation for playing back the first content by the first output device.

[0010] FIG. 1 is a block diagram of an electronic device within a network environment, according to one embodiment.

[0011] FIG. 2 is a block diagram of a system for controlling an external electronic device according to one embodiment.

[0012] FIG. 3A is a flowchart illustrating a method for processing a voice command according to one embodiment.

[0013] FIG. 3b is a diagram illustrating a method for identifying an output device based on at least one property of content and playback history.

[0014] FIG. 3c is a diagram illustrating a method for identifying an output device based on at least one property of content and playback history.

[0015] FIG. 3D is a diagram illustrating a method for identifying an output device based on at least one property and playback history of content.

[0016] FIG. 3e is a flowchart illustrating a method for processing a voice command according to one embodiment.

[0017] FIG. 4 is a flowchart illustrating a method for processing a voice command according to one embodiment.

[0018] FIG. 5 is a flowchart illustrating a method for processing a voice command according to one embodiment.

[0019] FIG. 6A is a flowchart illustrating a method for processing a voice command according to one embodiment.

[0020] FIG. 6b is a drawing for explaining a screen for selecting one of a plurality of output device candidates according to one embodiment.

[0021] FIG. 6c is a drawing for explaining a screen for selecting one of a plurality of output device candidates according to one embodiment.

[0022] FIG. 7A is a flowchart illustrating a method for processing a voice command according to one embodiment.

[0023] FIG. 7b is a diagram illustrating a method for selecting one of a plurality of output device candidates according to one embodiment.

[0024] FIG. 8a is a flowchart illustrating a method for checking attribute information of content according to one embodiment.

[0025] FIG. 8b is a flowchart illustrating a method for checking attribute information of content according to one embodiment.

[0026] FIG. 9A is a flowchart illustrating a method for processing a voice command according to one embodiment.

[0027] FIG. 9b is a flowchart illustrating a method for processing a voice command according to one embodiment.

[0028] FIG. 9c is a flowchart illustrating a method for processing a voice command according to one embodiment.

[0029] FIG. 10A is a flowchart illustrating a method for processing a voice command according to one embodiment.

[0030] FIG. 10b is a diagram illustrating a change in an entity that plays content according to one embodiment.

[0031] FIG. 10c is a flowchart illustrating a method for processing a voice command according to one embodiment.

[0032] FIG. 11A is a flowchart illustrating a method for processing a voice command according to one embodiment.

[0033] Figure 11b is a drawing for explaining a playback history according to one embodiment.

[0034] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In connection with the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and conciseness.

[0035] FIG. 1 is a block diagram of an electronic device (101) within a network environment (100), according to one embodiment. Referring to FIG. 1 , in the network environment (100), the electronic device (101) may communicate with the electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with the electronic device (104) or a server (108) via a second network (199) (e.g., a long-range wireless communication network). In one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).

[0036] The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (120) may store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the commands or data stored in the volatile memory (132), and store result data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (121). For example, when the electronic device (101) includes the main processor (121) and the auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a given function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as a part thereof.

[0037] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, in the electronic device (101) itself where artificial intelligence is performed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.

[0038] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).

[0039] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).

[0040] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0041] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.

[0042] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.

[0043] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).

[0044] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0045] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0046] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0047] The haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. According to one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.

[0048] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.

[0049] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).

[0050] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0051] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).

[0052] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) can support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.

[0053] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas, for example, by the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device via the at least one selected antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).

[0054] In one embodiment, the antenna module (197) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.

[0055] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).

[0056] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0057] FIG. 2 is a block diagram of a system for controlling an external electronic device according to one embodiment.

[0058] According to one embodiment, the electronic device (101) may transmit and / or receive data to and from a server (108). The server (108) may include a voice command processing server (200) and / or an IoT server (240). Here, the implementation that the voice command processing server (200) and the IoT server (240) are included in a single entity, the server (108), is merely exemplary, and at least one of the voice command processing server (200) and the IoT server (240) may not be included in a single server (108). For example, the voice command processing server (200) and the IoT server (240) may both be implemented as separate entities. Alternatively, the voice command processing server (200) may be implemented to include the IoT server (240), or the voice command processing server (200) may be implemented to support the functions performed by the IoT server (240) described in the present disclosure. Alternatively, the IoT server (240) may be implemented to include the voice command processing server (200), or the IoT server (240) may be implemented to support the functions performed by the voice command processing server (200) described in the present disclosure. The server (108) may include at least one processor. The at least one processor may include a central processing unit (CPU), a graphic processing unit (GPU), an NPU, an FPGA (field programmable gate array), an ASIC, and / or a SoC (system on chip), and there is no limitation on the form of implementation thereof. For example, according to an embodiment, one operation performed by the server (108) may be performed by any one of the at least one processor (e.g., a CPU, a GPU, an NPU, an FPGA, an ASIC, and / or an SoC), or by a linkage of two or more processors.For example, according to an embodiment, the plurality of operations performed by the server (108) may be performed by any one of at least one processor (e.g., a CPU, a GPU, an NPU, an FPGA, an ASIC, and / or a SoC), or some of the plurality of operations may be performed by one processor and the remaining some may be performed by another processor. For example, the server (108) may include at least one memory that stores at least one instruction. The at least one memory may include volatile memory and / or non-volatile memory, and there is no limitation on the form of implementation thereof. The at least one instruction, when executed by the at least one processor, may cause the server (108) to perform at least one operation (e.g., at least some of the operations performed by the server (108) described in the present disclosure). An instruction that causes one operation or multiple operations to be performed by the server (108) may be stored in one physically independent memory or may be stored distributed across multiple memories.

[0059] According to one embodiment, the voice command processing server (200) may process a voice command provided from a voice command processing client (230) of an electronic device (101) and perform at least one operation corresponding to the intent of the voice command. The electronic device (101) may execute the voice command processing client (230). The electronic device (101) may acquire a user's voice through a microphone and acquire a voice command based thereon. For example, the voice command may be an acoustic signal, or a signal in which an acoustic signal has been pre-processed (e.g., noise removed and / or amplified), but there is no limitation. If the voice command processing client (230) supports a pre-processing function, it may provide the pre-processed acoustic signal to the server (108) as a voice command. When the voice command processing client (230) supports a preprocessing function and an ASR (auto speech recognition) function, the electronic device (101) can provide the text identified based on the application of the ASR function to the acoustic signal as a voice command to the voice command processing server (200). When the voice command processing client (230) supports even an NLU (natural language understanding) function, the electronic device (101) can also provide the natural language understanding result identified by applying the NLU function to the text as a voice command to the server (108). As described above, those skilled in the art will understand that the voice command processing client (230) can support at least some of a plurality of functions for voice command processing, and that the voice command provided from the electronic device (101) to the server (108) can be set based on the supported functions.Meanwhile, depending on the implementation, at least some of the functions performed by the server (108) (or at least one entity included in the server (108)) described in the present disclosure may also be performed by the electronic device (101) that receives the user's voice. Depending on the implementation, all of the operations performed by the server (108) may also be performed by the electronic device (101), which may be referred to as on-device voice command processing. Meanwhile, depending on the implementation, at least some of the functions performed by the electronic device (101) described in the present disclosure may also be performed by the server (108) (or at least one entity included in the server (108).

[0060] According to one embodiment, the voice command processing server (200) can process (e.g., ASR and / or NLU) a voice command provided from the voice command processing client (230). For example, the voice command processing server (200) can execute (or include) at least one module (211, 212, 213, 220). The ASR module (211) can provide, for example, a text corresponding to an acoustic signal. The NLU module (212) can provide a natural language understanding result corresponding to the text. The TTS module (213) can provide a signal (e.g., an acoustic signal) for a voice output corresponding to the text. For example, the voice command processing server (200) can process a voice command provided from the electronic device (101) and provide a natural language understanding result. Meanwhile, as described above, those skilled in the art will understand that the ASR module (211) and / or the NLU module (212) may not be utilized depending on the functions supported by the voice command processing client (230) executed by the electronic device (101). The natural language understanding result may include, for example, keywords and / or intents, but there is no limitation on the implementation format thereof. Keywords may include, for example, parameters and / or slots, but there is no limitation on the implementation method thereof.

[0061] The output device verification module (220) can be provided with a natural language understanding result. The output device verification module (220) can include, for example, a content property verification module (221), an output device selection module (222), a function execution command module (223), a playback history management module (224), and / or a playback history database (225), but is not limited thereto. For example, the content property verification module (221) can verify at least one property corresponding to a keyword included in the natural language understanding result. The keyword can be, for example, a word associated with the content included in the voice command (for example, a title, a genre, an artist, copyright information, a review rating, but is not limited thereto). For example, at least one property of the content can include, but is not limited to, genre, title, review rating, copyright information, and / or artist-related information. The playback history database (225) can store the playback history of at least one output device (251, 252, 253). For example, the playback history may include, but is not limited to, the playback time of the content that was played, the corresponding intent, a keyword, at least one attribute information of the content, information for identifying the output device that played the content, and / or playback management information. The output device selection module (222) can check the content attribute corresponding to the keyword provided from the content attribute verification module. The output device selection module (222) can check the playback history from the playback history DB. The output device selection module (222) can select at least one output device (251, 252, 253) capable of playing the content based on the content attribute corresponding to the keyword and the playback history from the playback history DB.The function execution command module (223) can provide a command to the electronic device (101) and / or the selected output device (e.g., the first output device (251)) (e.g., via the IoT server (240) or directly) to cause the selected output device (e.g., the first output device (251)) to play the content. In one example, the function execution command module (223) can provide parameters (or, which may be named slots) for intent and / or execution, but is not limited thereto. The function execution command module (223) can also generate a response to be provided to the user (NLG: natural language generation). The playback history management module (224) can store information on playback of the content by the selected output device in the playback history database (225) through the above-described process, and / or update information in the playback history database (225).

[0062] According to one embodiment, the IoT server (240) may transmit and / or receive data to and from at least one output device (251, 252, 253) registered in connection with an account of a user who has accessed the device (101), for example. The IoT server (240) may provide the voice command processing server (200) with, but not limited to, the current status (e.g., whether connected to the IoT server (240)) and / or capability (e.g., the type of content that can be played, but is not limited). The voice command processing server (200) may identify an output device capable of playing content based on the current status and / or capability of the at least one output device (251, 252, 253). The IoT server (240) may provide, for example, a control command to the at least one output device (251, 252, 253). For example, a part of at least one output device (251, 252, 253) may be directly or indirectly connected to the IoT server (240). If at least one output device (251, 252, 253) is indirectly connected, a control command may be provided through an intermediate device (e.g., a mobile device), and sound output may be performed by an output device (e.g., an earphone connected to the mobile device) connected to the intermediate device. The at least one output device (251, 252, 253) that receives a control command from the IoT server (240) may perform at least one operation corresponding to the control command, for example, playing content. The at least one output device (251, 252, 253) may provide the result (e.g., whether success or failure, but there is no limitation) after performing the at least one operation to the IoT server (240). The IoT server (240) may provide the result of the performance to the voice command processing server (200).The voice command processing server (200) can store the performance results in a playback history database (225). The performance results can also be used for subsequent output device selection.

[0063] FIG. 3A is a flowchart illustrating a method for processing a voice command according to one embodiment. The embodiment of FIG. 3A will be described with reference to FIGS. 3B, 3C, and 3D. FIGS. 3B, 3C, and 3D are diagrams illustrating a method for identifying an output device based on at least one attribute and playback history of content. In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.

[0064] According to one embodiment, the server (108) may, in operation 301, determine that the analysis result of the voice command is intended to play content. For example, referring to the embodiment of FIG. 3B, the server (108) may determine the voice command (350) of "Play Banana Papa." Based on the natural language understanding result for the voice command (350) of "Play Banana Papa," the server (108) may determine that the keyword of the voice command (350) is "Banana Papa" and that the intent of the voice command (350) is "Play Content."

[0065] The server (108) can, in operation 303, check at least one attribute information of the content associated with the voice command (350). The at least one attribute information may include, for example, but is not limited to, genre, title, review rating, copyright information, and / or artist-related information of the content. In one example, the server (108) can, in operation 305, check at least one output device (e.g., output device (251, 252, 253)) capable of playing the content. The server (108) can check at least one output device (e.g., output device (251, 252, 253)) capable of playing the content based on whether at least one output device registered in response to the user's account is connected to the server (108) and / or its capability (e.g., but not limited to, the type of content that can be played), but the checking method is not limited.

[0066] According to one embodiment, the server (108), in operation 307, may identify a first output device (251) (e.g., a living room speaker (355) of FIG. 3B) for playing back the content among the at least one output device based on at least a portion of at least one attribute information of the content and a playback history of at least a portion of at least one output device (251, 252, 253). The server (108) may identify, for example, a playback history (225a) such as that of FIG. 3B. For example, referring to FIG. 3B, the playback history (225a) may be arranged, for example, in the order of the playback time (331) of the content, but is not limited thereto. The playback history (225a) may include, but is not limited to, for example, a playback time (331) of the content, an intent (332), a keyword (333), at least one attribute information (334, 335, 336) of the content, information (337) for identifying an output device that played the content, and / or playback management information (338). For example, the playback time (331) of the content in the first sub-playback history (341) may be "July 22, 2023 10:00". For example, the intent (332) of the first sub-playback history (341) may be "Play music", which is an example of content playback. The intent (332) may be, but is not limited to, an intent of a natural language understanding result that caused the content to be played. For example, the keyword (333) of the first sub-playback history (341) may be "grapes." For example, the keyword (333) may be a keyword of a natural language understanding result that caused content playback, but there is no limitation. For example, the genre information (334), which is an example of attribute information of the corresponding content of the first sub-playback history (341), may be "children's song, animation." For example, the artist information (335), which is an example of attribute information of the corresponding content of the first sub-playback history (341), may be "grapes."For example, the title information (336), which is an example of the attribute information of the corresponding content of the first sub-playback history (341), may be "grape". For example, the information (337) for identifying the output device that played back the corresponding content of the first sub-playback history (341) may be "mobile". For example, the playback management information (338) of the first sub-playback history (341) may be "additional". For example, the playback management information (338) may be set as "additional" based on the start of playback of the corresponding content, but there is no limitation. For example, other sub-playback histories (342, 343, 344, 345, 346, 347) may be stored in the playback history database (225). For example, the content titled "Grapes" may be started to be played by "Mobile" at 10:00 on July 22, 2023. Meanwhile, the output device that performs the playback of the content titled "Grapes" may be changed from "Mobile" to "Living Room Speaker" at 10:01 on July 22, 2023. Accordingly, a second sub-playback history (342) may be generated and stored. The information (337) for identifying the output device of the second sub-playback history (342) may be "Living Room Speaker", and the playback management information (338) may be "Changed." Meanwhile, the playback of the content titled "Grapes" may be terminated at 10:05 on July 22, 2023. Accordingly, a third sub-playback history (343) may be generated and stored. Information (337) for identifying the output device of the third sub-playback history (343) may be “living room speaker” and playback management information (338) may be “media end”.

[0067] For example, the playback time (331) of the content of the 4th sub-playback history (344) may be "July 22, 2023 10:05". For example, the intent (332) of the 4th sub-playback history (344) may be "Play music", which is an example of content playback. The intent (332) may be, for example, the intent of the natural language understanding result that caused the content to be played, but is not limited thereto. For example, the keyword (333) of the 4th sub-playback history (344) may be "banana papa". For example, the genre information (334), which is an example of attribute information of the corresponding content of the 4th sub-playback history (344), may be "nursery song". For example, the artist information (335), which is an example of attribute information of the corresponding content of the 4th sub-playback history (344), may be "grape". For example, the title information (336), which is an example of the attribute information of the corresponding content in the fourth sub-playback history (344), may be "Banana Papa." For example, the information (337) for identifying the output device that played back the corresponding content in the fourth sub-playback history (344) may be "Mobile." For example, the playback management information (338) in the fourth sub-playback history (344) may be "Add." "Add" may mean that the playback history by the output device for the corresponding content is added, but there is no limitation. Meanwhile, "Change" may mean a change of the playback device, but there is no limitation. Meanwhile, "End" may mean the end of playback of the corresponding content, but there is no limitation. For example, the content with the title "Banana Papa" may be started to be played by "Mobile" at 10:05 on July 22, 2023. Meanwhile, the output device that plays the content titled "Banana Papa" may be changed from "Mobile" to "Living Room Speaker" at 10:06 on July 23, 2023. Accordingly, a fifth sub-playback history (345) may be created and stored.Information (337) for identifying the output device of the 5th sub-playback history (345) may be "living room speaker", and playback management information (338) may be "changed". Meanwhile, playback of the content titled "Banana Papa" may end at 10:10 on July 23, 2023. Accordingly, the 6th sub-playback history (346) may be generated and stored. Information (337) for identifying the output device of the 3rd sub-playback history (346) may be "living room speaker", and playback management information (338) may be "media end". Meanwhile, the playback time (331) of the content of the 7th sub-playback history (347) may be "10:14 on July 23, 2023." For example, the intent (332) of the 7th sub-playback history (347) may be "Play music," which is an example of content playback. For example, the keyword (333) of the 7th sub-playback history (347) may be "Rocket." For example, the genre information (334), which is an example of attribute information of the corresponding content of the 7th sub-playback history (347), may be "Rock." For example, the artist information (335), which is an example of attribute information of the corresponding content of the 7th sub-playback history (347), may be "PatentRock." For example, the title information (336), which is an example of attribute information of the corresponding content of the 7th sub-playback history (347), may be "Rocket." For example, the information (337) for identifying the output device that played the corresponding content of the 7th sub-playback history (347) may be "in-room speaker." For example, the playback management information (338) of the 7th sub-playback history (347) may be “additional”.

[0068] As described above, the server (108) can confirm the intention of "play content" and the keyword "Banana Papa" as a result of the natural language understanding of the voice command (350) of "Play Banana Papa." The server (108) can confirm the attribute information (351) of the content associated with the keyword. As will be described later, the server (108) can confirm the attribute information (351) based on, for example, a linkage with an entity that supports NER (named entity recognition) and / or an entity that manages LLM (large language model), but there is no limitation on the method of confirmation. The attribute information (351) of the content can be, for example, genre information (352) of "Children's Song", artist information (353) of "Grape", and title information (354) of "Banana Papa", but there is no limitation on the type thereof. The server (108) can confirm the sub-playback history corresponding to the attribute information (351) of the content by referring to the playback history (225a). For example, it can be confirmed that the genre information (334), the artist information (335), and the title information (336) of the sub-playback history (344, 345, 346) are identical to the genre information (352), the artist information (353), and the title information (354) of the attribute information (351) of the content. Based on the identicalness of at least one piece of attribute information, one of the output devices corresponding to the sub-playback history (341, 342, 343) as the output device can be confirmed as the "mobile" or "living room speaker." For example, when multiple output devices are confirmed, the server (108) can select one of them based on, for example, a user's selection, a specified rule, and / or an inference result of an artificial intelligence model, which will be described later. For example, a given rule may determine the corresponding device as the output device after a change in the output device for content playback.In this case, the server (108) may identify the "living room speaker" as the output device based on the fact that the "living room speaker" is the corresponding device after the content playback change, but this is merely exemplary. In the embodiment of FIG. 3B, for example, it is assumed that the living room speaker (355) is selected. Meanwhile, it may be confirmed that the genre information (334) and the artist information (335) of the sub playback history (341, 342, 343) are identical to the genre information (352), the artist information (335), and the title information (336) of the attribute information (351) of the content. In one example, based on the identity of at least one attribute information, the output device corresponding to the sub playback history (341, 342, 343) may also be identified as a candidate for the output device. Alternatively, in another example, an output device of a sub-playback history (344, 345, 346) having a relatively larger number of attribute information identical to the attribute information (351) of the content may be identified as a candidate for the output device, and an output device of a sub-playback history (341, 342, 343) having a relatively smaller number of attribute information may not be identified as a candidate for the output device, but there is no limitation. In one example, the server (108) may preferentially consider specific attribute information (e.g., genre information (352)) over other attribute information, but there is no limitation. In one example, among a plurality of output devices that have played the content, an output device that has played the content most recently may be selected.

[0069] The server (108), in operation 309, may perform at least one operation for reproduction of content by the first output device (251) (e.g., the living room speaker (355) of FIG. 3B). For example, the server (108) may provide data to the first output device (251) that causes the output of content (e.g., content having a title of “Banana Papa”). Alternatively, for example, the server (108) may provide data to cause the transmission of data for reproduction of content to another external electronic device that may access a source that provides the content (e.g., content having a title of “Banana Papa”) or that stores the content. In this case, the external electronic device may establish a communication connection with the first output device (251) based on the received data, and may provide data for reproduction of the content to the first output device (251) based on the established communication connection. The first output device (251) may also reproduce content based on data for content reproduction received based on a communication connection.

[0070] The server (108) may also update the playback history (225a) based on the completion of content playback.

[0071] Meanwhile, referring to FIG. 3c, the server (108) can confirm the voice command (360) of "play jingle bells." The server (108) can confirm attribute information (361) of content associated with the voice command (360) (or keyword). The attribute information (361) may be, for example, genre information (362) of "children's songs," artist information (363) of "XXY," and title information (364) of "jingle bells." The server (108) can confirm sub-play history (341, 342, 343, 344, 345, 346) having genre information (334) corresponding to the genre information (362) of "children's songs." The server (108) can check the output devices corresponding to the sub-play history (341, 342, 343, 344, 345, 346), which are "mobile" and "living room speaker." The server (108) can select the living room speaker (365), for example, based on a user's selection, a specified rule, and / or an inference result of an AI model. The server (108) can perform at least one operation to cause the content of the title "Jingle Bells" to be played by the living room speaker (365).

[0072] Meanwhile, referring to FIG. 3D, the server (108) can confirm the voice command (370) of "Play Hostile." The server (108) can confirm attribute information (371) of content associated with the voice command (370) (or keyword). The attribute information (371) may be, for example, genre information (372) of "Rock," artist information (373) of "Pamtera," and title information (374) of "Hostile." The server (108) can confirm sub-play history (347) having genre information (334) corresponding to the genre information (372) of "Rock." The server (108) can confirm the "room speaker," which is an output device corresponding to the sub-play history (347). The server (108) can perform at least one action that causes the content of the title “Hostile” to be played by the room speaker (375).

[0073] FIG. 3e is a flowchart illustrating a method for processing a voice command according to one embodiment.

[0074] According to one embodiment, the server (108) may, in operation 391, determine that the analysis result of the voice command is intended for content playback. In operation 393, the server (108) may determine a category (or, may be referred to as a cluster or group) corresponding to the content associated with the voice command. For example, Table 1 is an example of a category corresponding to content.

[0075]

[0076] For example, the categories in Table 1 may be set based on genre information, but this is merely exemplary, and there is no limitation on the type and / or number of attribute information for distinguishing categories. For example, the server (108) may identify a category corresponding to content based on information such as Table 1, and / or identify a category corresponding to content based on clustering, and there is no limitation on the identification method. Category-related information such as Table 1 may be generated (or identified) based on, for example, K-means clustering and / or a knowledge graph, but there is no limitation on the method. Meanwhile, those skilled in the art will understand that the classification of categories by genre as in Table 1 is merely exemplary, and that categories may be classified based on content attributes other than categories.

[0077] The server (108) can, in operation 395, identify at least one output device capable of playing the content. In operation 397, the server (108) can identify a first output device for playing the content among the at least one output device based on the category and the playback history of at least some of the at least one output device. For example, a category for each playback history can be identified based on genre information of the playback history, or the playback history can be implemented to include information about the category. The server (108) can identify the first output device that played the content of the corresponding category based on the playback history, or identify the first output device among candidates for a plurality of output devices. In operation 399, the server (108) can perform at least one operation for playing the content by the first output device.

[0078] FIG. 4 is a flowchart illustrating a method for processing a voice command according to one embodiment.

[0079] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0080] According to one embodiment, the electronic device (101) may obtain a voice command in operation 401. The electronic device (101) may provide the voice command to the server (108) in operation 403. The voice command may be implemented as an acoustic signal, a text in which the acoustic signal is transformed, and / or a natural language understanding result for the text, as described above, and there is no limitation on the implementation thereof. The NLU module (212) included in the server (108) may confirm that the voice command is intended for content playback in operation 405. Those skilled in the art will understand that if the electronic device (101) provides a voice command including a natural language understanding result, the operation by the NLU module (212) may be omitted. The NLU module (212) may provide content-related information in operation 407. The content-related information may include, for example, keywords and / or intent (e.g., intent to play content), but is not limited thereto. The output device verification module (220) can verify at least one property of the content in operation 409. For example, the output device verification module (220) can verify at least one property of the content based on an inquiry to an entity that supports NER and / or an entity that manages LLM, but is not limited thereto. The output device verification module (220) can verify a playback history of at least one output device in operation 411. The output device verification module (220) can verify a first output device based on at least a portion of at least one property of the content and at least a portion of the playback history in operation 413. For example, the first output device can be verified based on at least a portion of at least one property of the content corresponding (identical and / or similar) to at least a portion of a content property in the playback history of the first output device.

[0081] According to one embodiment, the output device verification module (220) may, in operation 415, provide data that causes content playback to the first output device (251). The output device verification module (220) may provide the data to the first output device (251), for example, via the IoT server (240). However, this is exemplary and those skilled in the art will understand that the data may be transmitted without the intermediary of the IoT server (240). For example, the first data may include at least one command, a keyword, and / or at least one attribute of the verified content to achieve the intent of “content playback,” but there is no limitation on the implementation form thereof. The first output device (251), in operation 417, may play the content based on the received first data. For example, the first output device (251) can download information for playing the content and play it by connecting to the source of the content based on information included in the first data, or can play pre-stored content. Alternatively, the first output device (251) can establish a connection with an external electronic device (e.g., the electronic device (101) or another IoT device) capable of accessing the content, and can receive and play data for playing the content based on the established connection, and there is no limitation on the method of playing the content.

[0082] FIG. 5 is a flowchart illustrating a method for processing a voice command according to one embodiment.

[0083] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0084] According to one embodiment, the electronic device (101) may obtain a voice command in operation 501. The electronic device (101) may provide the voice command to the server (108) in operation 503. The NLU module (212) included in the server (108) may confirm that the voice command is intended for content playback in operation 505. The NLU module (212) may provide content-related information in operation 507. The output device verification module (220) may confirm at least one property of the content in operation 509. The output device verification module (220) may confirm the playback history of at least one output device in operation 511. The output device verification module (220) may confirm the first output device based on at least a portion of the at least one property of the content and at least a portion of the playback history in operation 513. The output device verification module (220) may, in operation 515, provide second data that causes transmission of first data for content playback to the electronic device (101) that provided the voice command. For example, the output device verification module (220) may verify whether the first output device (251) supports establishing a communication connection with the electronic device (101), receiving data for content playback through the communication connection, and / or supporting a content playback function based on the received data, and may select the first output device (251) based on supporting at least one function (which may be referred to as a wireless audio output function). The output device verification module (220) may also select the first output device (251) by verifying whether the first output device (251) supports the wireless audio output function and additionally whether the first output device (251) can establish a communication connection with the electronic device (101).For example, the output device verification module (220) can verify whether the electronic device (101) is located in the space where the first output device (251) is placed based on information provided from the IoT server (240), and can verify whether wireless sound output is currently possible based on the verification result. Alternatively, the output device verification module (220) can receive a list of connectable devices verified based on short-range communication from the electronic device (101), and can verify whether the first output device (251) is currently capable of wireless sound output based on the list, and there is no limitation on the verification method.

[0085] The electronic device (101) may establish a communication connection with the first output device (251) based on the received second data in operation 517. Meanwhile, those skilled in the art will understand that if the electronic device (101) has already established a communication connection with the first output device (251), the procedure for establishing the communication connection may be omitted. The second data may include, for example, information for identifying the first output device (251), information for requesting establishment of a communication connection with the first output device (251), and / or information related to content for which playback is requested, but there is no limitation in its implementation. The electronic device (101) may establish a communication connection with the first output device (251) based on the received second data in operation 517. The communication connection may be established based on, for example, short-range wireless communication, but is not limited thereto, and there is no limitation in the method of establishing the communication connection. The electronic device (101) may, in operation 519, provide first data for content playback via a communication connection. For example, the electronic device (101) may play content, and by setting the output device of the played content to the first output device (251), content playback by the first output device (251) may be possible, but there is no limitation. Alternatively, the electronic device (101) may provide only data for content playback to the first output device (251) without playing the content, thereby enabling content playback by the first output device (251).

[0086] FIG. 6A is a flowchart illustrating a method for processing a voice command according to one embodiment.

[0087] The embodiment of Fig. 6a will be described with reference to Figs. 6b and 6c. Fig. 6b is a drawing for explaining a screen for selecting one of a plurality of output device candidates according to one embodiment.

[0088] FIG. 6c is a drawing for explaining a screen for selecting one of a plurality of output device candidates according to one embodiment.

[0089] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0090] According to one embodiment, the electronic device (101) may obtain a voice command in operation 601. The electronic device (101) may provide the voice command to the server (108) in operation 603. The NLU module (212) included in the server (108) may confirm that the voice command is intended for content playback in operation 605. The NLU module (212) may provide content-related information in operation 607. The output device verification module (220) may confirm at least one property of the content in operation 609. The output device verification module (220) may confirm a playback history of at least one output device in operation 611. The output device verification module (220) may identify output device candidates based on at least a portion of the at least one property and at least a portion of the playback history in operation 613. For example, as in FIG. 3b, the genre information (352) of the content property (351) may be confirmed as “children’s songs.” Accordingly, “mobile” and “living room speaker” corresponding to the sub-play histories (341, 342, 343, 344, 345, 346) containing “children’s songs” in the genre information (333) may be confirmed as output device candidates. The output device confirmation module (220) may, in operation 615, provide data for representing the output device candidates to the electronic device (101) that provided the voice command. The electronic device (101) may, in operation 617, represent the output device candidates based on the received data. For example, as in FIG. 6b, the electronic device (101) may provide a pop-up window (631) on the screen (630) being displayed. Meanwhile, the expression of the pop-up window (631) is merely exemplary, and those skilled in the art will understand that various expression methods, such as screen switching, may be possible in addition to the pop-up window (631). The pop-up window (631) may include text intended to guide the selection of an output device capable of playing the content.The pop-up window (631) may include at least one object (632, 633, 634). The at least one object (632, 633, 634) may be expressed as a UI element, a UI object, an icon, text, an image, a visual element, a visual component, an avatar, a thumbnail, an animation, a key, and / or a button, but there is no limitation on the way it is expressed. The at least one object (632, 633, 634) may include text that can identify each of the plurality of output device candidates, but there is no limitation on the way it is expressed. Meanwhile, the electronic device (101) may replace or additionally provide identification information for the plurality of output device candidates and a voice response to select one of them. A user may select (for example, but not limited to, by tapping or inputting an additional voice command) one of at least one object (632, 633, 634). The electronic device (101) may, in operation 621, provide data related to the user selection to the output device verification module (220).

[0091] Alternatively, the electronic device (101) may provide a pop-up window (631) as in FIG. 6C based on the received data. In FIG. 6C, an additional object (636) providing a countdown animation effect may be expressed on an object (632) that is one of the objects (632, 633, 634) corresponding to a plurality of output device candidates. The additional object (636) may be expressed as being placed on an object (632) corresponding to an output device having the highest priority among the output device candidates, for example. For example, in FIG. 6C, “Living Room TV,” “My tab S8,” and “My S23+” may be set as output device candidates, and since “Living Room TV” has the highest priority, the additional object (636) may be expressed as being placed on an object (632) corresponding to “Living Room TV.” The priority may be set based on a specified rule, and / or determined based on an inference result (e.g., a confidence score, but not limited to) of an artificial intelligence model (e.g., a neural network-based artificial intelligence model, linear regression, and / or a support vector machine, but not limited to). For example, the specified rule may include a rule that assigns a higher priority to a content property the more times it has been played. For example, the specified rule may include a rule that assigns a higher priority to a content property the closer to a recent time point the content property has been played. For example, the specified rule may include a rule that assigns a higher priority to an output device after the change when an output device for playing the content property has been changed. Meanwhile, there is no limitation on the type and / or number of specified rules.When there are multiple rules, the priority of each output device candidate may be determined based on the result of applying each rule (e.g., a score for setting priorities) (e.g., based on a sum or weighted sum). For example, the priority may be set based on the inference results of an artificial intelligence model for each output device candidate. The artificial intelligence model may be trained to receive, for example, at least one attribute of the output device and content as input and output a fitness (or score). For example, the artificial intelligence model may be an LLM, but there are no limitations on its implementation.

[0092] The additional object (636) may be represented, for example, by the text "3" in the embodiment of FIG. 6C. Over time, the text "3" may sequentially change to the texts "2," "1," and "0," thereby displaying a countdown animation effect. Based on the display of the text "0," the electronic device (101) may determine that the output device corresponding to the object (632) associated with the location of the additional object (636) has been selected. Accordingly, the electronic device (101) may, in operation 621, provide data related to the user selection to the output device identification module (220). Referring again to FIG. 6A, the output device identification module (220) may, in operation 623, identify the first output device (251) based on the received data. The output device verification module (220) may, in operation 625, provide data to the first output device (251) that causes content playback. The output device verification module (220) may provide the data to the first output device (251) via, for example, the IoT server (240), but it will be understood by those skilled in the art that this is exemplary and that the data may be transmitted without the intermediary of the IoT server (240). The first output device (251) may, in operation 627, play back the content based on the received data. Meanwhile, in another example, as described in FIG. 5, those skilled in the art will understand that the output device verification module (220) may also provide second data to the electronic device (101) that causes transmission of the first data for content playback.

[0093] FIG. 7A is a flowchart illustrating a method for processing a voice command according to one embodiment. The embodiment of FIG. 7A will be described with reference to FIG. 7B.

[0094] FIG. 7b is a diagram illustrating a method for selecting one of a plurality of output device candidates according to one embodiment.

[0095] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0096] According to one embodiment, the server (108) may, in operation 701, confirm that the analysis result of the voice command is intended to play content. For example, the server (108) may confirm the voice command “Play Banana Papa.” Based on the natural language understanding result for the voice command “Play Banana Papa,” the server (108) may confirm that the keyword of the voice command is “Banana Papa” and the intent of the voice command (350) is “play content.” In operation 703, the server (108) may confirm at least one attribute information of the content associated with the voice command (350). In operation 705, the server (108) may confirm at least one output device capable of playing the content. In operation 707, the server (108) can identify candidates for output devices for playing back the content among at least one output device based on at least a portion of at least one attribute information of the content and at least a portion of the playback history of at least one output device. For example, it is assumed that genre information of "Children's Song", artist information of "Pododo", and title information of "Banana Papa" are identified as at least one attribute information of the keyword "Banana Papa". For example, the server (108) can identify sub-playback information (341, 342, 343, 344, 345, 346) having genre information (731) that includes "Children's Song". For example, the server (108) can identify sub-playback information (341, 342, 343, 344, 345, 346) having artist information (732) that includes artist information of "Pododo". For example, the server (108) can check sub-playback information (344, 345, 346) having artist information (733) of "Banana Papa." The server (108) can check the mobile (734) and living room speaker (735) corresponding to the sub-playback information (341, 342, 343, 344, 345, 346) as output device candidates.

[0097] Again, referring to FIG. 7A, the server (108) may, in operation 709, identify a first output device among output device candidates based on at least one rule and / or an artificial intelligence model. In operation 711, the server (108) may perform at least one operation for playback of content by the first output device. For example, the specified rule may include a rule that assigns a higher priority the more times the corresponding content property has been played. For example, the specified rule may include a rule that assigns a higher priority the closer the point in time the corresponding content property has been played to a recent point in time. For example, the specified rule may include a rule that assigns a higher priority to the output device after the change when the output device for playback of the corresponding content property has been changed. Meanwhile, there is no limitation on the type and / or number of the specified rules. When there are multiple rules, the priority of each output device candidate may be determined based on the result of applying each rule (e.g., a score for setting priorities) (e.g., based on a sum or weighted sum). For example, referring to FIG. 7B, the server (108) may set the priority of the living room speaker (735) to be higher than that of the mobile phone (734) based on a specified rule, based on the fact that the playback subject of the content titled "Banana Papa" has changed from the mobile phone (734) to the living room speaker (735). For example, the server (108) may determine the living room speaker (735), which has a relatively high priority, as the output device to play the content. Meanwhile, as described above, the priority may also be set based on the inference result of the artificial intelligence model of each output device candidate. As described above, the server (108) may determine any one of the multiple output device candidates without the user's selection.

[0098] FIG. 8a is a flowchart illustrating a method for checking attribute information of content according to one embodiment.

[0099] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0100] According to one embodiment, the server (108) may, in operation 801, confirm that the analysis result of the voice command indicates an intention to play content. The server (108) may, in operation 803, confirm a keyword. For example, the server (108) may, based on the natural language understanding result for the voice command “Play Banana Papa,” confirm that the keyword of the voice command is “Banana Papa” and the intention of the voice command is “Play Content.” In operation 805, the server (108) may inquire about at least one attribute information corresponding to the keyword to an entity that supports named entity recognition (NER) (e.g., it may be named a named entity recognition (NER) server, but is not limited thereto). The entity that supports NER may be implemented to be included in the server (108), for example, or may be implemented as a separate entity from the server (108), and there is no limitation on how it may be implemented. In operation 807, the server (108) may receive at least one attribute information corresponding to a keyword from an entity supporting NER in response to the inquiry. For example, Table 2 is an example of a response provided from an entity supporting NER according to one embodiment. NER may be performed, for example, as at least a part of an NLU operation, or may be performed independently of an NLU operation, and there are no limitations on the implementation of the performance thereof.

[0101]

[0102] As shown in Table 2, a response to a query may include identification information ("Id") of the corresponding message. The response may include a domain related to the corresponding content (e.g., "music" in Table 2). The response may include an entity related to the corresponding content (e.g., "song" in Table 2). The response may include a keyword related to the corresponding content (e.g., "textname" in Table 2, whose value may be "banana papa"). The response may include artist information as attribute information of the corresponding content (e.g., "artist_name", "artist_Id", and "type" in Table 2, whose values ​​may be "pododo", "1703695", and "artist"). The response may include, as attribute information of the content, information about the album in which the content is contained (for example, in Table 2, it is expressed as "album_name" and "album_id", and its values ​​may be "All About Bananas" and "10269480"). The response may include, as attribute information of the content, information about the genre (for example, in Table 2, it is expressed as "genre", and its value may be "kids"). The response may include, as attribute information of the content, information about the rank (for example, in Table 2, it is expressed as "genre_rank" and "rank", and its values ​​may be "4" and "1487"). The response may include information about an entity that supports NER (or a source referenced by the entity) (e.g., “streameverything.com” in Table 2). The server (108) may extract at least one attribute information from the response and, using the at least one attribute information extracted as described above, identify an output device that will play the content.

[0103] FIG. 8b is a flowchart illustrating a method for checking attribute information of content according to one embodiment.

[0104] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0105] According to one embodiment, the server (108) may, in operation 811, determine that the analysis result of the voice command indicates that the intent is to play content. The server (108) may, in operation 813, determine a keyword. For example, the server (108) may determine that the keyword of the voice command is "banana papa" and the intent of the voice command is "play content" based on the natural language understanding result for the voice command "play banana papa." In operation 815, the server (108) may provide prompting information for inquiring about at least one attribute information corresponding to the keyword to an entity managing a large language model (LLM). The entity managing the LLM may be implemented to be included in the server (108), for example, or may be implemented as a separate entity from the server (108), and there is no limitation on the implementation method thereof. In operation 817, the server (108) may receive at least one attribute information corresponding to a keyword from an entity managing the LLM in response to the inquiry. For example, Table 3 is an example of prompting information with the LLM according to one embodiment.

[0106]

[0107] For example, Table 3 may be default prompting information, and the server (108) may insert a keyword into "OOO" of Table 3 and provide the information to the entity managing the LLM. The server (108) may receive a response corresponding to the prompting information of Table 3, for example. The prompting information may include, for example, a sentence that inquires about at least one attribute information (e.g., genre, artist, title, author information, and review information). The prompting information may also include, for example, a sentence that requests a response in a format that can be processed without performing additional NLU on the server (108).

[0108] FIG. 9A is a flowchart illustrating a method for processing a voice command according to one embodiment.

[0109] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0110] According to one embodiment, the server (108) may obtain a voice command in operation 901. The server (108) may verify at least one attribute of the content in operation 903. The server (108) may verify at operation 905 whether an output device having a playback history corresponding to at least one attribute of the content exists. Based on the existence of an output device having a playback history corresponding to at least one attribute of the content (905-Yes), the server (108) may perform at least one operation for playback of the content by the verified output device in operation 907. As described above, those skilled in the art will understand that, based on the verification that there are multiple output device candidates corresponding to at least one attribute of the content, the server (108) may select any one of the output devices based on a user selection, a specified rule, and / or an artificial intelligence inference result. Based on the absence of an output device having a playback history corresponding to at least one attribute of the content (905-No), the server (108) may, in operation 909, perform at least one operation for playback of the content by the electronic device (101) that acquired the user voice. For example, the server (108) may cause the electronic device (101) to play the content by providing the electronic device (101) with at least one command for playback of the identified content. Here, the playback of the content by the electronic device (101) may not only mean playback of the content by the audio output module (155) and / or the display module (160) included in the electronic device (101), but may also mean playback of the content by an audio output device (for example, but not limited to, wired earphones, TWS (true wireless stereo) earphones, and Bluetooth-based speakers) connected to the electronic device (101).The server (108) may also update the playback history by adding the sub-playback history in which content is played by the electronic device (101) to the existing playback history.

[0111] FIG. 9b is a flowchart illustrating a method for processing a voice command according to one embodiment.

[0112] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0113] According to one embodiment, the server (108) may obtain a voice command in operation 921. The server (108) may verify at least one attribute of the content in operation 923. The server (108) may verify at operation 925 whether an output device having a playback history corresponding to at least one attribute of the content exists. Based on the existence of an output device having a playback history corresponding to at least one attribute of the content (925-Yes), the server (108) may verify at operation 927 whether the verified output device is capable of playing the content. For example, if the verified output device is not connected to the server (108) (e.g., is offline), data for playing the content may not be provided to the output device. Alternatively, if the verified output device is currently playing other content, data for playing the content may not be provided to the output device. Meanwhile, the offline state or the state of playing other content described above are merely exemplary, and there is no limitation on examples of states in which the output device cannot play the content. Based on the identified output device being capable of playing the content (927 - Yes), the server (108) may, in operation 929, perform at least one operation for playing the content by the identified output device. As described above, based on the existence of multiple output device candidates corresponding to at least one attribute of the content, those skilled in the art will understand that the server (108) may select any one of the output devices based on user selection, a specified rule, and / or an artificial intelligence inference result. Based on the identified output device being incapable of playing the content (927 - No), the server (108) may, in operation 931, perform at least one operation for playing the content by another output device corresponding to the identified output device.For example, the server (108) (e.g., IoT server (240)) can manage the arrangement location of the output device. For example, if the identified output device is not currently capable of playing content, the server (108) can perform at least one operation for playing the content by an output device arranged in the same space as the identified output device. For example, if the identified output device is not currently capable of playing content, the server (108) can perform at least one operation for playing the content by an output device with the next highest priority in relation to the identified output device. For example, if the identified output device is not currently capable of playing content, the server (108) can inquire of the user about other output device candidates and, based on the user's selection, perform at least one operation for playing the content by the selected output device. Based on the absence of an output device having a playback history corresponding to at least one attribute of the content (925-No), the server (108) may, in operation 933, perform at least one operation for playback of the content by the electronic device (101) that acquired the user voice.

[0114] FIG. 9c is a flowchart illustrating a method for processing a voice command according to one embodiment.

[0115] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0116] According to one embodiment, the server (108) may obtain a voice command in operation 941. The server (108) may verify at least one property of the content in operation 943. The server (108) may verify at operation 945 whether an output device having a playback history corresponding to at least one property of the content exists. Based on the existence of an output device having a playback history corresponding to at least one property of the content (945-Yes), the server (108) may verify at operation 947 whether the verified output device is already playing another content. The server (108) may verify at operation 949 whether the verified output device can change the playback content. For example, whether the verified output device can change the playback content may be set in advance by the user. For example, the server (108) may additionally ask the user whether to change the playback content, and based on the response, determine whether the identified output device can change the playback content. Based on the identified output device being able to change the playback content (949 - Yes), the server (108) may, in operation 951, perform at least one operation for changing the playback of the content by the identified output device. The server (108) may, for example, command to stop playback of the existing content and play the identified content, or may command only playback of the content without issuing a stop command. Based on the identified output device being unable to change the playback content (949 - No), the server (108) may, in operation 953, perform at least one operation for playback of the content by another output device corresponding to the identified output device. The other output device corresponding to the identified output device has been described with reference to FIG. 9B, and therefore, the description thereof will not be repeated here.Based on the absence of an output device having a playback history corresponding to at least one attribute of the content (945-No), the server (108) may, in operation 955, perform at least one operation for playback of the content by the electronic device (101) that acquired the user voice.

[0117] FIG. 10A is a flowchart illustrating a method for processing a voice command according to one embodiment. The embodiment of FIG. 10A will be described with reference to FIG. 10B.

[0118] FIG. 10b is a diagram illustrating a change in an entity that plays content according to one embodiment.

[0119] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0120] According to one embodiment, the server (108) may, in operation 1001, determine that the analysis result of the voice command provided from the electronic device (101) is a change in the content playback output device. For example, referring to FIG. 10B, the electronic device (101) may play the first content (e.g., title: Banana Papa, artist: Grape Do). The electronic device (101) may provide a screen (1030) indicating that the first content is played (e.g., may be an application execution screen for content playback, but is not limited thereto). The screen (1030) may include, for example, information about the title of the first content (1031) and / or information about the artist of the first content (1032), but there is no limitation in the implementation thereof. While playing the first content, the electronic device (101) may acquire a user voice (1042) from the user (1041). For example, the user voice (1042) may be "Play Banana Papa on another device." The electronic device (101) may provide a voice command corresponding to the user voice (1042) to the voice command processing server (200). The voice command may include an acoustic signal, text, and / or a natural language understanding result corresponding to the user voice (1042), as described above, but there is no limitation on the form of its implementation. The server (108) (e.g., the voice processing server (200)) may determine that the intent of the voice command is "change the content playback output device" and the keyword may be "Banana Papa."

[0121] Again, referring to FIG. 10A, the server (108) may, in operation 1003, identify at least one attribute information of the content associated with the voice command. The server (108), in operation 1005, may identify at least one output device capable of playing the content. In operation 1007, the server (108) may identify a first output device (251) for playing the content among the at least one output device based on at least a portion of the at least one attribute information of the content and a playback history of at least a portion of the at least one output device. In operation 1009, the server (108) may perform at least one operation for playing the content by the first output device (251). For example, in the example of FIG. 10B, the server (108) (e.g., the IoT server (240)) may provide data that causes the content to be played to the first output device (251). In this case, the server (108), depending on the implementation, may instruct the electronic device (101) to stop playing the content. Meanwhile, the provision of data to the first output device (251) is exemplary. As described above, the server (108) may provide data to the electronic device (101) that causes transmission of data for playing the content, and the electronic device (101) may transmit data for playing the content to the first output device (251) based on this. For example, the first output device (251) may be implemented to play the content from the beginning, but this is exemplary. The first output device (251) may be implemented to play the content from the point in time at which the electronic device (101) stopped playing, for example.

[0122] FIG. 10c is a flowchart illustrating a method for processing a voice command according to one embodiment.

[0123] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0124] According to one embodiment, the server (108) may, in operation 1041, confirm an event of changing a playback output device for content being played on the electronic device (101). Confirmation of the intention to change the playback output device in the voice command described in FIG. 10A may be an example of an event. For example, the event may be set as the electronic device (101) entering a designated location. For example, a user may set the content being output to be played by another output device based on the electronic device (101) entering the user's home. Accordingly, the occurrence of the event may be notified to the server (108) based on confirmation that the electronic device (101) has entered the designated location. Meanwhile, the output device may be determined based on attribute information of the content being played on the electronic device (101). For example, in terms of protecting the user's privacy, the output device needs to be determined based on the review rating, which is one of the attribute information of the content. Content with an adult rating is preferably not played on an output device that can be viewed by minors, and may be preferably played on an output device set for personal use. The server (108) may, in operation 1043, check at least one attribute information of the content associated with the voice command. The server (108), in operation 1045, may check at least one output device capable of playing the content. In operation 1047, the server (108) may check a first output device (251) for playing the content among the at least one output device based on at least a portion of the at least one attribute information of the content and a playback history of at least a portion of the at least one output device. In operation 1049, the server (108) may perform at least one operation for playing the content by the first output device (251).

[0125] FIG. 11A is a flowchart illustrating a method for processing a voice command according to one embodiment. The embodiment of FIG. 11A will be described with reference to FIG. 11B.

[0126] Figure 11b is a drawing for explaining a playback history according to one embodiment.

[0127] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0128] According to one embodiment, the server (108) may, in operation 1101, confirm that the analysis result of the voice command is intended for content playback. In operation 1103, the server (108) may confirm a first review rating as at least one attribute information of the content associated with the voice command. In operation 1105, the server (108) may confirm an output device having a playback history associated with the confirmed first review rating among at least one output device. In operation 1107, the server (108) may perform at least one operation for playback of the content by the confirmed output device. For example, the server (108) may manage a playback history (1125) as in FIG. 11B . The playback history (1125) may be arranged, for example, in the order of the playback time (1131) of the content, but is not limited thereto. The playback history (1125) may include, but is not limited to, for example, the playback time (1131) of the content, keywords (1132), at least one attribute information (1133, 1134, 1135, 1136) of the content, information for identifying the output device that played the content (1137), and / or playback management information (1138). Compared to the playback history (225a) described in FIG. 3b, the attribute information of the content in the playback history (1125) may further include a review rating (1136). The server (108) may confirm, for example, that the review rating of the sub-playback history (1141, 1142, 1143) is "15 years old" and that the review rating of the sub-playback history (1144, 1145, 1146) is "18 years old". The server (108) can confirm that the output devices corresponding to the sub-play history (1141, 1142, 1143) corresponding to the deliberation rating of, for example, "15 years old" are "mobile" and "projector." The server (108) can confirm that the output devices corresponding to the sub-play history (1144, 1145, 1146) corresponding to the deliberation rating of "18 years old" are "mobile" and "home TV."The server (108) can confirm, for example, a voice command such as "Play Horror Mansion." The server (108) can confirm that the keyword of the voice command is "Horror Mansion" and the intent is "Play Content." The server (108) can confirm at least one attribute information of the content associated with the keyword "Horror Mansion." For example, the rating of "Horror Mansion" may be "18 years old" as at least one attribute information. Based on the rating of "Horror Mansion" being "18 years old," the server (108) can refer to sub-play histories (1144, 1145, 1146) having a rating of "18 years old." The server (108) can confirm "mobile" and "TV in my room," which are output devices of the sub-play histories (1144, 1145, 1146), as candidates for output devices. For example, the server (108) may identify "My Room TV" as an output device among the candidates based on a specified rule (e.g., a rule that gives higher priority to a device after changing content playback). In one example, the server (108) may give (or manage) a higher priority to the attribute information of "review rating" than other attribute information. For example, the server (108) may be configured to primarily select an output device based on the "review rating" and, if no output device is identified based on the "review rating" and / or if multiple output devices are identified, to select an output device based on other attribute information, but this is exemplary. The server (108) may also, for example, give a relatively higher priority to attribute information other than the "review rating". Alternatively, the server (108) may be configured to assign weights to each attribute information and select an output device based on the sum of the weights confirmed based on whether the attribute information of the confirmed content corresponds to the attribute information in the playback history.In this case, weights corresponding to each attribute information may also be determined based on the priority of the attribute information, but there is no limitation. Alternatively, before playing content with a "moderate age rating" of "18+," the server (108) may inquire of the user whether to play the content through the identified output device. If the response to the inquiry is positive, the server (108) may be configured to perform at least one action for playing the content through the identified output device.

[0129] For example, the server (108) may manage public output devices and personal output devices separately. The server (108) may select a personal output device when the content rating is a designated rating (e.g., adult), and may select a public output device when the content rating is not a designated rating. Those skilled in the art will understand that the public output devices and personal output devices may be managed separately, for example, based on user settings or based on analysis results regarding their placement.

[0130] According to one embodiment, a method for processing a voice command may be provided. The method may include an operation of confirming that an analysis result of the voice command is intended to play content. The method may include an operation of confirming at least one attribute information of content associated with the voice command. The method may include an operation of confirming at least one output device capable of playing the content. The method may include an operation of confirming a first output device for playing the content among the at least one output device based on at least a portion of the at least one attribute information of the content and a playback history of at least a portion of the at least one output device. The method may include an operation of performing at least one operation for playing the content by the first output device.

[0131] According to one embodiment, the operation of identifying a first output device for playing back the content among the at least one output device may include an operation of identifying, based on at least a portion of the playback history, at least one first content played back by the first output device among the at least one output device and at least one piece of first attribute information corresponding to each of the at least one piece of first content. The operation of identifying a first output device for playing back the content among the at least one output device may include an operation of identifying the first output device as an output device for playing back the content based on the identity of at least a portion of the at least one piece of first attribute information and at least a portion of the at least one piece of attribute information of the content.

[0132] According to one embodiment, the act of performing at least one operation for reproduction of the content by the first output device may include the act of providing, to the first output device, first data that causes reproduction of the content.

[0133] According to one embodiment, the act of performing at least one operation for reproduction of the content by the first output device may include the act of providing second data to an external electronic device capable of communicating with the first output device, the second data causing transmission of data for reproduction of the content.

[0134] According to one embodiment, the operation of identifying a first output device for playing back the content among the at least one output device may include identifying candidates of a plurality of output devices based on at least a portion of the at least one attribute information of the content and at least a portion of the playback history. The operation of identifying a first output device for playing back the content among the at least one output device may include identifying the first output device as an output device for playing back the content based on identifying information about a user's selection of the first output device among the plurality of output devices.

[0135] According to one embodiment, the operation of identifying a first output device for playing back the content among the at least one output device may include identifying candidates of a plurality of output devices based on at least a portion of the at least one attribute information of the content and the playing history. The operation of identifying a first output device for playing back the content among the at least one output device may include identifying the first output device among the plurality of output devices as an output device for playing back the content based on at least one rule and / or an artificial intelligence model.

[0136] According to one embodiment, at least one attribute information of the content may include a genre of the content, a title of the content, a review rating of the content, authorship information of the content, and / or artist-related information of the content.

[0137] According to one embodiment, the operation of verifying at least one attribute information of the content associated with the voice command may include an operation of inquiring about the at least one attribute information of the content to an entity that supports named entity recognition (NER). The operation of verifying the at least one attribute information of the content associated with the voice command may include an operation of verifying the at least one attribute information provided from an entity that supports NER.

[0138] According to one embodiment, the operation of verifying at least one attribute information of the content associated with the voice command may include providing prompting information for inquiring about the at least one attribute information of the content to an entity managing a large language model (LLM). The operation of verifying at least one attribute information of the content associated with the voice command may include verifying the at least one attribute information provided from the entity managing the LLM.

[0139] According to one embodiment, the method may further include receiving, from an entity managing the at least one output device, information for identifying the at least one output device and / or the playback history of the at least one output device. The operation of identifying the at least one output device capable of playing the content may include an operation of identifying the at least one output device based on the information for identifying the at least one output device.

[0140] According to one embodiment, the operation of identifying a first output device among the at least one output device for playing back the content may include an operation of identifying, based on at least a portion of the playback history, at least one second content played back by a second output device, different from the first output device among the at least one output device, and at least one piece of second attribute information corresponding to each of the at least one second content. The operation of identifying a first output device among the at least one output device for playing back the content may include an operation of identifying whether the second output device can play back the content based on the identity of at least a portion of the at least one piece of second attribute information and at least a portion of the at least one piece of attribute information of the content. The operation of identifying a first output device among the at least one output device for playing back the content may include an operation of identifying the first output device corresponding to the second output device as an output device for playing back the content based on the identification that the second output device cannot play back the content.

[0141] According to one embodiment, the operation of identifying a first output device among the at least one output device for playing back the content may include an operation of identifying, based on the playback history, at least one first content played back by the first output device among the at least one output device and at least one piece of first attribute information corresponding to each of the at least one first content. The operation of identifying a first output device among the at least one output device for playing back the content may include an operation of identifying whether the first output device is playing back other content based on whether at least a portion of the at least one piece of first attribute information and at least a portion of the at least one piece of attribute information of the content are identical. The operation of identifying a first output device among the at least one output device for playing back the content may include an operation of identifying whether the first output device can stop playing back the other content and play back the content based on identifying that the first output device is playing back the other content. The operation of identifying a first output device for playing back the content among the at least one output device may include an operation of identifying the first output device as an output device for playing back the content based on the first output device being identified as being capable of stopping playing back the other content and playing back the content.

[0142] According to one embodiment, a storage medium storing at least one computer-readable instruction may be provided. The at least one instruction, when executed by one or more processors including processing circuitry of the electronic device, may cause the electronic device to perform at least one operation. The at least one operation may include an operation of determining that an analysis result of a voice command is intended to play content. The at least one operation may include an operation of determining, based on determining that the analysis result of the voice command is intended to play content, at least one piece of attribute information of content associated with the voice command. The at least one operation may include an operation of determining at least one output device capable of playing the content. The at least one operation may include an operation of determining, based on at least a portion of the at least one piece of attribute information of the content and a playback history of at least a portion of the at least one output device, a first output device for playing the content from among the at least one output device. The at least one operation may include an operation of performing at least one operation for reproduction of the content by the first output device.

[0143] According to one embodiment, the operation of identifying a first output device for playing back the content among the at least one output device may include an operation of identifying, based on at least a portion of the playback history, at least one first content played back by the first output device among the at least one output device and at least one piece of first attribute information corresponding to each of the at least one piece of first content. The operation of identifying a first output device for playing back the content among the at least one output device may include an operation of identifying the first output device as an output device for playing back the content based on the identity of at least a portion of the at least one piece of first attribute information and at least a portion of the at least one piece of attribute information of the content.

[0144] According to one embodiment, the act of performing at least one operation for reproduction of the content by the first output device may include the act of providing, to the first output device, first data that causes reproduction of the content.

[0145] According to one embodiment, the act of performing at least one operation for reproduction of the content by the first output device may include the act of providing second data to an external electronic device capable of communicating with the first output device, the second data causing transmission of data for reproduction of the content.

[0146] According to one embodiment, the operation of identifying a first output device for playing back the content among the at least one output device may include identifying candidates of a plurality of output devices based on at least a portion of the at least one attribute information of the content and at least a portion of the playback history. The operation of identifying a first output device for playing back the content among the at least one output device may include identifying the first output device as an output device for playing back the content based on identifying information about a user's selection of the first output device among the plurality of output devices.

[0147] According to one embodiment, the operation of identifying a first output device for playing back the content among the at least one output device may include identifying candidates of a plurality of output devices based on at least a portion of the at least one attribute information of the content and the playing history. The operation of identifying a first output device for playing back the content among the at least one output device may include identifying the first output device among the plurality of output devices as an output device for playing back the content based on at least one rule and / or an artificial intelligence model.

[0148] According to one embodiment, at least one attribute information of the content may include a genre of the content, a title of the content, a review rating of the content, authorship information of the content, and / or artist-related information of the content.

[0149] According to one embodiment, the operation of verifying at least one attribute information of the content associated with the voice command may include an operation of inquiring about the at least one attribute information of the content to an entity supporting NER. The operation of verifying the at least one attribute information of the content associated with the voice command may include an operation of verifying the at least one attribute information provided from the entity supporting NER.

[0150] According to one embodiment, the operation of checking at least one attribute information of the content associated with the voice command may include providing prompting information for inquiring about the at least one attribute information of the content to an entity managing an LLM. The operation of checking at least one attribute information of the content associated with the voice command may include checking the at least one attribute information provided from the entity managing the LLM.

[0151] According to one embodiment, the at least one operation may further include receiving, from an entity managing the at least one output device, information for identifying the at least one output device and / or the playback history of the at least one output device. The operation of identifying the at least one output device capable of playing the content may include identifying the at least one output device based on the information for identifying the at least one output device.

[0152] According to one embodiment, the operation of identifying a first output device among the at least one output device for playing back the content may include an operation of identifying, based on at least a portion of the playback history, at least one second content played back by a second output device, which is different from the first output device among the at least one output device, and at least one piece of second attribute information corresponding to each of the at least one second content. The operation of identifying a first output device among the at least one output device for playing back the content may include an operation of identifying whether the second output device can play back the content based on the identity of at least a portion of the at least one piece of second attribute information and at least a portion of the at least one piece of attribute information of the content. The operation of identifying a first output device among the at least one output device for playing back the content may include an operation of identifying the first output device corresponding to the second output device as an output device for playing back the content based on the identification that the second output device cannot play back the content.

[0153] According to one embodiment, the operation of identifying a first output device among the at least one output device for playing back the content may include an operation of identifying, based on the playback history, at least one first content played back by the first output device among the at least one output device and at least one piece of first attribute information corresponding to each of the at least one first content. The operation of identifying a first output device among the at least one output device for playing back the content may include an operation of identifying whether the first output device is playing back other content based on whether at least a portion of the at least one piece of first attribute information and at least a portion of the at least one piece of attribute information of the content are identical. The operation of identifying a first output device among the at least one output device for playing back the content may include an operation of identifying whether the first output device can stop playing back the other content and play back the content based on identifying that the first output device is playing back the other content. The operation of identifying a first output device for playing back the content among the at least one output device may include an operation of identifying the first output device as an output device for playing back the content based on the first output device being identified as being capable of stopping playing back the other content and playing back the content.

[0154] According to one embodiment, an electronic device may include a memory storing at least one instruction. The electronic device may include one or more processors, each processor including processing circuitry. The at least one instruction, when executed by the one or more processors, may cause the electronic device to perform at least one operation. The at least one operation may include an operation of confirming that an analysis result of a voice command is intended to play content. The at least one operation may include an operation of confirming at least one attribute information of content associated with the voice command, based on confirmation that the analysis result of the voice command is intended to play content. The at least one operation may include an operation of confirming at least one output device capable of playing the content. The at least one operation may include an operation of identifying a first output device for playing back the content among the at least one output device based on at least a portion of at least one attribute information of the content and a playback history of at least a portion of the at least one output device. The at least one operation may include an operation of performing at least one operation for playing back the content by the first output device.

[0155] According to one embodiment, a method for processing a voice command may be provided. The method may include an operation of confirming that an analysis result of a voice command provided by an electronic device reproducing first content is intended to change a playback device of the first content. The method may include an operation of confirming at least one attribute information of content associated with the voice command based on the confirmation that the analysis result of the voice command is intended to change a playback device of the first content. The method may include an operation of confirming at least one output device capable of reproducing the first content. The method may include an operation of confirming a first output device for reproducing the first content among the at least one output device based on at least a portion of the at least one attribute information of the first content and a playback history of at least a portion of the at least one output device. The method may include an operation of performing at least one operation for stopping playback of the first content by the electronic device and reproducing the first content by the first output device.

[0156] According to one embodiment, a storage medium storing at least one computer-readable instruction may be provided. The at least one instruction, when executed by one or more processors including processing circuitry of the electronic device, may cause the electronic device to perform at least one operation. The at least one operation may include an operation of determining that an analysis result of a voice command provided by the electronic device reproducing first content is intended to change a playback device of the first content. The at least one operation may include an operation of determining at least one attribute information of content associated with the voice command based on determining that the analysis result of the voice command is intended to change a playback device of the first content. The at least one operation may include an operation of determining at least one output device capable of reproducing the first content. The at least one operation may include an operation of identifying a first output device for playing back the first content among the at least one output device based on at least a portion of at least one attribute information of the first content and a playback history of at least a portion of the at least one output device. The at least one operation may include an operation of stopping playback of the first content by the electronic device and performing at least one operation for playing back the first content by the first output device.

[0157] According to one embodiment, an electronic device may include a memory storing at least one instruction. The electronic device may include one or more processors, each processor including processing circuitry. The at least one instruction, when executed by the one or more processors, may cause the electronic device to perform at least one operation. The at least one operation may include an operation of determining that an analysis result of a voice command provided by an electronic device reproducing a first content is intended to change a playback device of the first content. The at least one operation may include an operation of determining at least one attribute information of content associated with the voice command based on determining that the analysis result of the voice command is intended to change a playback device of the first content. The at least one operation may include an operation of determining at least one output device capable of reproducing the first content. The at least one operation may include an operation of identifying a first output device for playing back the first content among the at least one output device based on at least a portion of at least one attribute information of the first content and a playback history of at least a portion of the at least one output device. The at least one operation may include an operation of stopping playback of the first content by the electronic device and performing at least one operation for playing back the first content by the first output device.

[0158] Electronic devices according to the embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments disclosed in this document are not limited to the aforementioned devices.

[0159] The embodiments of this document and the terminology used herein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.

[0160] The term "module" used in the embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0161] One embodiment of the present document may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.

[0162] According to one embodiment, the method according to one embodiment disclosed in the present document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0163] According to one embodiment, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and arranged in other components. According to one embodiment, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to one embodiment, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

1. In the method of processing voice commands, An action that verifies that the analysis of a voice command is intended to play content; An operation of checking at least one attribute information of content associated with the voice command based on determining that the analysis result of the voice command is intended for content playback; An action of identifying at least one output device capable of playing the above content; An operation of identifying a first output device for playing back the content among the at least one output device, based on at least a portion of at least one attribute information of the content and at least a portion of a playback history of the at least one output device; and An operation for performing at least one operation for reproduction of said content by said first output device. A method including:

2. In paragraph 1, The operation of identifying a first output device for playing back the content among the at least one output device is: An operation of identifying at least one first content played by said first output device among said at least one output device and at least one first attribute information corresponding to each of said at least one first content, based on at least a part of said playback history; and An operation of confirming the first output device as an output device for playing back the content based on at least a part of the at least one first attribute information and at least a part of the at least one attribute information of the content being identical. A method including:

3. In any one of paragraphs 1 and 2, An operation of performing at least one operation for reproduction of said content by said first output device is: A method comprising the action of providing first data to said first output device, said first data causing reproduction of said content.

4. In any one of paragraphs 1 to 3, An operation of performing at least one operation for reproduction of said content by said first output device is: A method comprising the action of providing second data to an external electronic device capable of communicating with said first output device, said second data causing transmission of data for reproduction of said content.

5. In any one of paragraphs 1 to 4, The operation of identifying a first output device for playing back the content among the at least one output device is: An operation of identifying candidates for a plurality of output devices based on at least a portion of the at least one attribute information of the content and at least a portion of the playback history; and An operation of confirming the first output device as an output device for playing the content based on information about a user's selection of the first output device among the plurality of output devices. A method including:

6. In any one of paragraphs 1 to 5, The operation of identifying a first output device for playing back the content among the at least one output device is: An operation of identifying candidates of a plurality of output devices based on at least a portion of at least one attribute information of the content and the playback history; and An operation of identifying the first output device among the plurality of output devices as an output device for playing the content based on at least one rule and / or artificial intelligence model. A method including:

7. In any one of paragraphs 1 to 6, A method wherein at least one attribute information of the content includes a genre of the content, a title of the content, a review rating of the content, authorship information of the content, and / or artist-related information of the content.

8. In any one of paragraphs 1 to 7, An operation of checking at least one attribute information of content associated with the above voice command is: An operation of querying an entity supporting NER (named entity recognition) for at least one attribute information of the content; and An action to verify at least one attribute information provided from an entity supporting the above NER. A method including:

9. In any one of paragraphs 1 to 8, An operation of checking at least one attribute information of content associated with the above voice command is: An operation of providing prompting information for querying at least one attribute information of said content to an entity managing a large language model (LLM); and An action to verify at least one attribute information provided from an entity managing the LLM. A method including:

10. In any one of paragraphs 1 to 9, An operation of receiving, from an entity managing at least one output device, information for identifying the at least one output device and / or the playback history of the at least one output device. Including, The operation of checking at least one output device capable of playing the above content is: A method comprising an action of identifying said at least one output device based on information for identifying said at least one output device.

11. In any one of paragraphs 1 to 10, The operation of identifying a first output device for playing back the content among the at least one output device is: An operation of identifying at least one second content played by a second output device different from the first output device among the at least one output device based on at least a portion of the playback history and at least one second attribute information corresponding to each of the at least one second content; An operation of determining whether the second output device can reproduce the content based on at least a part of the at least one second attribute information and at least a part of the at least one attribute information of the content being identical; and An operation of confirming the first output device corresponding to the second output device as an output device for reproducing the content based on confirmation that the second output device cannot reproduce the content. A method including:

12. In any one of paragraphs 1 to 11, The operation of identifying a first output device for playing back the content among the at least one output device is: An operation of checking at least one first content played by the first output device among the at least one output device based on the above playback history and at least one first attribute information corresponding to each of the at least one first content; An operation of determining whether the first output device is playing back another content based on at least a portion of the at least one first attribute information and at least a portion of the at least one attribute information of the content being identical; An operation of determining whether the first output device can stop playing the other content and play the content based on determining that the first output device is playing the other content; and An operation of confirming the first output device as an output device for reproducing the content, based on the confirmation that the first output device is capable of stopping reproduction of the other content and reproducing the content. A method including:

13. A storage medium storing at least one computer-readable instruction, wherein the at least one instruction, when executed by one or more processors including processing circuitry of an electronic device (108), causes the electronic device to perform at least one operation, At least one of the above actions: An action that verifies that the analysis of a voice command is intended to play content; An operation of confirming at least one attribute information of content associated with the voice command, based on determining that the analysis result of the voice command is intended to play content; An action of identifying at least one output device capable of playing the above content; An operation of identifying a first output device for playing back the content among the at least one output device, based on at least a portion of at least one attribute information of the content and at least a portion of a playback history of the at least one output device; and An operation for performing at least one operation for reproduction of said content by said first output device. A storage medium containing.

14. In paragraph 13, The operation of identifying a first output device for playing back the content among the at least one output device is: An operation of checking at least one first content played by the first output device among the at least one output device based on the playback history and at least one first attribute information corresponding to each of the at least one first content; and An operation of confirming the first output device as an output device for playing back the content based on at least a part of the at least one first attribute information and at least a part of the at least one attribute information of the content being identical. A storage medium containing.

15. In an electronic device (108), memory for storing at least one instruction; and comprising one or more processors, including processing circuitry; wherein said at least one instruction, when executed by said one or more processors, causes said electronic device to perform at least one operation; At least one of the above actions: An action that verifies that the analysis of a voice command is intended to play content; An operation of checking at least one attribute information of content associated with the voice command based on determining that the analysis result of the voice command is intended for content playback; An action of identifying at least one output device capable of playing the above content; An operation of identifying a first output device for playing back the content among the at least one output device, based on at least a portion of at least one attribute information of the content and at least a portion of a playback history of the at least one output device; and An operation for performing at least one operation for reproduction of said content by said first output device. An electronic device comprising:

Citation Information

Patent Citations

  • Hotword detection on multiple devices

    KR101832648B1

  • Intelligent device arbitration and control

    KR1020180133526A

  • Voice control method for a media playback system

    KR102095250B1

  • Systems and methods for routing content to associated output devices

    KR102360589B1

  • Audio response playback

    KR102422270B1