Emotion recognition method and electronic device supporting same

Customized learning data for emotion recognition models in electronic devices addresses the issue of low accuracy by improving individual emotion recognition through user-specific training, enhancing precision.

WO2025264057A1PCT designated stage Publication Date: 2025-12-26SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/008633
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-25
Filing Date
2025-06-20
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Existing emotion recognition models in electronic devices suffer from low accuracy due to variations in speech styles among individuals, leading to inconsistent emotion recognition for different users.

Method used

Utilizing a pre-trained emotion recognition model with customized learning data based on user-specific biometric and speech input to improve individual emotion recognition accuracy.

Benefits of technology

Enhances emotion recognition accuracy by training the model with user-specific data, allowing for more precise emotion inference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025008633_26122025_PF_FP_ABST
    Figure KR2025008633_26122025_PF_FP_ABST
Patent Text Reader

Abstract

An electronic device according to various embodiments comprises at least one processor and a memory storing at least one instruction, wherein the at least one instruction may be configured to, when executed by the at least one processor, cause the electronic device to: obtain first data related to biometric information of a user and second data related to an utterance input of the user; identify emotion of the user related to the first data and the second data; generate training data labeled with the emotion of the user for at least a part of the first data and the second data; use the training data to train an emotion recognition model stored in the memory; and infer emotion of the user on the basis of the emotion recognition model trained with the training data, third data related to the biometric information of the user, and fourth data related to the utterance input of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Emotion recognition method and electronic device supporting the same

[0001] Embodiments disclosed in this document relate to an emotion recognition method and an electronic device supporting the same.

[0002] With the advancement of digital technology, a variety of electronic devices capable of communicating and processing personal information on the go, such as mobile terminals, electronic organizers, smartphones, tablet PCs, and wearable devices, have been released. These devices have evolved from simple voice calls and messaging to include video calls, electronic organizers, document processing, email, internet access, and photography.

[0003] Additionally, electronic devices can provide emotion recognition capabilities that estimate (or recognize) a user's emotions (e.g., joy, sadness, fear, or anger) and apply them to various services. For example, electronic devices can estimate the emotions felt by a specific subject (e.g., a user) through voice analysis, image analysis, or text analysis.

[0004] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above-described matters constitute prior art related to the present disclosure.

[0005] In general, electronic devices can provide AI-based emotion recognition capabilities. For example, the electronic device can provide voice, video, or text data obtained from the user as input to a pre-trained emotion recognition model, and the user's emotions can be captured through the output of the emotion recognition model.

[0006] The pre-trained emotion recognition model described above may be an artificial intelligence model built by learning from training data. For example, the training data may consist of speech data obtained from a variety of users (e.g., users of various ages, or users in occupations characterized by high emotional expression).

[0007] However, emotion recognition models trained with the aforementioned training data suffer from a problem of low emotion recognition accuracy for individuals. For example, because each user's speech style (e.g., speaking rate, accent, or volume) varies, even when multiple users utter the same utterance (e.g., "Let's meet tomorrow"), the emotion recognized for at least some users (e.g., "happy") may differ from that for others (e.g., "angry").

[0008] Accordingly, at least one example among various embodiments may utilize a pre-trained emotion recognition model with customized learning data for emotion recognition.

[0009] The technical problems to be achieved in this document are not limited to the technical problems mentioned above, and other technical problems not mentioned can be clearly understood by a person having ordinary skill in the technical field to which the present invention belongs from the description below.

[0010] According to various embodiments, an electronic device includes at least one processor and a memory operatively connected to the at least one processor and storing at least one instruction, wherein the at least one instruction, when individually or collectively executed by the at least one processor, causes the electronic device to: obtain first data related to biometric information of a user and second data related to a speech input of the user, identify the user emotion related to the first data and the second data, generate learning data in which the user emotion is labeled for at least a portion of the first data and the second data, use the learning data to train an emotion recognition model stored in the memory, and infer the emotion of the user based on the emotion recognition model trained with the learning data, third data related to the biometric information of the user, and fourth data related to the speech input of the user.

[0011] A method of operating an electronic device according to various embodiments may include an operation of acquiring first data related to biometric information of a user and second data related to a speech input of the user, an operation of identifying the user emotion related to the first data and the second data, an operation of generating learning data in which the user emotion is labeled for at least a portion of the first data and the second data, an operation of utilizing the learning data to learn an emotion recognition model stored in the electronic device, and an operation of inferring the emotion of the user based on the emotion recognition model learned with the learning data, third data related to the biometric information of the user, and fourth data related to the speech input of the user.

[0012] A computer-readable recording medium according to various embodiments may include instructions configured to acquire first data related to biometric information of a user and second data related to a speech input of the user, identify the user emotion related to the first data and the second data, generate learning data in which the user emotion is labeled for at least a portion of the first data and the second data, utilize the learning data to learn an emotion recognition model stored in the electronic device, and infer the emotion of the user based on the emotion recognition model learned with the learning data, third data related to the biometric information of the user, and fourth data related to the speech input of the user.

[0013] Electronic devices according to various embodiments disclosed in this document can use customized learning data to train an emotion recognition model.

[0014] Additionally, the electronic device according to various embodiments disclosed in this document can improve the accuracy of emotion recognition for an individual by using an emotion recognition model pre-trained with user-customized learning data.

[0015] The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned can be clearly understood by a person having ordinary skill in the art to which the present disclosure belongs from the description below.

[0016] In connection with the description of the drawings, the same or similar reference numerals may be used for the same or similar components.

[0017] FIG. 1 is a block diagram of an exemplary electronic device capable of performing the operations described in this document.

[0018] FIG. 2 is a diagram illustrating an electronic device according to various embodiments.

[0019] FIG. 3a is a diagram schematically illustrating the configuration of an emotion recognition system according to various embodiments.

[0020] FIG. 3b is a diagram for explaining an emotion recognition procedure of an emotion recognition system according to various embodiments.

[0021] FIG. 4a is a diagram schematically illustrating the configuration of an emotion recognition system according to various embodiments.

[0022] FIG. 4b is a diagram for explaining an emotion recognition procedure of an emotion recognition system according to various embodiments.

[0023] FIG. 4c is a diagram illustrating a procedure for processing input data of a first emotion recognition model according to various embodiments.

[0024] FIG. 5a is a diagram schematically illustrating the configuration of an emotion recognition system according to various embodiments.

[0025] FIG. 5b is a diagram for explaining the emotion recognition process of an emotion recognition system according to various embodiments.

[0026] FIG. 6 is a diagram for explaining the learning procedure of the first emotion recognition model according to various embodiments.

[0027] FIG. 7a is a diagram illustrating a process for generating user-customized learning data according to various embodiments.

[0028] FIG. 7b is a diagram for explaining a query message according to various embodiments.

[0029] Figure 7c is a diagram illustrating an example of a process for generating user-customized learning data.

[0030] FIG. 8 is a diagram for explaining the learning period of the first emotion recognition model according to various embodiments.

[0031] FIG. 9a is a diagram for explaining a procedure for providing a query message according to various embodiments.

[0032] Figure 9b is a diagram illustrating an example of a query message generation procedure.

[0033] FIG. 10 is a diagram illustrating an example of a learning procedure of a first emotion recognition model according to various embodiments.

[0034] FIG. 11 is another diagram exemplifying a learning procedure of a first emotion recognition model according to various embodiments.

[0035] Figure 12 is a diagram illustrating the process of updating the weights of the first emotion recognition model.

[0036] FIGS. 13A to 13C are diagrams for explaining content corresponding to final inference results according to various embodiments.

[0037] Figure 14 is a diagram for explaining the process of creating content corresponding to the final inference result.

[0038] FIG. 15 is a flowchart illustrating the operation of an electronic device according to various embodiments.

[0039] Figure 16 is a flowchart illustrating a learning data generation operation of an emotion recognition system according to various embodiments.

[0040] Figure 17 is a flowchart illustrating the learning operation of an emotion recognition system according to various embodiments.

[0041] Figure 18 is a flowchart illustrating the emotion recognition operation of an emotion recognition system according to various embodiments.

[0042] Figure 19 is a diagram schematically illustrating the configuration of an emotion recognition system according to various embodiments.

[0043] FIG. 20 is a diagram for explaining diagnostic results provided through an emotion recognition system according to various embodiments.

[0044] Hereinafter, various embodiments of this document are described with reference to the attached drawings. However, this is not intended to limit the technology described in this document to specific embodiments, and it should be understood that various modifications, equivalents, and / or alternatives of the embodiments of this document are included. In connection with the description of the drawings, similar reference numerals may be used for similar components.

[0045]

[0046] FIG. 1 is a block diagram of an exemplary electronic device (100) capable of performing the operations described in this document.

[0047] Referring to FIG. 1, the electronic device (100) may be one of various forms of electronic devices, such as a notebook (190), smartphones (191) having various form factors (e.g., a bar-type smartphone (191-1), a foldable-type smartphone (191-2), or a sliderable (or rollable) type smartphone (191-3)), a tablet (192), a cellular phone (not shown), and other similar computing devices (not shown). The components, their relationships, and their functions illustrated in FIG. 1 are exemplary only and do not limit the implementations described or claimed in this document. The electronic device (100) may be referred to as a mobile device, a user device, a multi-function device, a portable device, or a server.

[0048] The electronic device (100) may include components including at least one processor (110) (hereinafter referred to as processor (110)), at least one memory (120) (hereinafter referred to as memory (120)), at least one display (140) (hereinafter referred to as display (140)), at least one image sensor (150) (hereinafter referred to as image sensor (150)), at least one communication circuit (160) (hereinafter referred to as communication circuit (160)), and / or at least one sensor (170) (hereinafter referred to as sensor (170)). The above components are merely exemplary. For example, the electronic device (100) may include other components (e.g., power management integrated circuitry (PMIC), audio processing circuitry, an antenna, a rechargeable battery, or an input / output interface). For example, some components may be omitted from the electronic device (100). For example, some components may be integrated into one component.

[0049] The processor (110) may be implemented as one or more IC (integrated circuit (or circuitry)) chips and may perform various data processing. The processor (110) may include at least one electrical circuit and may individually or collectively perform distributed processing of instructions (or programs, data) stored in the memory (120). The processor (110) may include a processor assembly including one or more processing circuits. The processor (110) may include any processing circuit operative to control the performance and operations of one or more components (e.g., the memory (120), the display (140), the image sensor (150), the communication circuit (160), and / or the sensor (170)) of the electronic device (100). For example, the processor (110) (e.g., the application processor (AP)) may be implemented as a system on chip (SoC) (e.g., a single chip or a chipset). For example, the processor (110) may be implemented with multiple cores (or at least one core circuit), multiple chips, or multiple chipsets. For example, the processor (110) may include one or more processing circuits. For example, the processor (110) may include one or more processing circuits configured to individually and / or collectively perform various functions of the present disclosure. As a non-limiting example, at least a portion of the processor (110) may be included in a first chip of the electronic device (100), and at least another portion of the processor (110) may be included in a second chip of the electronic device (100) that is different from the first chip of the electronic device (100).

[0050] For example, the processor (110) may include a central processing unit (CPU) (111), a graphics processing unit (GPU) (112), a neural processing unit (NPU) (113), an image signal processor (ISP) (114), a display controller (115), a memory controller (116), a storage controller (117), a communication processor (CP) (118), and / or a sensor interface (119). These components of the processor (110) are merely exemplary. For example, the processor (110) may further include other components. For example, some components of the processor (110) may be omitted from the processor (110). For example, some components of the processor (110) may be included as separate components of the electronic device (100) outside the processor (110). For example, some components of the processor (110) (e.g., memory controller (116)) may be included within other components (e.g., at least a portion of memory (120), an interface (e.g., available for connection to at least one component of the electronic device (100)), a display (140) and / or an image sensor (150)).

[0051] The processor (110) may cause other components of the electronic device (100) to perform various operations by executing instructions stored in the memory (120). The CPU (111) (or central processing circuit) may be configured to control components of the processor (110) based on the execution of instructions stored in the memory (120) (e.g., volatile memory (121) and / or non-volatile memory (122)). The GPU (112) (or graphics processing circuit) may be configured to execute parallel operations (e.g., rendering). The NPU (113) (or neural processing circuit, or artificial intelligence (AI) chip) may be configured to execute operations for an artificial intelligence model (e.g., convolution computation). The ISP (114) (or image signal processing circuit) may be configured to process a raw image acquired through the image sensor (150) into a format suitable for a component within the electronic device (100) or a component of the processor (110). The display controller (115) (or display control circuit, or display processing unit (DPU)) may be configured to process an image acquired from the CPU (111), the GPU (112), the ISP (114), or the memory (120) (e.g., the volatile memory (121)) into a format suitable for the display (140). The memory controller (116) (or memory control circuit) may be configured to control reading data from the volatile memory (121) and writing data to the volatile memory (121). The storage controller (117) (or storage control circuit) may be configured to control reading data from the nonvolatile memory (122) and writing data to the nonvolatile memory (122).The CP (118) (communication processing circuit) may be configured to process data acquired from a component of the processor (110) into a format suitable for transmission to another electronic device via the communication circuit (160), or to process data acquired from another electronic device via the communication circuit (160) into a format suitable for processing by the component of the processor (110). For example, the communication circuit (160) may include one or more communication circuits. The sensor interface (119) (or sensing data processing circuit, sensor hub) may be configured to process data about the state of the electronic device (100) and / or the state of the surroundings of the electronic device (100), acquired via the sensor (170), into a format suitable for the component of the processor (110).

[0052] The processor (110) can control the operations of the electronic device (100) by executing instructions stored in the memory (120). For example, the processor (220) can correspond to a plurality of processors that collectively perform a plurality of operations by dividing them among the processors.

[0053] The memory (120) may include one or more storage media (or one or more storage devices). For example, the memory (120) may include a memory assembly including one or more storage media. For example, the one or more storage media may include permanent memory (e.g., non-volatile memory (122)) such as a hard drive, flash memory, read-only memory (ROM), semi-permanent memory (e.g., volatile memory (121)) such as random access memory (RAM), any other suitable type of storage (or storage assembly), or any combination thereof. The memory (120) may include cache memory, which is one or more different types of memory used to temporarily store data for a function or feature of the electronic device (100). As a non-limiting example, the cache memory may be included within the processor (110). The memory (120) may be fixedly embedded within the electronic device (100) or incorporated into one or more suitable types of components (e.g., a subscriber identity module (SIM) card and / or a secure digital (SD) card) that may be repeatedly inserted into and removed from the electronic device (100).

[0054] For example, the memory (120) may store one or more software applications, such as an operating system (or system) software application, a firmware software application, a driver software application, a plug-in (e.g., add-in, add-on, and / or applet) software application, and / or any other suitable software applications. For example, the one or more software applications may include instructions executable by the processor (110). For example, the memory (120) may store instructions callable by an application programming interface (API). For example, the memory (120) may store instructions within a library.

[0055]

[0056] FIG. 2 is a diagram illustrating an electronic device according to various embodiments.

[0057] Referring to FIG. 2, an electronic device (200) (e.g., electronic device (100)) according to various embodiments may be implemented as a wearable device (e.g., a ring-shaped electronic device) worn on the body (e.g., a finger). However, this is merely an example, and various embodiments are not limited thereto. For example, the electronic device (200) may be implemented as an electronic device in the form of glasses, a watch, or a band, or may be implemented as a portable medical device, a camera, or a home appliance.

[0058] According to various embodiments, the electronic device (200) may include an outer ring member (210), an inner ring member (220) disposed along an inner surface of the outer ring member (210), a cover member (230) disposed along an outer surface of the outer ring member (210), a display (240) (e.g., display (140)), a sensor (or sensor module) (250) (e.g., sensor (170)), and a microphone (260). Depending on the embodiment, the electronic device (200) may be implemented to have more or fewer components than the components illustrated in FIG. 2.

[0059] According to various embodiments, the outer ring member (210) may have an inner diameter that can be fitted to a user's body.

[0060] According to one embodiment, at least some of the components described above (e.g., a display (240), a sensor (250), and a microphone (260)) may be disposed on the outer ring member (210). For example, a groove may be formed on the outer surface of the outer ring member (210) into which at least some of the components may be seated. For example, at least some of the components seated on the outer surface of the outer ring member (210) may not be exposed to the outside by the cover member (230).

[0061] According to one embodiment, the outer ring member (210), the inner ring member (220), and the cover member (230) may form the exterior of the electronic device (200). In this case, the inner ring member (220) may be in direct contact with the user's body.

[0062] According to an embodiment, the inner ring member (220) may be formed in a circular shape having a predetermined thickness and may be detachably coupled along the inner circumferential surface of the outer ring member (210).

[0063] According to various embodiments, the display (240) may include at least one of a first display (241) and a second display (242) arranged circumferentially along the outer side of the outer ring member (210) with the cover member (230) interposed therebetween. Depending on the embodiment, at least a portion of the display (240) may be exposed through the cover member (230).

[0064] According to one embodiment, the display (240) may be configured to provide visual information (e.g., text, images, videos, icons, or symbols) to a user and receive user input (e.g., touch input).

[0065] However, this is merely an example, and various embodiments are not limited thereto. Depending on the embodiment, a transparent display configured as a transparent or optically transparent type may be provided along the outer circumference of the outer ring member (210) (or cover member (230)). Depending on the embodiment, a transparent display configured as a transparent or optically transparent type may form the exterior of the electronic device (200) together with or instead of the cover member (230).

[0066] Additionally or optionally, the display (240) may output information indicating the user's emotional state. For example, when outputting information indicating the user's emotional state, the display (240) may emit light in a designated light pattern.

[0067] In this regard, the outer surface of the outer ring member (210) may be provided with a light-emitting element (e.g., one light-emitting element or a plurality of light-emitting elements (e.g., an LED array)) configured to irradiate a designated light. For example, the outer surface of the outer ring member (210) equipped with a transparent display may be provided with a light-emitting element configured to irradiate a designated light.

[0068] According to an embodiment, the light-emitting element may be provided in the form of a band extending a certain length along the circumference of the outer ring member (210) or extending a certain length in the width direction of the outer ring member (210). In this case, the light-emitting band implemented by the light-emitting element may emit light in different light-emitting patterns (e.g., different brightness, different colors, different lighting patterns, or a combination thereof) depending on the emotional state of the user.

[0069] Depending on the embodiment, the light emitting element may be provided in a form that covers all or part of the outer surface of the outer ring member (210).

[0070] According to various embodiments, at least a portion of the sensor (250) may be exposed through the inner ring member (220). For example, the inner ring member (220) may include a plurality of openings formed to expose at least a portion of the light emitting portion (251) and the light receiving portion (252) of the sensor (250). Accordingly, when the electronic device (200) is worn on the body, the sensor (250) may be exposed toward the body through the inner ring member (220).

[0071] According to one embodiment, the light emitting unit (251) may be composed of a first light emitting element that radiates light of a first wavelength (e.g., a red wavelength) and a second light emitting element that radiates light of a second wavelength (e.g., an infrared wavelength). For example, the first light emitting element and the second light emitting element may be implemented using a light emitting diode, an organic light emitting diode, a quantum dot light emitting diode, a laser diode, or a phosphor. However, this is merely an example, and various embodiments are not limited thereto. For example, the light emitting unit (251) may further include one or more additional light emitting elements having a wavelength that is the same as or different from the wavelengths of the first light emitting element and the second light emitting element (e.g., a blue wavelength, a green wavelength).

[0072] According to one embodiment, the light receiving unit (252) can receive light and generate a current signal by photoelectrically converting the received light. For example, the light receiving unit (252) can include at least one light receiving element for detecting light of a first wavelength irradiated from a first light emitting element and light of a second wavelength irradiated from a second light emitting element. For example, the at least one light receiving element can include a photo detector or a photo diode.

[0073] For example, the light receiving unit (252) can detect at least some of the light irradiated from the light emitting unit (251) and reflected from the subject. In this regard, the light receiving unit (214) and the light emitting unit (212) can be arranged on the same surface.

[0074] As another example, the light receiving unit (252) can detect light that has been irradiated from the light emitting unit (251) and has passed through the subject. In this regard, the light receiving unit (252) and the light emitting unit (251) can be arranged to face each other. However, this is merely an example, and various embodiments are not limited thereto, and the light receiving unit (252) and the light emitting unit (251) can be arranged in various forms.

[0075] According to an embodiment, the aforementioned sensor (250) can acquire biometric information about a part of the body. According to one embodiment, the biometric information can be used to predict the user's emotions. This will be described in more detail with reference to FIGS. 3A and 3B below.

[0076] According to various embodiments, the microphone (260) may be configured to convert external sounds into electrical audio signals and output them. According to one embodiment, the microphone (260) may be applied with various noise removal algorithms to remove noise generated during the process of receiving external sounds.

[0077] According to various embodiments, the aforementioned electronic device (200) may provide an emotion recognition function that estimates a user's emotions and applies them to various services. This will be described in detail with reference to FIGS. 3A to 20 below. Furthermore, at least one of the various embodiments described with reference to FIGS. 3A to 20 below may be combined with other embodiments.

[0078]

[0079] Figure 3a is a schematic diagram illustrating the configuration of an emotion recognition system according to various embodiments. Furthermore, Figure 3b is a diagram illustrating the emotion recognition process of the emotion recognition system according to various embodiments.

[0080] Referring to FIGS. 3A and 3B , an emotion recognition system (30) according to various embodiments may be configured with an electronic device (310) (e.g., electronic device (200)) and a first external device (320). According to one embodiment, the electronic device (310) and the first external device (320) may communicate with each other via a network (e.g., a short-range communication network or a long-range communication network).

[0081] According to an embodiment, the electronic device (310) may be referred to as a wearable device (e.g., a ring-shaped electronic device) worn on a part of the body (e.g., a finger), and the first external device (320) may be referred to as a server that is communicatively connected to the electronic device (310). However, this is merely an example, and various embodiments are not limited thereto. For example, the electronic device (310) may be referred to as a first smartphone, and the first external device (320) may be referred to as a second smartphone.

[0082] An electronic device (310) according to various embodiments may provide an emotion recognition function through collaboration with a first external device (320).

[0083] According to one embodiment, the electronic device (310) may obtain a first prediction result (303) regarding the user's emotion. In addition, the first external device (320) may obtain a second prediction result (304) regarding the user's emotion, and may make a final inference (305) regarding the user's emotion based on the first prediction result (303) and the second prediction result (304). For example, the first prediction result (303) and the second prediction result (304) may include at least one of the type of emotion (valence) (e.g., positive emotion or negative emotion), the intensity of emotion (arousal), or the degree of control or dominance over the emotion (dominance). For example, the final inference (305) may be a result of more specifically inferring the user's emotion (e.g., emotional state) (e.g., at least one of anger, contempt, disgust, fear, happiness, sadness, surprise, confusion, shame, alertness, fatigue, relaxation, dissatisfaction, boredom, and embarrassment).

[0084] The configuration of an exemplary electronic device (310) related thereto will be described in more detail below.

[0085] An electronic device (310) according to various embodiments may be configured with a sensor (311), a first memory (313), a first communication circuit (315), and at least one first processor (317) (hereinafter referred to as the first processor (317)), as illustrated in FIG. 3A. Depending on the embodiment, the electronic device (310) may be implemented to have more or fewer components than the aforementioned components.

[0086] According to various embodiments, the sensor (311) may include at least one sensor. For example, the sensor (311) may acquire (or detect) information related to the user's status and generate an electrical signal or data value corresponding to the acquired information.

[0087] According to one embodiment, information related to the user's status may include biometric information (300). In this regard, the sensor (311) may include a biometric sensor configured to acquire biometric information (300). The biometric information (300) may include a pulse signal acquired from a subject (e.g., a body part such as a wrist or a finger). For example, the pulse signal may be a photoplethysmogram (PPG) signal, and according to an embodiment, the sensor (311) may be configured with a light-emitting unit (e.g., light-emitting unit (251)) and a light-receiving unit (e.g., light-receiving unit (252)).

[0088] However, this is merely an example, and various embodiments are not limited thereto. For example, the sensor (311) may additionally or alternatively include a sensor configured to obtain information related to the posture of the electronic device (310) (e.g., at least one of an acceleration sensor, a gyro sensor, a gesture sensor, or a barometric pressure sensor).

[0089] According to various embodiments, the first memory (313) may store commands or data related to at least one other component of the electronic device (310). For example, the first memory (313) may store programs, algorithms, routines, and instructions related to emotion recognition.

[0090] According to one embodiment, a first emotion recognition model (3130) configured to recognize a user's emotion based on data acquired by a sensor (311) may be stored in the first memory (313).

[0091] According to an embodiment, the first emotion recognition model (3130) may include a feature extraction model (3131) and a first emotion prediction model (3133), as illustrated in FIG. 3b.

[0092] According to one embodiment, the feature extraction model (3131) can extract feature data from biometric information (300) (e.g., pulse signal). The feature data can include important features that can be used to derive various information from the biometric information (300). For example, the feature extraction model (3131) can extract various feature data that are correlated with the user's emotional state using the amplitude value or time value of the input biometric information (300).

[0093] For example, the feature extraction model (3131) may output first data (301) related to the biometric information (300) based on feature data extracted from the biometric information (300). The first data (301) may be related to heart rate variability indicating a change in the user's heartbeat. For example, the feature extraction model (3131) may extract various heart rate variability indices based on a plurality of peaks detected from a PPG signal. For example, the feature extraction model (3131) may extract at least one of low frequency (LF), very low frequency (VLF), which indicate sympathetic nerve activity, high frequency (HF), RR interval (RR), LF / HF ratio, standard deviation of RR intervals (SDRR), or total power (TP), which indicate parasympathetic nerve activity.

[0094] In this regard, the feature extraction model (3131) can extract feature data using various known techniques, such as linear predictive coefficient technique, cepstrum technique, mel frequency cepstral coefficient technique, and filter bank energy technique.

[0095] However, this is merely an example, and various embodiments are not limited thereto. For example, the first data (301) may include at least one of stress information, blood pressure information, blood sugar information, or sleep information.

[0096] According to one embodiment, the first emotion prediction model (3133) may output a first prediction result (303) related to the user's emotion based on the first data (301). As described above, the first prediction result (303) may include at least one of the type of emotion (valence) (e.g., positive emotion or negative emotion), the intensity of emotion (arousal), or the degree of control or dominance over the emotion. For example, when comparing the first prediction result (303) for a user who feels sad and the first prediction result (303) for a user who feels angry, the type of emotion (e.g., negative emotion) and the intensity of emotion (arousal) may be similar, but the degree of control or dominance over the emotion may be different.

[0097] The first emotion recognition model (3130) described above (and / or models included in the first emotion recognition model (3130) (e.g., at least one of the feature extraction model (3131) and the first emotion prediction model (3133)) may be an artificial intelligence model generated through machine learning. For example, the first emotion recognition model (3130) may be machine-learned using user-customized learning data (or a personalized learning model). For example, the first emotion recognition model (3130) machine-learned using user-customized learning data (or user-customized training data) may be available (or accessible) only to the electronic device (310), and use by devices other than the electronic device (310) may be restricted. The learning of the first emotion recognition model (3130) will be described in detail with reference to FIGS. 6 to 12 below.

[0098] According to one embodiment, the first emotion recognition model (3130) described above may include a plurality of artificial neural network layers. The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), deep Q-networks, a transformer network, or a combination of two or more thereof, but is not limited to the examples described above. The first emotion recognition model (3130) may additionally or alternatively include a hardware structure in addition to a software structure.

[0099] According to an embodiment, the first emotion recognition model (3130) may include first-level artificial neural network layers with relatively low performance. This limits the performance of the first emotion recognition model (3130) due to the processing power of the electronic device (310), which is limited by its miniaturization characteristics. However, various embodiments are not limited thereto. For example, the first emotion recognition model (3130) may also include second-level artificial neural network layers with relatively high performance.

[0100] According to various embodiments, the first communication circuit (315) may support wireless communication with the first external device (320). According to one embodiment, the first communication circuit (315) may be a device including hardware and software for transmitting and receiving signals (e.g., commands or data) between the electronic device (310) and the first external device (320).

[0101] For example, the first communication circuit (315) may communicate with the first external device (320) via a first network (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)).

[0102] According to various embodiments, the first processor (317) may be operatively connected to the sensor (311), the first memory (313) and the first communication circuit (315), and may control various components (e.g., hardware or software components) of the electronic device (310).

[0103] According to one embodiment, the first processor (317) can obtain a first prediction result (303) related to the user's emotion based on biometric information (300) obtained through a sensor (311) and a first emotion recognition model (3130) stored in the first memory (313).

[0104] According to one embodiment, the first processor (317) may provide the first prediction result (303) obtained based on the first emotion recognition model (3130) to the first external device (320). For example, the first prediction result (303) provided to the first external device (320) may be used to ultimately infer the user's emotion.

[0105] In this regard, the first processor (317) may provide the first data (or biometric information) (301) utilized to obtain the first prediction result (303) to the first external device (320). This first data (301) may be used by the first external device (320) to obtain the second prediction result (304) regarding the user's emotion.

[0106] The configuration of an exemplary first external device (320) related thereto will be described in more detail below.

[0107] According to various embodiments, a first external device (320) may be configured with a second memory (323), a second communication circuit (325), and at least one second processor (327) (hereinafter referred to as the second processor (327)), as illustrated in FIG. 3A. Depending on the embodiment, the first external device (320) may be implemented to have more or fewer components than the aforementioned components.

[0108] The second memory (323), the second communication circuit (325), and the second processor (327) described above may be similar to or identical to the first memory (313), the first communication circuit (315), and the first processor (317) of the electronic device (310) described above, and thus a detailed description thereof may be omitted.

[0109] According to various embodiments, the second memory (323) may store commands or data related to at least one other component of the first external device (320). For example, the second memory (323) may store programs, algorithms, routines, and commands related to emotion recognition.

[0110] According to one embodiment, a second emotion recognition model (3230) configured to infer a user's emotion using first data (301) acquired from an electronic device (310) may be stored in the second memory (323). The first data (301) may include the user's heart rate variability derived from biometric information (300) acquired by the electronic device (310) (e.g., sensor (311)).

[0111] According to an embodiment, the second emotion recognition model (3230) may include a second emotion prediction model (3233) and an inference model (3235), as illustrated in FIG. 3B . In FIG. 3B , dotted lines indicate data (and / or information) provided from outside the first external device (320), and solid lines indicate data provided from within the first external device (320).

[0112] According to one embodiment, the second emotion prediction model (3233) may output a second prediction result (304) related to the user's emotion based on the first data (301) provided from the electronic device (310). As described above, the second prediction result (304) may include at least one of the type of emotion (valence) (e.g., positive emotion or negative emotion), the intensity of the emotion (arousal), or the degree of control or dominance over the emotion (dominance). For example, the second emotion prediction model (3233) may obtain the second prediction result (304) by using the first data (301) utilized to obtain the first prediction result (303). As described above, since the artificial neural network layer included in the second emotion recognition model (3230) is different from the artificial neural network layer included in the first emotion recognition model (3130), the second emotion prediction model (3233) can output a second prediction result (304) that is different from the first prediction result (303) even if the first data (301) used to obtain the first prediction result (303) is used.

[0113] According to one embodiment, the inference model (3235) may make a final inference of the user's emotion based on the first prediction result (303) obtained by the electronic device (310) (e.g., the first emotion prediction model (3130)) and the second prediction result (304) obtained by the first external device (320) (e.g., the second emotion prediction model (3230)). As described above, the final inference (305) may be a result of inferring the user's emotion (e.g., emotional state) more specifically (e.g., at least one of anger, contempt, disgust, fear, happiness, sadness, surprise, confusion, shame, alertness, fatigue, relaxation, dissatisfaction, boredom, and embarrassment). For example, the inference model (3235) may include an activation function (e.g., a softmax function) that takes the first prediction result (303) and the second prediction result (304) as inputs and outputs the final inference result (305) of the user's emotion.

[0114] As described above, the configuration of outputting the final inference result (305) using the second emotion prediction model (3233) and the inference model (3235) is only one embodiment, and various embodiments are not limited thereto. For example, the second emotion prediction model (3233) and the inference model (3235) may be integrated into one. In this case, the second emotion prediction model (3233) (or the inference model (3235)) may take the first prediction result (303) and the first data (301) as inputs, and output the final inference result (305) regarding the user's emotion.

[0115] The second emotion recognition model (3230) described above (and / or models included in the second emotion recognition model (3230) (e.g., at least one of the second emotion prediction model (3233) and the inference model (3235)) may be an artificial intelligence model generated through machine learning. For example, the second emotion recognition model (3230) may be machine-learned using public learning data (or public training data).

[0116] In one embodiment, the shared learning data may be statistical learning data obtained by various users (e.g., users of various ages, multiple users in occupations skilled in emotional expression) and labeled with the emotions felt by each user. For example, a second emotion recognition model (3230) machine-learned using the shared learning data may be utilized by other devices capable of communicating with the first external device (320).

[0117] According to one embodiment, the second emotion recognition model (3230) may include a plurality of artificial neural network layers. According to an embodiment, the second emotion recognition model (3230) may include second-level artificial neural network layers with relatively high performance. For example, the second emotion recognition model (3230) may include more artificial neural network layers than the first emotion recognition model (3130). In other words, the second emotion recognition model (3230) may utilize a wider variety of input data to infer the user's emotions than the first emotion recognition model (3130).

[0118] According to various embodiments, the second communication circuit (325) may support wireless communication with the electronic device (310). According to one embodiment, the second communication circuit (325) may be a device including hardware and software for transmitting and receiving signals (e.g., commands or data) between the first external device (320) and the electronic device (310).

[0119] According to various embodiments, the second processor (327) may be operatively connected to the second memory (323) and the second communication circuit (325) and may control various components (e.g., hardware or software components) of the first external device (320).

[0120] According to one embodiment, the second processor (327) may make a final inference about the user's emotion based on a first prediction result (303) obtained through the electronic device (310) (e.g., a first emotion prediction model (3133)) and a second prediction result (304) obtained by the first external device (320) (e.g., a second emotion prediction model (3233)). For example, the second processor (327) may provide the final inference result (305) to the electronic device (310).

[0121] As described above, the electronic device (310) according to various embodiments can utilize the first data (301) related to the biometric information (300) for user emotion recognition. Additionally or alternatively, the electronic device (310) according to various embodiments can further enhance the emotion recognition performance by utilizing additional information other than the biometric information (300) for emotion recognition. For example, the additional information may include the user's speech. Such user's speech can be utilized to more specifically predict at least the user's emotion type (valence) (e.g., positive emotion or negative emotion), thereby increasing the reliability of the user's emotion recognition result. This will be described in detail with reference to FIGS. 4A and 4B below.

[0122]

[0123] FIG. 4a is a schematic diagram illustrating the configuration of an emotion recognition system according to various embodiments. FIG. 4b is a diagram illustrating the emotion recognition procedure of an emotion recognition system according to various embodiments, and FIG. 4c is a diagram illustrating the procedure for processing input data of a first emotion recognition model according to various embodiments.

[0124] Referring to FIGS. 4A and 4B , an emotion recognition system (40) according to various embodiments may be configured with an electronic device (310) (e.g., electronic device (200)) and a first external device (320). Compared to the emotion recognition system (30) illustrated in FIG. 3A , the emotion recognition system (40) illustrated in FIG. 4A differs in that the electronic device (310) acquires additional information other than biometric information (300) and utilizes it for emotion recognition. The configuration of an exemplary emotion recognition system (40) related thereto will be described in more detail below.

[0125] An electronic device (310) according to various embodiments may be configured with a sensor (311), a first memory (313), a first communication circuit (315), a first processor (317), and a microphone (411), as illustrated in FIG. 4A. Depending on the embodiment, the electronic device (310) may be implemented to have more or fewer components than the aforementioned components. In addition, among the components of the electronic device (310) illustrated in FIG. 4A, the sensor (311), the first memory (313), the first communication circuit (315), and the first processor (317) may be similar to or identical to the components of the electronic device (310) illustrated in FIG. 3A, and thus, a detailed description thereof may be omitted.

[0126] According to various embodiments, the microphone (411) (e.g., the microphone (260)) may be configured to convert external sounds into electrical audio signals and output them. According to one embodiment, the microphone (411) may be configured to receive a user's speech input (400). This speech input (400) may be utilized to recognize the user's emotions.

[0127] In this regard, the feature extraction model (3131) can extract feature data from the speech input (400), as illustrated in FIG. 4B. The feature data may include important features that can be used to identify the speaker and / or the speaker's speech style (e.g., speech rate, accent, or speech volume). For example, the feature extraction model (3130) can extract various feature data using the amplitude value or time value of the input speech input (400).

[0128] For example, the feature extraction model (3131) can output first data (301) related to biometric information (300) based on feature data extracted from the biometric information (300). Additionally, the feature extraction model (3131) can output second data (401) related to the speech input (400) based on feature data extracted from the speech input (400).

[0129] According to an embodiment, the feature extraction model (3131) may be a model that supports multiple modalities (e.g., a multimodal model). For example, when outputting the first data (301) and the second data (401), the feature extraction model (3131) may concatenate the first data (301) and the second data (401) to output the first emotion prediction model (3133).

[0130] According to one embodiment, as illustrated in FIG. 4c, the feature extraction model (3131) can perform an operation of converting one-dimensional data into two-dimensional data by connecting (480) the first data (301) and the second data (401).

[0131] For example, a feature extraction model (3131) can obtain purified biometric information by preprocessing (461) biometric information (300) in a manner such as filtering and sampling, and obtain first data (301) by converting (463) the biometric information into two-dimensional biometric information expressed in a relationship between time and frequency (e.g., scalograms).

[0132] Similarly, the feature extraction model (3131) can obtain refined speech input (400) by preprocessing (471) the speech input (400) in a manner such as filtering and sampling, and obtain second data (401) by converting (473) the speech input into a two-dimensional expression of the relationship between time and frequency (e.g., log-mel spectrogram).

[0133] These first data (301) and second data (401) are merged (480) and provided as input to the first emotion prediction model (3133), and the first emotion prediction model (3133) can output a first prediction result (303) related to the user's emotion based on the first data (301) and the second data (401).

[0134] In this regard, the first processor (317) may provide the first prediction result (303) obtained based on the first emotion recognition model (3130) to the first external device (320). Additionally, the first processor (317) may also provide the first data (or biometric information) (301) and the second data (or speech input) (401) utilized in obtaining the first prediction result (303) to the first external device (320). These first data (301) and second data (401) may be used by the first external device (320) to obtain the second prediction result (304) regarding the user's emotion.

[0135] According to one embodiment, the second emotion recognition model (3230) (e.g., the second emotion prediction model (3233)) may output a second prediction result (304) related to the user's emotion based on the first data (301) and the second data (401) provided from the electronic device (310), as illustrated in FIG. 4b.

[0136] In addition, the second emotion recognition model (3230) (e.g., inference model (3235)) can finally infer the user's emotion based on the first prediction result (303) obtained by the electronic device (310) (e.g., first emotion prediction model (3130)) and the second prediction result (304) obtained by the first external device (320) (e.g., second emotion prediction model (3230)).

[0137] Additionally or optionally, the first processor (317) may acquire second data (401) when the user is on a call or conversing with another user. This is designed to take into account situations where the user's emotional changes become more pronounced when conversing with at least two other people. However, depending on the embodiment, the first processor (317) may also acquire second data (401) when the user is talking to himself rather than conversing with another person.

[0138] Additionally or optionally, the first processor (317) may acquire the first data (301) while the second data (401) is being acquired or after it has been acquired (e.g., when the user is on a call or when the user is conversing with another user). In this regard, the first processor (317) may also change the sensor (311) from an inactive state to an active state.

[0139] Additionally or optionally, the first processor (417) may learn and store the user's voice in advance and use it to determine whether the user is conversing with others or talking to himself.

[0140] As described above, the first emotion prediction model (3133) can output the first prediction result (303) based on the first data (301) related to the biometric information (300) and the second data (401) related to the speech input (400). However, this is merely an example, and various embodiments are not limited thereto. For example, as described above with reference to FIGS. 3A and 3B , the first emotion prediction model (3133) can utilize only the first data (301) related to the biometric information (300) in outputting the first prediction result (303). In this case, the second emotion prediction model (3233) may output a second prediction result (304) based on the first data (301) related to the biometric information (300) and the second data (401) related to the speech input (400), and the inference model (3225) may output a final inference result (305) based on the first prediction result (303) output by the first emotion recognition model (3130) and the second prediction result (304) output by the second emotion prediction model (3233). Additionally or optionally, depending on the embodiment, the second emotion prediction model (3233) and the inference model (3235) may be integrated into one, in which case the second emotion prediction model (3233) (or the inference model (3235)) may take the first prediction result (303), the first data (301), and the second data (401) as inputs, and output the final inference result (305) regarding the user's emotion.

[0141] As described above, the electronic device (310) according to various embodiments may recognize the user's emotions using speech input (400) as additional information in addition to biometric information (300). Depending on the embodiment, the electronic device (310) may utilize other additional information for emotion recognition instead of speech input (400), or may utilize other additional information together with speech input (400) for emotion recognition.

[0142] In this regard, the electronic device (310) needs to be equipped with a separate component for acquiring additional information in addition to the biometric information (300). However, due to the miniaturization of the electronic device (310), it may be somewhat difficult to equip the electronic device (310) with a separate component for acquiring additional information. Accordingly, the electronic device (310) according to various embodiments may acquire additional information utilized for emotion recognition from an external source. This will be described in detail with reference to FIGS. 5A and 5B below.

[0143]

[0144] Figure 5a is a schematic diagram illustrating the configuration of an emotion recognition system according to various embodiments. Furthermore, Figure 5b is a diagram illustrating the emotion recognition process of an emotion recognition system according to various embodiments.

[0145] Referring to FIGS. 5A and 5B , an emotion recognition system (50) according to various embodiments may be configured with an electronic device (310) (e.g., electronic device (200)), a first external device (320), and a second external device (510). Compared to the emotion recognition system (40) illustrated in FIG. 4A , the emotion recognition system (50) illustrated in FIG. 5A differs in that the electronic device (310) acquires second data (501) related to speech input from the second external device (510). The configuration of an exemplary emotion recognition system (50) related thereto will be described in more detail below.

[0146] According to various embodiments, the second external device (510) may be configured with a microphone (511), a third communication circuit (515), and a third processor (517), as illustrated in FIG. 5A. Depending on the embodiment, the second external device (510) may be implemented to have more or fewer components than the aforementioned components.

[0147] According to various embodiments, the microphone (511) (e.g., the microphone (411)) may be configured to convert external sounds into electrical audio signals and output them. According to one embodiment, the microphone (511) may be configured to receive a user's speech input (400). This speech input (400) may be utilized to recognize the user's emotions.

[0148] According to various embodiments, the third communication circuit (515) may support wireless communication with the electronic device (310). In one embodiment, the third communication circuit (515) may be a device including hardware and software for transmitting and receiving signals (e.g., commands or data) between the second external device (510) and the electronic device (310). In another embodiment, the third communication circuit (515) may also support wireless communication with the first external device (320).

[0149] According to various embodiments, the third processor (517) may be operatively connected to the microphone (511) and the third communication circuit (515) and may control various components (e.g., hardware or software components) of the third external device (510).

[0150] According to one embodiment, the third processor (517) may obtain second data (501) related to a speech input (400) obtained through a microphone (511). Additionally, the third processor (517) may provide the second data (501) to the electronic device (310). This second data (501) may be used by the electronic device (310) to obtain a first prediction result (303) regarding the user's emotion.

[0151] In some embodiments, the third processor (517) may provide the speech input (400) to the electronic device (310) instead of the second data (501). In this case, the electronic device (310) may obtain the second data (501) by using the speech input (400) obtained from the second external device (510) and the feature extraction model (3131).

[0152] According to various embodiments, the electronic device (310) may obtain a first prediction result (303) related to the user's emotion based on the first data (301) and the second data (501). According to one embodiment, the first processor (317) may obtain the first prediction result (303) by providing the first data (301) obtained by the electronic device (310) and the second data (501) obtained from the second external device (510) as inputs to the first emotion prediction model (3133), as illustrated in FIG. 5B.

[0153] Additionally, the first processor (317) may provide the first data (301) acquired by the electronic device (310) and the second data (501) acquired from the second external device (510) to the first external device (320). The first data (301) and the second data (501) may be used by the first external device (320) to acquire a second prediction result (304) regarding the user's emotion.

[0154]

[0155] FIG. 6 is a diagram illustrating a learning procedure of a first emotion recognition model according to various embodiments. FIGS. 7A to 7C are diagrams illustrating a generation procedure of user-customized learning data (601) according to various embodiments, and FIG. 8 is a diagram illustrating a learning period of a first emotion recognition model according to various embodiments.

[0156] Referring to FIG. 6, the first emotion recognition model (3130) according to various embodiments can be machine-learned (600) by customized learning data (or personalized learning model) (601). For example, multiple weight values ​​of multiple artificial neural network layers included in the first emotion recognition model (3130) can be optimized by the learning results. For example, multiple weight values ​​can be updated (w → w') so that the loss value or cost value acquired from the first emotion recognition model (3130) is reduced or minimized during learning (610).

[0157] According to one embodiment, the user-customized learning data (601) is data in which emotions felt by the user are labeled (or mapped) with respect to data obtained from a user who owns (or uses) the electronic device (310), and may be used to learn the first emotion recognition model (3130) stored in the electronic device (310). For example, biometric information (or posture information or speech input of the electronic device (310)) obtained through the sensor (311) and the emotion felt by the user while the biometric information (or posture information or speech input of the electronic device (310)) is being obtained (or after the biometric information is obtained) may be used as the user-customized learning data (601). This user-customized learning data (601) may be distinguished from the common learning data used to learn the second emotion recognition model (3230) stored in the first external device (320). In addition, the first emotion recognition model (3130) machine-learned with user-customized learning data (601) is an emotion recognition model optimized for the user, can be used (or accessed) only by the electronic device (310), and can improve the emotion recognition accuracy for the user.

[0158] The learning procedure of the exemplary first emotion recognition model (3130) related to this will be described in more detail below.

[0159] The first emotion recognition model (3130) according to various embodiments may include a feature extraction model (3131), a first emotion prediction model (3133), and a learning data generation model (610), as illustrated in FIG. 6.

[0160] Depending on the embodiment, the first emotion recognition model (3130) may be implemented with more or fewer components than the aforementioned components. In addition, the feature extraction model (3131) and the first emotion prediction model (3133) illustrated in FIG. 6 may be similar to or identical to the feature extraction model (3131) and the first emotion prediction model (3133) described with reference to FIGS. 3b, 4b, and 5b, and thus, a detailed description thereof may be omitted.

[0161] According to various embodiments, the learning data generation model (610) can generate user-customized learning data used to train the first emotion recognition model (3130) (e.g., feature extraction model (3131) and / or first emotion prediction model (3133)).

[0162] According to one embodiment, the learning data generation model (610) can generate learning data by labeling emotions felt by the user on data collected (or acquired) by the electronic device (310). For example, as illustrated in FIG. 7A, data (701) collected by the electronic device (310) (e.g., at least one of biometric information (300), speech input (400), or posture information of the electronic device (310)) and emotions (703) felt by the user while the data is being collected (or after the data is being collected) can be provided as inputs to the learning data generation model (610), and the learning data generation model (610) can output user-customized learning data (601) with the data and emotions labeled.

[0163] In this regard, the electronic device (310) may provide a query message asking about the emotions felt by the user while data is being collected (or after data is collected).

[0164] According to one embodiment, the query message may be visual information output through an electronic device (310) (e.g., display (240)), as illustrated in 730 of FIG. 7b.

[0165] Additionally or alternatively, the query message may be visual information output through another device (e.g., a smartphone) that is connected to the electronic device (310) via communication, as illustrated in 740 of FIG. 7B . In such a case, the electronic device (310) may directly obtain the user's input in response to the query message or may obtain it through another device connected to the communication.

[0166] However, this is merely an example, and various embodiments are not limited thereto. For example, depending on the embodiment, the query message may be provided through auditory information, tactile information, or a combination thereof.

[0167] According to various embodiments, the electronic device (310) may identify the user's emotions based on user input in response to a query message. In one embodiment, the electronic device (310) may identify the user's emotions based on user input such as touch input, speech input, or gesture input, and label the emotions in the collected data (701).

[0168] An exemplary operation related to the generation of the aforementioned user-customized learning data (601) is described with reference to FIG. 7c.

[0169] Referring to FIG. 7c, an electronic device (310) according to various embodiments can detect a user's speech input (e.g., "Let's meet tomorrow") (751). For example, the electronic device (310) can analyze audio data containing voice signals to detect a situation in which a call is in progress or a situation in which the user is conversing with another user.

[0170] In one embodiment, after detecting a user's speech input, the electronic device (310) may provide a query message (e.g., "Please enter the emotion you are feeling now") asking about the user's feelings (753). For example, the electronic device (310) may provide the query message while detecting the speech input or after the speech input is terminated.

[0171] In one embodiment, after providing a query message, the electronic device (310) may receive a user input (e.g., "joy") in response to the query message (755). For example, the electronic device (310) may receive a user's spoken input in response to the query message. However, depending on the embodiment, the electronic device (310) may also receive a user response in the form of a touch input.

[0172] In one embodiment, after receiving a user input responding to a query message, the electronic device (310) may generate user-customized learning data (601) labeled with the user's utterance data and emotions (757). For example, the electronic device (310) may label the utterance itself (e.g., "Let's meet tomorrow") with an emotion (e.g., "joy") or may label at least one utterance word (e.g., "Let's meet tomorrow") with an emotion (e.g., "joy").

[0173] As described above, the performance of the first emotion recognition model (3130) according to various embodiments can be improved through machine learning. This machine learning may require a certain level of resources from the electronic device (310). However, if machine learning is performed in a situation where the resources of the electronic device (310) exceed a certain level, the learning results for the first emotion recognition model (3130) may deteriorate.

[0174] In this regard, the electronic device (310) according to various embodiments can train the first emotion recognition model (3130) in a situation where the use of resources is relatively low.

[0175] According to one embodiment, the electronic device (310) can perform machine learning operations while being removed from the body. For example, as illustrated in FIG. 8, the machine learning operations can be performed while the electronic device (310) is connected to a charging system (830) (e.g., a charging cradle). According to an embodiment, as described below with reference to FIGS. 10 and 11 , the electronic device (310) can perform machine learning operations through collaboration with a first external device (320). In this case, the electronic device (310) can activate the first communication circuit (315) based on the connection with the charging system (830), and then transmit and receive data related to machine learning with the first external device (320), as described below.

[0176]

[0177] Figure 9a is a diagram illustrating a query message provision process according to various embodiments. Figure 9b is a diagram exemplarily illustrating a query message generation process.

[0178] According to various embodiments, the electronic device (310) may provide a query message that inquires about the emotions felt by the user when generating user-tailored learning data (601). According to one embodiment, the electronic device (310) may analyze the situation while data (701) is being collected by the electronic device (310) and provide an appropriate query message. For example, when the user is on a call, the electronic device (310) may analyze the content of the call and provide an appropriate query message based on the content of the call.

[0179] The configuration of an exemplary first emotion recognition model (3230) related to this will be described in more detail below.

[0180] Referring to FIG. 9A, a first emotion recognition model (3230) according to various embodiments may include a second emotion prediction model (3233), an inference model (3235), and an emotion query generation model (910). Depending on the embodiment, the second emotion prediction model (3233) and the inference model (3235) illustrated in FIG. 9A may be similar to or identical to the configurations described with reference to FIGS. 3B, 4B, and 5B, and thus, a detailed description thereof may be omitted.

[0181] According to various embodiments, the emotional query generation model (910) may be provided with data (701) collected by the electronic device (310) (e.g., biometric information (300), speech input (400), or posture information of the electronic device (310)) as input.

[0182] According to one embodiment, the emotional query generation model (910) can analyze data (701) collected by the electronic device (310) and output an appropriate query message to the user.

[0183] For example, as illustrated in FIG. 9b, in a situation where a user is talking to the other party, the electronic device (310) may collect incoming voice (e.g., “I have good news. Let’s meet up and talk”) and outgoing voice (e.g., “Can’t you tell me now? Then let’s meet up tomorrow”) and provide them to an emotional query generation model (910). In this regard, the emotional query generation model (910) may select and output a query message that matches the mood of the conversation from among various query messages (920).

[0184] As another example, in a situation where biometric information is acquired, the electronic device (310) may select and output a query message (e.g., “What emotion did you feel at the current point when your heart rate variability is increasing?”) that matches the biometric information (e.g., increasing heart rate variability) among various query messages.

[0185]

[0186] Figure 10 is a diagram illustrating an exemplary learning procedure for a first emotion recognition model according to various embodiments. Furthermore, Figure 12 is a diagram illustrating a process for updating the weights of the first emotion recognition model.

[0187] As described above, the electronic device (310) according to various embodiments may store a first emotion recognition model (3130) configured to recognize the user's emotions. This first emotion recognition model (3130) may be learned by the electronic device (310) as described above with reference to FIG. 6 . However, due to the limited processing power of the electronic device (310) due to its miniaturization, it may be somewhat difficult to learn the first emotion recognition model (3130). Accordingly, the electronic device (310) according to various embodiments may learn the first emotion recognition model (3130) through the first external device (320). The learning procedure of an exemplary first emotion recognition model (3130) related thereto will be described in more detail below.

[0188] An electronic device (310) according to various embodiments (e.g., a first emotion recognition model (3130)) may include a feature extraction model (3131) and a first emotion prediction model (3133). Depending on the embodiment, the second emotion prediction model (3233) and the inference model (3235) illustrated in FIG. 10 may be similar to or identical to the configurations described with reference to FIGS. 3b, 4b, and 5b, and thus, a detailed description thereof may be omitted.

[0189] According to one embodiment, the electronic device (310) may request learning of the first emotion prediction model (3133) by providing the user-customized learning data (601) and the first emotion prediction model (3133) to the first external device (320) (1010). However, this is merely exemplary, and various embodiments are not limited thereto. For example, the electronic device (320) may also request learning of not only the first emotion prediction model (3133), but also the feature extraction model (3131) and / or the learning data generation model (610).

[0190] According to one embodiment, in response to a learning request from an electronic device (310), a first external device (320) may store a first emotion prediction model (3133) in a second memory (323) and learn it using user-customized learning data (601) (1020). For example, the first external device (320) may update (w → w') (1001) the weights of the first emotion prediction model (3133) through learning.

[0191] For example, the first external device (320) can derive a first prediction value (1211) using user-customized learning data (601) and a first emotion prediction model (3133), as illustrated in 1200 of FIG. 12, and derive a second prediction value (1213) using user-customized learning data (601) and a second emotion prediction model (3233) (e.g., performing forward propagation). In addition, the first external device (320) can update (1001) weights by combining (1215) the first prediction value (1211) and the second prediction value (1213) to reduce or minimize an error with the actual value (e.g., performing back propagation).

[0192] According to an embodiment, the first external device (320) may exclude an operation of updating the weights of the second emotion prediction model (3233). This is because training the second emotion prediction model (3233) including relatively high-performance second-level artificial neural network layers significantly requires the resources of the first external device (320), so the first external device (320) can efficiently manage the resources of the first external device (320) by excluding training the second emotion prediction model (3233). The shaded second emotion prediction model (3233) in 1210 of FIG. 12 represents a state in which it is excluded from training.

[0193] Additionally, the first external device (320) can provide updated weights (1001) to the electronic device (310). Accordingly, the electronic device (310) can recognize the emotions of the user using the weights (1001) provided from the first external device (320).

[0194] The aforementioned user-customized learning data (601) may be personal information that can identify the user. Accordingly, the electronic device (310) can prevent personal information leakage by processing (e.g., encrypting) the user-customized learning data (601) and providing it to the first external device (320). Furthermore, after learning the first emotion prediction model (3133), the first external device (320) may delete the user-customized learning data (601) used for learning.

[0195]

[0196] FIG. 11 is another diagram exemplifying a learning procedure of a first emotion recognition model according to various embodiments.

[0197] As described above, the electronic device (310) according to various embodiments may provide user-customized learning data (601) and a first emotion prediction model (3133) to the first external device (320) as part of a learning request. This relates to a situation where the first external device (320) does not store the first emotion prediction model (3133). However, unlike what was described with reference to FIG. 10, the first external device (320) illustrated in FIG. 11 may store an emotion recognition model (1101) corresponding to each of a plurality of electronic devices.

[0198] In this regard, the electronic device (310) according to various embodiments may, as illustrated in FIG. 11, provide only learning data (601) to the first external device (320) as part of a learning request (1110). Additionally, the electronic device (310) may also provide identification information of the electronic device (310) to the first external device (320).

[0199] According to one embodiment, in response to a learning request from the electronic device (310), the first external device (320) may select one emotion prediction model (1101-1) corresponding to the electronic device (310) from among the stored emotion prediction models (1101) and learn it using user-customized learning data (601) (1120). For example, the first external device (320) may select, as a learning target, an emotion prediction model corresponding to (or assigned to) the identification information of the electronic device (310) from among the stored emotion prediction models (1101).

[0200] Additionally, the first external device (320) can update (w → w') the weights of the selected emotion prediction model (3133) through learning (1103) and provide the updated weights (1103) and the learned emotion prediction model (1101-1) to the electronic device (310). Accordingly, the electronic device (310) can recognize the emotion of the user by using the learned emotion prediction model (1101-1) and weights (1103) provided from the first external device (320).

[0201] As described above, the first external device (320) according to various embodiments can make a final inference about the user's emotions based on the first prediction result (303) obtained by the electronic device (310) and the second prediction result (304) obtained by the first external device (320). Additionally or optionally, the first external device (320) according to various embodiments can generate and provide content corresponding to the final inference result. This will be described in detail with reference to FIGS. 13A and 14 below.

[0202]

[0203] Figures 13a to 13c are diagrams illustrating content corresponding to the final inference result according to various embodiments. Furthermore, Figure 14 is a diagram illustrating the process of creating content corresponding to the final inference result.

[0204] Referring to FIGS. 13A to 13C, the first external device (320) according to various embodiments can finally infer emotions for the user and provide content related thereto.

[0205] In this regard, the first external device (320) according to various embodiments may include a second emotion prediction model (3233), an inference model (3235), and a content generation model (1410), as illustrated in FIG. 14. For example, the second emotion prediction model (3233), the inference model (3235), and the content generation model (1410) may be part of a second emotion recognition model (3230) stored in the first external device (320). Depending on the embodiment, the second emotion prediction model (3233) and the inference model (3235) illustrated in FIG. 14 may be similar to or identical to the configuration described with reference to FIGS. 3b, 4b, and 5b, and thus a detailed description thereof may be omitted.

[0206] According to various embodiments, the content generation model (1410) may receive as input the final inference result (305) regarding the user's emotions output from the inference model (3235).

[0207] According to one embodiment, the content generation model (1410) can analyze the final inference result (305) and recommend (1411) suitable content to the user.

[0208] According to one embodiment, as illustrated in FIG. 13A, if the emotion finally inferred by the first external device (320) is a first emotion (e.g., surprise), the first external device (320) may generate first image content (1311) visually expressing the first emotion. Similarly, if the emotion finally inferred by the first external device (320) is a second emotion (e.g., happiness) or a third emotion (e.g., depression), the first external device (320) may generate second image content (1313) or third image content (1315) visually expressing the second emotion or the third emotion.

[0209] According to one embodiment, as illustrated in FIG. 13b, the first external device (320) may provide content generated based on the final inference result (350) to the electronic device (310) and / or another electronic device that is connected to the electronic device (310) through communication. For example, the generated content may be provided as a background image of a watch-type wearable, as illustrated in 1320 of FIG. 13b, or as a background image of a smartphone, as illustrated in 1330 of FIG. 13b.

[0210] Additionally or optionally, as illustrated in FIG. 13c, the first external device (320) may provide information helpful in improving the user's emotions based on the final inferred emotions. For example, if the final inferred emotions by the first external device (320) are the second emotions (e.g., happiness), the first external device (320) may provide guide information recommending food menus that can maintain or further increase the emotions. As another example, if the final inferred emotions by the first external device (320) are the second emotions (e.g., happiness), the first external device (320) may recommend movies or music that help maintain or improve the emotions.

[0211]

[0212] An electronic device (310) according to various embodiments may include at least one processor (317) and a memory (313) operatively connected to the at least one processor (317) and storing at least one command. According to one embodiment, the at least one instruction, when individually or collectively executed by the at least one processor (317), may be configured to cause the electronic device (310) to: obtain first data (301) related to biometric information of the user and second data (401) related to a speech input of the user; identify the user emotion related to the first data (301) and the second data (401); generate learning data (601) in which the user emotion is labeled for at least a portion of the first data (301) and the second data (401); use the learning data (601) to train an emotion recognition model (3130) stored in the memory (313); and infer the emotion of the user based on the emotion recognition model (3130) trained with the learning data (601), third data related to the biometric information of the user, and fourth data related to the speech input of the user.

[0213] According to various embodiments, the electronic device (310) may further include a display (240) and a communication circuit (315). According to one embodiment, the at least one instruction, when individually or collectively executed by the at least one processor (317), may be configured to cause the electronic device (310) to: obtain a query message querying the user emotion from the external device (320) through the communication circuit (315), output the query message through the display (240), and identify the user emotion related to the first data (301) and the second data (401) based on an input detected after outputting the query message.

[0214] According to various embodiments, the at least one instruction, when individually or collectively executed by the at least one processor (317), may be configured to cause the electronic device (310) to: provide the stored emotion recognition model (3130) and the learning data (601) to an external device (320) via the communication circuit (315), and acquire the learned emotion recognition model (3130) from the external device (320) via the communication circuit (315).

[0215] According to various embodiments, the at least one instruction, when executed individually or in combination by the at least one processor (317), may be configured to cause the electronic device (310) to: obtain a plurality of weight values ​​of artificial neural network layers included in the learned emotion recognition model (3130) from the external device (320) through the communication circuit (315).

[0216] According to various embodiments, the at least one instruction, when executed individually or in combination by the at least one processor (317), may be configured to cause the electronic device (310) to: provide a first prediction result (303) obtained based on the learned emotion recognition model (3130), the third data and the fourth data to the external device (320) through the communication circuit (315), and obtain, from the external device (320) through the communication circuit (315), the user's emotion inferred based on the first prediction result (303) and a second prediction result (304) obtained from the external device (320) based on the third data and the fourth data.

[0217] According to various embodiments, the at least one instruction, when individually or collectively executed by the at least one processor (317), may be configured to cause the electronic device (310) to: obtain content (1311-1315) related to the inferred user emotion from the external device (320) via the communication circuit (315).

[0218] According to various embodiments, the at least one instruction, when individually or collectively executed by the at least one processor (317), may be configured to cause the electronic device (315) to: obtain one of the first data (301) and the second data (401) from another external device (510) via the communication circuit (315).

[0219] According to various embodiments, the at least one instruction, when individually or collectively executed by the at least one processor (317), may be configured to cause the electronic device (310) to: obtain the user emotion related to the first data (301) and the second data (401) from the other external device (510) via the communication circuit (315).

[0220] According to various embodiments, the electronic device (310) may be implemented in a body-wearable form.

[0221] According to various embodiments, the at least one instruction, when individually or collectively executed by the at least one processor (317), may be configured to cause the electronic device (310) to: acquire the learned emotion recognition model (3130) from an external device (320) via the communication circuit (315) while the electronic device (310) is unworn by the user.

[0222] A computer-readable recording medium according to various embodiments may include instructions configured to acquire first data related to biometric information of a user and second data related to a speech input of the user, identify the user emotion related to the first data and the second data, generate learning data in which the user emotion is labeled for at least a portion of the first data and the second data, utilize the learning data to learn an emotion recognition model stored in the electronic device, and infer the emotion of the user based on the emotion recognition model learned with the learning data, third data related to the biometric information of the user, and fourth data related to the speech input of the user.

[0223]

[0224] Figure 15 is a flowchart illustrating the operation of an electronic device according to various embodiments. While the operations in the following embodiments may be performed sequentially, they are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel. Furthermore, at least one of the aforementioned operations may be omitted depending on the embodiment.

[0225] According to one embodiment, operations 1501 to 1509 may be understood to be performed in a processor (e.g., the first processor (317) of FIG. 3A) of an electronic device (e.g., the electronic device (310) of FIG. 3A).

[0226] Referring to FIG. 15, an electronic device (310) according to various embodiments may obtain user-related data in operation 1501. According to one embodiment, the user-related data may include at least one of biometric information, posture information of the electronic device (310), or speech input.

[0227] According to various embodiments, the electronic device (310) may, in operation 1503, identify the user's emotions while data is being acquired. In one embodiment, the electronic device (310) may identify the user's emotions while data is being acquired (or at the time data acquisition is completed). In this regard, the electronic device (310) may also provide a query message querying the user's emotions.

[0228] According to various embodiments, the electronic device (310) may generate user-customized learning data (e.g., user-customized learning data (601)) based on acquired data and user emotions in operation 1505. The user-customized learning data is data in which emotions felt by a user are labeled (or mapped) to data acquired from a user who owns (or uses) the electronic device (310), and may be used to train an emotion recognition model (e.g., a first emotion recognition model (3130)) stored in the electronic device (310).

[0229] According to various embodiments, the electronic device (310) may, in operation 1507, learn a stored emotion recognition model using user-customized learning data. For example, multiple weight values ​​of multiple artificial neural network layers included in the emotion recognition model may be optimized through learning.

[0230] According to various embodiments, the electronic device (310) may recognize the user's emotions based on a learned emotion recognition model at operation 1509. For example, an emotion recognition model learned using user-tailored learning data is an emotion recognition model optimized for the user, and is only available (or accessible) to the electronic device (310) and may improve the accuracy of emotion recognition for the user.

[0231]

[0232] Figure 16 is a flowchart illustrating the learning data generation operation of an emotion recognition system according to various embodiments. The operations of Figure 16 described below may represent various embodiments of operation 1505 of Figure 15.

[0233] Referring to FIG. 16, an electronic device (1601) (e.g., electronic device (310)) according to various embodiments may obtain data related to a user in operation 1611. According to one embodiment, the electronic device (1601) may obtain first data related to the user's biometric information (e.g., first data (301)) and second data related to speech (e.g., second data (401) or second data (501)).

[0234] An electronic device (1601) according to various embodiments may provide first data and second data to a first external device (1603) (e.g., the first external device (320)) in operation 1613.

[0235] According to various embodiments, the first external device (1603) may, in operation 1615, generate a query (e.g., a query message) for emotion verification based on the first data and the second data. In this regard, reference may be made to the descriptions of FIGS. 9A and 9B described above.

[0236] According to various embodiments, the first external device (1603) may provide the generated query to the electronic device (1601) in operation 1617.

[0237] An electronic device (1601) according to various embodiments may, in operation 1619, identify a user's emotions based on a query. In this regard, reference may be made to the descriptions of FIGS. 7B and 7C described above.

[0238] An electronic device (1601) according to various embodiments may generate user-customized learning data (e.g., user-customized learning data (601)) labeling the user's emotions on first data and second data in operation 1621.

[0239]

[0240] Figure 17 is a flowchart illustrating the learning operations of an emotion recognition system according to various embodiments. The operations of Figure 17 described below may represent various embodiments of operation 1507 of Figure 15.

[0241] Referring to FIG. 17, an electronic device (1601) (e.g., electronic device (310)) according to various embodiments may, in operation 1711, provide a first emotion recognition model (e.g., first emotion recognition model (3130)) and learning data (e.g., user-customized learning data (601)) to a first external device (e.g., first external device (320)). According to one embodiment, the electronic device (1601) may provide learning data, in which emotions felt by a user are labeled (or mapped) with respect to data obtained from a user who owns (or uses) the electronic device (1601), to the first external device (1603). In this regard, reference may be made to the description of FIG. 10 described above.

[0242] A first external device (1603) according to various embodiments may perform a learning operation for a first emotion recognition model (3130) (e.g., a first emotion prediction model (3133)) in operations 1703 to 1707.

[0243] According to one embodiment, the first external device (1603) may derive a first predicted value (e.g., a first predicted value (1211)) using learning data and a first emotion recognition model (3130) (e.g., a first emotion prediction model (3133)) (operation 1703), and derive a second predicted value (e.g., a second predicted value (1213)) using learning data and a second emotion recognition model (3230) (e.g., a second emotion prediction model (3233)) (e.g., performing forward propagation) (operation 1705). In addition, the first external device (1603) may update weights (e.g., performing back propagation) (operation 1707) by combining the first predicted value and the second predicted value so as to reduce or minimize an error with the actual value. In this regard, reference may be made to the description of FIG. 12 described above.

[0244] According to various embodiments, the first external device (1603) may provide the learning results to the electronic device (1601) in operation 1709. According to one embodiment, the first external device (1603) may provide the weights of the first emotion recognition model updated through the learning operation to the electronic device (1601). Accordingly, the electronic device (1601) may recognize the emotion of the user using the weights provided from the first external device (1603).

[0245]

[0246] Figure 18 is a flowchart illustrating the emotion recognition operation of an emotion recognition system according to various embodiments. The operations of Figure 18 described below may represent various embodiments of operation 1509 of Figure 15.

[0247] Referring to FIG. 18, an electronic device (1601) (e.g., electronic device (310)) according to various embodiments may obtain data related to a user in operation 1811. According to one embodiment, the electronic device (1601) may obtain third data related to the user's biometric information and fourth data related to a speech input.

[0248] According to various embodiments, the electronic device (1601) may obtain a first prediction result (e.g., a first prediction result (303)) for the user's emotion based on the third data, the fourth data, and the first emotion recognition model (e.g., a first emotion prediction model (3133)) in operation 1813.

[0249] An electronic device (1601) according to various embodiments may provide third data, fourth data, and a first prediction result to a first external device (1603) (e.g., the first external device (320)) in operation 1815.

[0250] According to various embodiments, the first external device (1603) may obtain a second prediction result (e.g., a second prediction result (304)) for the user's emotion based on the third data, the fourth data, and the second emotion recognition model (e.g., a second emotion prediction model (3233)) in operation 1817.

[0251] According to various embodiments, the first external device (1603) can make a final inference about the emotion of the user based on the first prediction result obtained by the electronic device (1601) and the second prediction result obtained by the first external device (1603) in operation 1819.

[0252] According to various embodiments, the first external device (1603) may provide the final inference result (e.g., the final inference result (305)) to the electronic device (1601) in operation 1821.

[0253]

[0254] As described above, the electronic device (310) according to various embodiments can generate user-customized learning data (601) labeled (or mapped) with the emotions felt by the user and utilize this to train an emotion recognition model (e.g., a first emotion prediction model (3133)). Additionally or optionally, the electronic device (310) according to various embodiments can provide more specialized diagnostic results regarding the emotions felt by the user. This will be described in detail with reference to FIGS. 19 and 20 below.

[0255]

[0256] Figure 19 is a schematic diagram illustrating the configuration of an emotion recognition system according to various embodiments. Figure 20 is a diagram illustrating diagnostic results provided by an emotion recognition system according to various embodiments.

[0257] Referring to FIG. 19, an emotion recognition system (190) according to various embodiments may be composed of an electronic device (310) (e.g., electronic device (200)), a first external device (320), and a third external device (1910).

[0258] An electronic device (310) according to various embodiments may be composed of a sensor (311), a first memory (313), a first communication circuit (315), and a first processor (317). According to one embodiment, the configuration of the electronic device (310) illustrated in FIG. 19 may be similar or identical to the configuration illustrated in FIG. 3A, and thus a detailed description thereof may be omitted.

[0259] According to various embodiments, the first external device (320) may be composed of a second memory (323), a second communication circuit (325), and a second processor (327). According to one embodiment, the configuration of the first external device (320) illustrated in FIG. 19 may be similar or identical to the configuration illustrated in FIG. 3A, and thus a detailed description thereof may be omitted.

[0260] According to various embodiments, the electronic device (310) may provide an emotion recognition function through collaboration with a first external device (320). According to one embodiment, the electronic device (310) may obtain a first prediction result (303) regarding the user's emotion. In addition, the first external device (320) may obtain a second prediction result (304) regarding the user's emotion, and may make a final inference (305) regarding the user's emotion based on the first prediction result (303) and the second prediction result (304).

[0261] According to various embodiments, the electronic device (310) may utilize a first emotion recognition model (3130) learned by user-customized learning data (601) to obtain a first prediction result. According to one embodiment, the user-customized learning data (601) may be data obtained from a user who owns (or uses) the electronic device (310) and in which emotions felt by the user are labeled (or mapped).

[0262] According to various embodiments, the electronic device (310) can obtain the emotions felt by the user through a query message when generating user-customized learning data.

[0263] In this regard, the electronic device (310) can request a diagnosis for the user by providing data related to the emotions felt by the user to a third external device (1910).

[0264] According to one embodiment, the third external device (1910) may be at least one other electronic device capable of communicating with the electronic device (310) and operated by a professional authorized to treat the patient's condition (e.g., a qualified medical professional such as a doctor or nurse).

[0265] According to various embodiments, the third external device (1910) may provide diagnostic results for the user to the electronic device (310) based on data obtained from the electronic device (310).

[0266] In this regard, the electronic device (310) may provide a third external device (1910) with information regarding the need for a diagnosis of the user's emotions. Accordingly, the third external device (1910) may generate a diagnosis result including an opinion on the inquiry of the electronic device (310), as illustrated in FIG. 20, and provide the same to the electronic device (310).

[0267]

[0268] According to various embodiments, a method of operating an electronic device (310) may include an operation of acquiring first data (301) related to biometric information of a user and second data (401) related to a speech input of the user, an operation of identifying the user emotion related to the first data (301) and the second data (401), an operation of generating learning data (601) in which the user emotion is labeled for at least a portion of the first data (301) and the second data (401), an operation of utilizing the learning data (601) to learn an emotion recognition model (3130) stored in the electronic device (310), and an operation of inferring the emotion of the user based on the emotion recognition model (3130) learned with the learning data (601), third data related to the biometric information of the user, and fourth data related to the speech input of the user.

[0269] According to various embodiments, the operating method of the electronic device (310) may include an operation of obtaining a query message querying the user emotion from an external device (320) connected to the electronic device (310) through communication, an operation of outputting the query message through the electronic device (310), and an operation of identifying the user emotion related to the first data (301) and the second data (401) based on an input detected after outputting the query message.

[0270] According to various embodiments, the operating method of the electronic device (310) may include an operation of providing the stored emotion recognition model (3130) and the learning data (601) to an external device (320) that is connected to the electronic device (310) through communication, and an operation of acquiring the learned emotion recognition model (3130) from the external device (320).

[0271] According to various embodiments, the operating method of the electronic device (310) may include an operation of obtaining a plurality of weight values ​​of artificial neural network layers included in the learned emotion recognition model (3130) from the external device (320).

[0272] According to various embodiments, the operating method of the electronic device (310) may include an operation of providing a first prediction result (303) obtained based on the learned emotion recognition model (3130), the third data, and the fourth data to an external device (320) connected to the electronic device (310) through communication, and an operation of obtaining, from the external device (320), the user's emotion inferred based on the first prediction result and a second prediction result (304) obtained from the external device (320) based on the third data and the fourth data.

[0273] According to various embodiments, the method of operating the electronic device (310) may include an operation of obtaining content (1311-1315) related to the inferred user's emotion from an external device (320) that is connected to the electronic device (310) through communication.

[0274] According to various embodiments, the method of operating the electronic device (310) may include an operation of acquiring one of the first data (301) and the second data (401) from another external device (320) that is connected to the electronic device (310) through communication.

[0275] According to various embodiments, the method of operating the electronic device (310) may include an operation of obtaining the user emotion related to the first data (301) and the second data (401) from another external device (510) that is connected to the electronic device (310) through communication.

[0276] According to various embodiments, the electronic device (310) is implemented in a body-wearable form, and the operating method of the electronic device (310) may include an operation of acquiring the learned emotion recognition model (3130) from an external device (320) that is connected to the electronic device (310) through communication while the electronic device (310) is unworn by the user.

Claims

1. In an electronic device (310), At least one processor (317); and A memory (313) operatively connected to at least one processor and storing at least one instruction, The at least one instruction, when individually or collectively executed by the at least one processor, causes the electronic device to: Obtaining first data (301) related to the user's biometric information and second data (401) related to the user's speech input, Identifying the user sentiment related to the first data and the second data, Generating learning data (601) labeled with the user emotion for at least a portion of the first data and the second data, The above learning data is used to train the emotion recognition model (3130) stored in the above memory, An electronic device configured to infer the user's emotion based on an emotion recognition model learned with the above learning data, third data related to the user's biometric information, and fourth data related to the user's speech input.

2. In paragraph 1, display (240); and Further comprising a communication circuit (315), The at least one instruction, when individually or collectively executed by the at least one processor, causes the electronic device to: Obtaining a query message querying the user's emotions from an external device (320) through the above communication circuit, Output the above query message through the above display, An electronic device configured to identify the user emotion associated with the first data and the second data based on input detected after outputting the query message.

3. In paragraph 1, Further comprising a communication circuit (315), The at least one instruction, when individually or collectively executed by the at least one processor, causes the electronic device to: The above-mentioned stored emotion recognition model and the above-mentioned learning data are provided to an external device (320) through the above-mentioned communication circuit, An electronic device configured to obtain the learned emotion recognition model from the external device through the communication circuit.

4. In paragraph 3, The at least one instruction, when individually or collectively executed by the at least one processor, causes the electronic device to: An electronic device configured to obtain multiple weight values ​​of artificial neural network layers included in the learned emotion recognition model from the external device through the communication circuit.

5. In paragraph 1, Further comprising a communication circuit (315), The at least one instruction, when individually or collectively executed by the at least one processor, causes the electronic device to: The first prediction result (303) obtained based on the learned emotion recognition model, the third data, and the fourth data is provided to an external device (320) through the communication circuit, An electronic device configured to obtain, from the external device through the communication circuit, the user's emotion inferred based on the first prediction result and the second prediction result (304) obtained from the external device based on the third data and the fourth data.

6. In paragraph 1, Further comprising a communication circuit (315), The at least one instruction, when individually or collectively executed by the at least one processor, causes the electronic device to: An electronic device configured to obtain content (1311-1315) related to the user's emotions inferred above from an external device (320) through the communication circuit.

7. In paragraph 1, Further comprising a communication circuit (315), The at least one instruction, when individually or collectively executed by the at least one processor, causes the electronic device to: An electronic device configured to obtain one of the first data and the second data from another external device (510) through the communication circuit.

8. In paragraph 1, Further comprising a communication circuit (315), The at least one instruction, when individually or collectively executed by the at least one processor, causes the electronic device to: An electronic device configured to obtain the user emotion related to the first data and the second data from another external device (510) through the communication circuit.

9. In at least one of paragraphs 1 to 8, The above electronic device is an electronic device implemented in a body-wearable form.

10. In paragraph 3, The above electronic device is implemented in a body-wearable form, The at least one instruction, when individually or collectively executed by the at least one processor, causes the electronic device to: An electronic device configured to acquire the learned emotion recognition model from an external device through the communication circuit while the electronic device is unworn by the user.

11. In the operating method of the electronic device (310), An operation of acquiring first data (301) related to the user's biometric information and second data (401) related to the user's speech input; An operation of identifying the user emotion related to the first data and the second data; An operation of generating learning data (601) in which the user emotion is labeled for at least a portion of the first data and the second data; An operation of utilizing the above learning data to learn an emotion recognition model (3130) stored in the electronic device; A method comprising an operation of inferring the user's emotion based on an emotion recognition model learned with the above learning data, third data related to the user's biometric information, and fourth data related to the user's speech input.

12. In paragraph 11, An operation of obtaining a query message querying the user's emotions from an external device (320) connected to the electronic device through communication; An action of outputting the above query message through the electronic device; and A method comprising an action of identifying the user emotion related to the first data and the second data based on the input detected after outputting the query message.

13. In paragraph 11, An operation of providing the stored emotion recognition model and the learning data to an external device (320) connected to the electronic device through communication; and A method comprising an action of obtaining the learned emotion recognition model from the external device.

14. In paragraph 13, A method including an operation of obtaining multiple weight values ​​of artificial neural network layers included in the learned emotion recognition model from the external device.

15. In paragraph 11, An operation of providing a first prediction result (303) obtained based on the learned emotion recognition model, the third data, and the fourth data to an external device (320) connected to the electronic device through communication; and A method including an action of obtaining, from the external device, the user's emotion inferred based on the first prediction result and the second prediction result (304) obtained from the external device based on the third data and the fourth data.

Citation Information

Patent Citations

  • Emotion recognition method, device, computer equipment and storage medium

    CN112418059B

  • Artificial intelligence-based emotion monitoring method, system, device and storage medium

    CN116649980B

  • Method and apparatus for providing emotional digital twin

    KR102472786B1

  • Smart mirror

    KR102513289B1

  • A transformable chair

    KR102515741B1