Display device, machine learning device, inference device, information processing method, machine learning method, and inference method
The display device uses sensors and machine learning to generate and display emotion information, improving the natural interaction between humans and AI devices.
Patent Information
- Application Number
- JP2025037247
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2045-03-10
AI Technical Summary
Existing AI devices interact primarily through text, voice, or touch, lacking natural and familiar communication with humans.
A display device equipped with sensors that detect physical quantities, generating emotion information through machine learning and displaying it on a display unit, allowing for more natural interaction.
Enhances the feeling of familiarity and interaction with AI devices by dynamically changing facial expressions based on environmental and user interactions.
Smart Images

Figure 0007796446000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a display device, a machine learning device, an inference device, an information processing method, a machine learning method, and an inference method. [Background technology]
[0002] In recent years, AI (artificial intelligence) has become widely used as a tool in many fields, and is also used in devices such as smartphones, smart speakers, and home appliances (for example, Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2000-227799 Summary of the Invention [Problem to be solved by the invention]
[0004] The above-mentioned devices interact with people through text input, voice input, or touch operation, but communication using the five human senses is not necessarily sufficient. As AI becomes more widespread, there is a demand for interaction between people and AI to be more natural and feel more familiar.
[0005] The present invention has been made in light of the above-mentioned problems, and has an object to provide a device that can be made to feel more familiar. [Means for solving the problem]
[0006] In order to achieve the above object, a display device according to one aspect of the present invention includes a display unit, a control unit, and one or more sensors that detect a predetermined amount of physical quantity, wherein the control unit includes an emotion information generation unit that generates emotion information corresponding to the physical quantity based on the physical quantity detected by the sensor, and a display control unit that causes the emotion information generated by the emotion information generation unit to be displayed on the display unit. [Effects of the Invention]
[0007] According to the display device according to one aspect of the present invention, it is possible to create a more familiar feeling.
[0008] Problems, configurations, and effects other than those described above will become apparent from the detailed description of the invention that follows. [Brief explanation of the drawings]
[0009] [Figure 1] 1 is an external view showing an example of a wireless earphone case 1 equipped with a display device 10 according to an embodiment. [Figure 2] 1 is a block diagram showing an example of a functional configuration of a display device 10. FIG. [Figure 3] 4 is a flowchart showing an example of the operation of the display device 10. [Figure 4] 3 is a schematic diagram showing an example of a facial expression of an electronic pet displayed on the display unit 12. FIG. [Figure 5] 3 is a schematic diagram showing an example of a facial expression of an electronic pet displayed on the display unit 12. FIG. [Figure 6] FIG. 10 is a block diagram showing an example of a functional configuration of a display device 10 according to a second embodiment. [Figure 7] 10 is a flowchart showing an example of the operation of the display device 10 according to the second embodiment. [Figure 8] 2 is a schematic diagram showing examples of facial expressions of a user U and an electronic pet displayed on a display unit 12. FIG. [Figure 9] FIG. 9 is a hardware configuration diagram showing an example of a computer 900. DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, an embodiment for carrying out the present invention will be described with reference to the drawings. The scope necessary for the explanation to achieve the object of the present invention will be schematically shown, and the scope necessary for explaining the relevant part of the present invention will be mainly explained, and the parts that are omitted from the explanation will be based on publicly known techniques.
[0011] (First embodiment) An example of a wireless earphone case 1 equipped with a display device 10 according to the first embodiment will be described. Fig. 1 is an external view showing an example of a wireless earphone case 1 equipped with a display device 10 according to this embodiment. Fig. 2 is a block diagram showing an example of the functional configuration of the display device 10.
[0012] 1, the wireless earphone case 1 according to this embodiment includes wireless earphones 2 (2a, 2b) for the right and left ears, a roughly rectangular parallelepiped housing 3 with a housing section 3a that houses each of the wireless earphones 2 in a predetermined position, and a lid section 4 that slides open and close relative to the opening of the housing 3. As will be described later, the lid section 4 also functions as a display section 12 of the display device 10. The display section 12 displays facial expression information related to the facial expression of a virtual creature (hereinafter referred to as an electronic pet) that can interact with the user as emotional information indicating the emotion of the electronic pet.
[0013] The earphones 2 include a receiving unit that receives audio data transmitted from an electronic device such as a smartphone, and a playback unit that plays back the received data. The receiving unit of the earphones 2 may receive audio data transmitted from the earphone case 1.
[0014] The housing 3 has a built-in case battery (not shown) inside. This built-in case battery functions as a power source for charging the earphone built-in battery (not shown) built into the earphone 2. The storage section 3a of the housing 3 is provided with an output terminal for the built-in case battery, and the earphone 2 is provided with an input terminal for the built-in earphone battery. Every time the earphone 2 is stored in the storage section 3a of the housing 3, the output terminal in the storage section 3a comes into contact with the input terminal of the earphone built-in battery, and the earphone built-in battery is charged from the built-in case battery. The housing 3 may also have inside it a transmitting section that transmits audio data to the earphone 2 and a storage section that stores audio data.
[0015] In addition, the display device 10 mounted in the earphone case 1 includes a sensor 11, a display unit 12, a battery 13, a control unit 14, a data storage unit 15, and a trained model storage unit 16, as shown in FIG. 2.
[0016] Sensor 11 is a collection of at least one or more sensors that detect physical quantities in the external environment of earphone case 1. As will be described later, sensor 11 can be selected appropriately depending on the desired function, and examples include a temperature sensor that detects the temperature of housing 3, a humidity sensor that detects the humidity around earphone case 1, a light sensor that detects the amount of light around earphone case 1, a vibration sensor (including a six-axis sensor) that detects vibrations of earphone case 1, a pressure sensor that detects pressure applied to earphone case 1, and a human presence sensor that detects the presence or movement of a person around earphone case 1. The physical quantities detected by sensor 11 are stored in database 151 of data storage unit 15.
[0017] Display unit 12 includes a display device such as a liquid crystal display, an organic EL display, or a mini LED display, and is configured as the exterior of cover unit 4. As will be described later, display unit 12 displays facial expression information as emotional information.
[0018] Battery 13 is a power source for driving display device 10, and supplies power to sensor 11, display unit 12, and control unit 14. Battery 13 may be composed of a single battery that also functions as a power source for charging the earphone built-in battery described above, or may be composed of two or more batteries.
[0019] Control unit 14 of display device 10 comprises inference model learning unit 140, emotion information generation unit 141 that generates emotion information corresponding to physical quantities based on the physical quantities detected by sensor 11, and display control unit 142 that displays the generated emotion information on display unit 12. By executing information processing program 152 recorded in data storage unit 15, control unit 14 functions as inference model learning unit 140, emotion information generation unit 141, and display control unit 142, and displays emotion information according to the external environment on display unit 12.
[0020] The inference model learning unit 140 includes a learning data acquisition unit 140A and a machine learning unit 140B, and performs inference based on the acquired data. The learning data acquisition unit 140A references the database 151 and acquires multiple sets of learning data each consisting of input data and output data. As an example, the input data constituting the learning data is inference model input information based on a predetermined amount of physical quantity detected by the sensor 11. The output data constituting the learning data is emotion information corresponding to the physical quantity detected by the sensor 11.
[0021] The machine learning unit 140B performs machine learning to make the feature inference model 161 learn the correspondence between input data and output data using multiple sets of learning data acquired by the learning data acquisition unit 140A. The learned feature inference model 161 is stored in the learned model storage unit 16. The number of feature inference models 161 stored in the learned model storage unit 16 is not limited to one. For example, multiple feature inference models 161 with different conditions, such as differences in machine learning methods or data, may be stored and used selectively or in parallel. Note that, although the data storage unit 15 and the learned model storage unit 16 are shown as two storage units in FIG. 2, they may be configured as a single storage unit or three or more storage units.
[0022] The feature inference model 161 may be, for example, an inference model that applies a deep neural network, a convolutional neural network, a recurrent neural network, a support vector machine, or the like, but is not limited to these.
[0023] The learning method of the feature inference model 161 may be supervised learning, which uses data labeled with emotional information such as happiness or loneliness, or unsupervised learning, which automatically finds patterns from unlabeled data. Furthermore, reinforcement learning may be used, in which the model receives rewards through interactions with the user and improves its behavior.
[0024] Emotion information generation unit 141 generates emotion information corresponding to the detected physical quantities based on a trained inference model that has been trained by machine learning to learn the correspondence between the physical quantities detected by sensor 11 and emotion information. As emotion information, for example, smiling facial expression information corresponding to "happy" can be generated. In this way, by utilizing machine learning, it is possible to generate emotion information that cannot be captured by simple rules, thereby improving the diversity of emotional expression.
[0025] Emotion information generation section 141 may generate emotion information (or facial expression information) using a rule-based algorithm for acquiring emotion information based on multiple correspondences between physical quantities detected by sensor 11 and emotion information stored in data storage section 15. Generating emotion information using a rule-based system has the advantage of being able to instantly recognize specific patterns. Furthermore, by using machine learning and rule-based algorithms together, it is possible to generate emotion information appropriate to the situation.
[0026] Display control section 142 converts the emotion information generated by emotion information generation section 141 into visual information and displays it on display section 12. As an example, display control section 142 may represent the emotion information using patterns, figures, symbols, letters, or colors, or may represent it using animation, and display it on display section 12. For example, facial expression information as emotion information may be represented and displayed using patterns or figures that resemble eyes or a mouth, as will be described later.
[0027] The operation of the control unit 14 of the display device 10 according to this embodiment will be described below. FIG. 3 is a flowchart showing an example of the operation of the display device 10. As shown in FIG. 3, in step S101, the control unit 14 of the display device 10 according to this embodiment causes the learning data acquisition unit 140A of the inference model learning unit 140 to acquire multiple sets of learning data, each set of learning data being made up of input data consisting of a predetermined amount of physical quantity detected from the sensor 11 and output data consisting of emotion information corresponding to the physical quantity (learning data acquisition step). Then, the machine learning unit 140B uses the multiple sets of acquired learning data to cause the feature quantity inference model 161 to learn the correlation between the input data and the output data (machine learning step). The feature quantity inference model 161 is stored in the learned model storage unit 16 (learned model storage step).
[0028] In step S102 (physical quantity acquisition step), the control unit 14 acquires a predetermined amount of physical quantity detected by the sensor 11. At this time, preprocessing may be performed to ensure the quality of the acquired information by removing noise from the detection result from the sensor 11 or complementing missing data.
[0029] In step S103 (emotion information generation step), emotion information generation unit 141 extracts feature quantities necessary for emotion estimation from the physical quantities acquired in step S102 (inference processing step), and generates emotion information by inputting these into feature quantity inference model 161. At this time, emotion information generation unit 141 may, under certain circumstances, generate emotion information using a pre-defined rule-based system. Using a machine learning model and a rule-based system together also enables emotion estimation appropriate to the situation.
[0030] In step S104 (display step), display control section 142 converts the emotion information generated in step S103 (emotion information generation step) into visual emotion information and displays it on display section 12.
[0031] 4 and 5 are schematic diagrams showing examples of facial expressions of the electronic pet displayed on the display unit 12. For example, if it is determined that the user has touched the earphone case 1 based on the detection results of the human presence sensor, temperature sensor, etc. obtained in step S101 (information acquisition step), facial expression information corresponding to "woke up" is generated in step S103 (emotion information generation step). Then, as shown in FIG. 4(a), in step S104 (display step), the emotional information is visually converted, and the electronic pet with an expression that looks like it has woken up is displayed on the display unit 12.
[0032] Furthermore, if it is determined from the detection results of the pressure sensor, vibration sensor, etc. that the earphone case 1 has received a strong impact, emotional information corresponding to "knocked" is generated in step S103 (emotional information generation step). Then, in step S104 (display step), an electronic pet with its eyes closed and a patient expression is displayed on the display unit 12, as shown in Fig. 4(b).
[0033] Furthermore, if it is determined that it is night based on the detection results of the light sensor, etc., emotion information corresponding to "it has become dark" is generated in step S103 (emotion information generation step). Then, in step S104 (display step), the electronic pet with a tearful expression is displayed on the display unit 12, as shown in Fig. 4(c).
[0034] If the temperature sensor, humidity sensor, etc. detect that the environment is hot and humid, emotion information corresponding to "hot" is generated in step S103 (emotion information generation step). Then, in step S104 (display step), the electronic pet with a sweating expression is displayed on the display unit 12, as shown in Fig. 5(a).
[0035] Furthermore, if it is determined from the detection results of the Hall sensor, voltage sensor, etc. that the remaining charge of the earphone 2 is low, emotional information corresponding to "troubled" is generated in step S103 (emotional information generation process), and an electronic pet with a frowning expression is displayed on the display unit 12 in step S104 (display process), as shown in FIG. 5(b).
[0036] In this way, with the earphone case 1 (display device 10) of this embodiment, the facial expression of the electronic pet on the display unit 12 changes depending on the usage status of the earphone case 1, allowing the user U to feel more familiar with it when using it.
[0037] (Second embodiment) The earphone case 1 equipped with the display device 10 according to the second embodiment includes an imaging unit 17 that captures an image of the user's face, and an emotion information generation unit 141 generates emotion information based on the physical quantities detected by the sensor 11 and the captured facial image.
[0038] Fig. 6 is a block diagram showing the configuration of a display device 10 according to the second embodiment. As shown in Fig. 6, the display device 10 according to this embodiment includes a sensor 11, a display unit 12, an imaging unit 17, a battery 13, a control unit 14, a data storage unit 15, and a trained model storage unit 16. The configurations and operations other than the imaging unit 17 are the same as those in the above embodiment, and therefore the same reference numerals are used and detailed description thereof will not be repeated.
[0039] The display device 10 according to the second embodiment is characterized by including an imaging unit 17, an AI face recognition system, and a face tracking function.
[0040] For example, the imaging unit 17 captures an image in front of the earphone case 1 at a predetermined timing when a sensor (e.g., a motion sensor) detects a person. The configuration of the imaging unit 17 is not particularly limited, and a conventionally known camera may be used. Note that the imaging unit 17 may capture still images or videos.
[0041] Emotion information generation unit 141 is equipped with an AI face recognition system that detects a face area from the video captured by imaging unit 17 and extracts features from the detected face, comparing the extracted features with the facial features of user U pre-registered in data storage unit 15 to determine whether the detected user is the same person as the registered user U. Emotion information generation unit 141 also has a face tracking function that tracks the position of user U's face detected by imaging unit 17, and generates emotion information from the direction and angle of the face.
[0042] Display control section 142 visually converts the emotion information generated by emotion information generation section 141, and displays the movement of the eyes of the electronic pet so as to follow the movement of the eyes in the user U's facial image.
[0043] 7 is a flowchart showing an example of the operation of the display device 10. FIG. 7 is a schematic diagram showing an example of facial expressions of the user U and the electronic pet displayed on the display unit 12.
[0044] As shown in FIG. 7, first, in step S101, the sensor 11 detects the movement and presence of people in the vicinity.
[0045] In step S102 (imaging step), the imaging unit 17 captures an image of the area in front of the earphone case 1 at the timing when the presence of a nearby person is detected in step S110.
[0046] In step S103 (emotion information generation process), the emotion information generation unit 141 detects a face area from the video captured by the imaging unit 17, extracts features (facial landmarks and embedded vectors by deep learning) from the detected face, and compares these with the facial features pre-registered in the data storage unit 15 to determine whether the detected face is the same person as the registered user U.
[0047] If it is determined that the user is a registered user U, emotion information generation unit 141 tracks the position of the face detected in step S102, and generates emotion information consisting of the eye movements of the electronic pet from the direction and angle of the face.
[0048] In step S104 (display step), the display control unit 142 visually converts the emotion information into the eye movement of the electronic pet so as to follow the eye movement of the user U, and displays it on the display unit 12, for example.
[0049] For example, as shown in Fig. 8(a), when the user U is looking at the earphone case 1 from the left, the display control unit 142 displays on the display unit 12 an expression of the electronic pet looking to the left where the user U is. Also, as shown in Fig. 8(b), when the user U is looking at the earphone case 1 from the right, the display control unit 142 displays on the display unit 12 an expression of the electronic pet looking to the right where the user U is.
[0050] In this way, with the earphone case 1 (display device 10) of this embodiment, the facial expression of the electronic pet on the display unit 12 changes to follow the eye movements of the user U, whose facial image has been registered in advance, allowing the user U to feel more familiar with the device when using it.
[0051] (Device hardware configuration) 9 is a hardware configuration diagram showing an example of a computer 900. The display device 10 in this embodiment is implemented by a general-purpose or dedicated computer 900.
[0052] 7, the computer 900 includes, as its main components, a bus 910, a processor 912, a memory 914, an input device 916, an output device 917, a display device 918, and a storage device 920. Note that the above components may be omitted as appropriate depending on the application of the computer 900.
[0053] The processor 912 is composed of one or more arithmetic processing devices (such as a central processing unit (CPU), a micro-processing unit (MPU), a digital signal processor (DSP), or a graphics processing unit (GPU)), and operates as a control unit that controls the entire computer 900. The memory 914 stores various data and programs 930, and is composed of, for example, a volatile memory (such as a DRAM or SRAM) that functions as a main memory, a non-volatile memory (ROM), a flash memory, etc.
[0054] In the input device 916, the various sensors 11 function as input units, but in addition, for example, a keyboard, numeric keypad, electronic pen, microphone, etc. may function as input units. In the display device 918, the display unit 12 functions as an output device, but in addition, for example, a sound (audio) output device, vibration device, etc. may function as output units. The input device 916 and the display device 918 may be integrally configured, such as a touch panel display. The storage device 920 is configured, for example, with an HDD, SSD, etc., and functions as a storage unit. The storage device 920 stores various data required for the operating system and program execution.
[0055] In the computer 900 having the above configuration, the processor 912 loads a program 930 stored in the storage device 920 into the memory 914 to execute the program, and controls each unit of the computer 900 via the bus 910 .
[0056] In the computer 900 having the above configuration, the processor 912 loads a program 930 stored in the storage device 920 into the memory 914, executes the program, and controls each unit of the computer 900 via the bus 910. Note that the program 930 may be stored in the memory 914 instead of the storage device 920.
[0057] (Other embodiments) A machine learning device according to another embodiment includes a learning data acquisition unit 140A that acquires multiple sets of learning data each composed of input data and output data, a machine learning unit 140B that causes an inference model to learn a correlation between the input data and the output data using the multiple sets of learning data acquired by the learning data acquisition unit 140A, and a learned model storage unit 16 that stores a feature inference model 161 in which the correlation correspondence has been learned by the machine learning unit 140B. The input data is inference model input information based on a predetermined amount of physical quantity detected by one or more sensors 11, and the output data is emotion information.
[0058] A machine learning method according to another embodiment includes a learning data acquisition step of acquiring multiple sets of learning data each composed of input data and output data, a machine learning step of causing an inference model to learn a correlation between the input data and the output data using the multiple sets of learning data acquired in the learning data acquisition step, and a learned model storage step of storing a feature inference model 161 that has learned the correlation in the machine learning step. The input data is a predetermined amount of physical quantity detected by one or more sensors 11, and the output data is emotion information corresponding to the physical quantity.
[0059] This method includes step S103 (emotion information generation step) of generating emotion information corresponding to a predetermined physical quantity based on the physical quantity detected by one or more sensors 11, and step S104 (display step) of displaying the generated emotion information on a display unit.
[0060] An information processing method according to another embodiment comprises step S103 (emotion information generation step) of generating emotion information corresponding to a predetermined amount of physical quantity based on the physical quantity detected by one or more sensors 11, and step S104 (display step) of displaying the generated emotion information on display unit 12.
[0061] (Inference device and inference method) The present invention can be provided not only in the form of the display device 10 according to the above embodiment, but also in the form of an inference device or inference method used to infer feature quantities. In this case, the inference device or inference method can include a memory 914 and a processor 912, with the processor 912 executing a series of processes. The series of processes includes step S101 (information acquisition processing step) of acquiring a predetermined amount of physical quantity from one or more sensors 11, and step S102 (emotion information generation step) which is an inference process of inferring emotion information based on the predetermined amount of physical quantity acquired in step S101 (information acquisition processing step).
[0062] By providing it in the form of an inference device or an inference method, it can be applied to various devices more easily than when an information processing device is implemented. It will be naturally understood by those skilled in the art that when an inference device or an inference method infers features of emotion information, the inference method implemented by the feature inference model 161 may be applied using a trained inference model generated by the machine learning device and machine learning method according to the above embodiments.
[0063] Furthermore, in the above embodiment, display device 10 includes inference model learning section 140 and emotion information generation section 141, and emotion information generation section 141 generates emotion information using a trained model that has undergone machine learning, but this is not limited to this example. That is, emotion information generation section 141 may generate emotion information using a rule-based algorithm without using a trained model that has undergone machine learning. In this case, for example, the correspondence between the physical quantity detected by sensor 11 and the emotion information may be stored in advance in data storage section 15. Then, emotion information generation section 141 may be configured to refer to data storage section 15 when a physical quantity is detected by sensor 11, and generate the corresponding emotion information.
[0064] Furthermore, in the above embodiment, the display device 10 is mounted in the earphone case 1, but this is not limiting. That is, the display device 10 may be mounted in an item other than an earphone case. Specific examples include, but are not limited to, a pen case, a card case, or a wearable device.
[0065] The present invention is not limited to the above-described embodiment, and various additions and modifications can be made without departing from the spirit of the present invention, and all of these are included in the technical concept of the present invention. [Explanation of symbols]
[0066] 1...earphone case, 2...earphone, 3...casing, 4...lid, 10...display device, 11...sensor, 12...display unit, 13...battery, 14...control unit, 15...data storage unit, 16... trained model storage unit, 140... inference model learning unit, 140A... learning data storage unit, 140B... machine learning unit, 141... emotion information generation unit, 142...display control unit, 151...database, 152...information processing program, 161...Feature inference model
Claims
1. a display unit that displays a virtual creature that can interact with a user; A control unit; one or more sensors for detecting a predetermined amount of physical quantity in an external environment of the display unit; a storage unit that stores a correspondence between the physical quantity and emotion information that indicates an emotion of the virtual creature; The control unit an emotion information generation unit that, when it is determined that an impact has been applied to the display unit based on the physical quantity detected by the sensor, generates emotion information relating to a facial expression of the virtual creature corresponding to the physical quantity based on the physical quantity and the correspondence relationship stored in the storage unit; a display control unit that causes the generated emotion information to be displayed on the display unit.
2. a display unit that displays a virtual creature that can interact with a user; A control unit; one or more sensors for detecting a predetermined amount of physical quantity in an external environment of the display unit; a storage unit that stores a correspondence between the physical quantity and emotion information that indicates an emotion of the virtual creature; The control unit an emotion information generation unit that, when it is determined that an impact has been applied to the display unit based on the physical quantity detected by the sensor, generates emotion information relating to a facial expression of the virtual creature corresponding to the physical quantity based on a learned inference model that has learned the correspondence between the physical quantity and the storage unit by machine learning; a display control unit that causes the generated emotion information to be displayed on the display unit.
3. The display device according to claim 1 , wherein the sensor includes at least one of a temperature sensor, a humidity sensor, a light sensor, a vibration sensor, a pressure sensor, and a human presence sensor.
4. Equipped with an imaging unit, the imaging unit captures a facial image of the user; The display device according to claim 1 , wherein the emotion information generator generates the emotion information based on the physical quantity detected by the sensor and the captured face image.
5. The display device according to claim 4 , wherein the display control unit causes the display unit to display, as the emotion information, eye movements of the virtual creature that follow eye movements in the facial image of the user.
6. Further comprising a face image storage unit, The display device according to claim 4 , wherein, when a user in the captured face image is stored in advance in the face image storage section, the emotion information generation section generates emotion information based on the face image.
Citation Information
Patent Citations
Method for raising virtual pet through vehicle cloud, mobile terminal, storage medium and vehicle
CN117339220A
Communication robot system
JP2007181888A
Emotion determination apparatus and emotion determination method
JP2022006253A
Information processor, method for processing information, computer program, and learning system
JP2022064115A
Robot system, control method for robot system, and program
JP2022067744A