Data processing device and data processing program
The AI earphones with integrated biometric and emotional data analysis facilitate real-time emotional sharing, addressing the limitations of conventional AI earphones by providing effective communication tools for long-distance relationships and family members.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2025-01-07
- Publication Date
- 2026-07-17
Smart Images

Figure 2026119668000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to an AI earphone with an emotion sharing communication function, and particularly relates to a data processing device and a data processing program for sharing biometric information and emotional states in real time between users.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, in the prior art, when generating an utterance in response to the words spoken by a user, the actions taken by the user in daily life are not reflected, so there is room for improvement in proposing information corresponding to the content of the user's utterance.
[0005] For example, conventional AI earphones mainly provide information based on the life logs of individual users, and it is difficult to share real-time emotional states and exchange biometric information with others. More specifically, there is a lack of means for sharing emotions and mental states that are difficult to convey in words among long-distance lovers and family members living apart, and it has been a problem to reduce the psychological distance.
[0006] The present invention aims to solve the above problems by providing a data processing device and a data processing program that can engage in deep communication with multiple users wearing AI earphones by analyzing and sharing biometric information and emotional states in real time. [Means for solving the problem]
[0007] The data processing device relating to the technology disclosed herein comprises an input unit that inputs user data collected by an earphone including a microphone, speaker, biosensor, and camera; a processing unit that performs identification processing using a data generation model that analyzes the user's emotional state based on the user data; an output unit that transmits the results of the identification processing to the other user and outputs feedback corresponding to the other user's emotional state; and a communication unit that performs wireless communication with the other user's earphone. The processing unit further generates at least one of a vibration pattern, an audio message, and visual feedback according to the other user's emotional state, and the output unit outputs at least one of the generated vibration pattern, audio message, and visual feedback.
[0008] In this disclosure, the processing unit is characterized by generating at least one of the vibration pattern, voice message, and visual feedback by inputting a prompt into a data generation model that instructs the model to generate at least one of the vibration pattern, voice message, and visual feedback according to the emotional state of the other party user.
[0009] In this disclosure, the user data includes biometric information, audio data, and video data, and the processing unit analyzes the user's emotional state by inputting a prompt to the data generation model instructing it to integrate the biometric information, audio data, and video data to analyze the user's emotional state in a multidimensional manner.
[0010] The data processing program relating to the technology disclosed herein is characterized by causing a computer to operate as the above-mentioned data processing device.
[0011] According to the technology disclosed herein, the data processing device comprises an input unit that receives user data collected by earphones including a microphone, speaker, biosensor, and camera; a processing unit that performs specific processing using a data generation model; and an output unit that outputs the results of the specific processing. The input unit acquires the user's biometric information, voice, and video data, and the processing unit analyzes this data to estimate the emotional state. Furthermore, it transmits and receives data from another user via wireless communication and generates and outputs feedback corresponding to the other user's emotional state.
[0012] This allows users to share emotions and feelings that are difficult to put into words in real time, deepening communication in long-distance relationships and among family members, and bridging psychological distance. Furthermore, the technology, which combines biometric information analysis with wireless communication, can provide a new means of communication. [Brief explanation of the drawing]
[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the main functions of a data processing device and earphone according to the first embodiment. [Figure 3A] This figure shows an example of the configuration of earphones according to the first embodiment. [Figure 3B] This figure shows the state in which a user is wearing earphones according to the first embodiment. [Figure 3C] This is a diagram illustrating the field of view of the camera 42 according to the first embodiment. [Figure 3D] This figure shows the state in which a user is wearing earphones according to the first embodiment. [Figure 3E] This figure shows the state in which a user is wearing earphones according to the first embodiment. [Figure 3F] FIG. 1 shows a state where a user wears earphones according to the first embodiment. [Figure 4] FIG. 4 schematically shows a functional configuration of a specific processing unit of a data processing apparatus according to the first embodiment. [Figure 5] FIG. 7 schematically shows an example of an operation flow of specific processing by a data processing apparatus according to the first embodiment. [Figure 6] FIG. 10 is a conceptual diagram showing an example of a configuration of a data processing system according to the second embodiment. [Figure 7] FIG. 13 is a conceptual diagram showing an example of main functions of a data processing apparatus and earphones according to the second embodiment. [Figure 8] FIG. 16 shows an example of a configuration of earphones according to the second embodiment. [Figure 9] FIG. 19 shows a state where a user wears earphones according to the second embodiment. [Figure 10] FIG. 22 schematically shows a functional configuration of a specific processing unit of a data processing apparatus according to the second embodiment. [Figure 11] FIG. 25 is an operation flowchart of specific processing by a data processing apparatus according to the second embodiment, where (A) shows data collection processing and (B) shows feedback output processing.
MODE FOR CARRYING OUT THE INVENTION
[0014] (First Embodiment) Hereinafter, an example of a first embodiment of a data processing apparatus, a data processing method, and a program according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following first embodiment, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be one type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), or an APU (Accelerated Processing Unit).
[0017] In the first embodiment described below, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0018] In the first embodiment described below, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0019] In the following first embodiment, the coded communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0020] In the first embodiment described below, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0021] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0022] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and an AI earphone 14 (hereinafter simply referred to as earphone 14). An example of the data processing device 12 is a server. In the first embodiment, the data processing device 12 is an example of a "data processing device" related to the technology of this disclosure.
[0023] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0024] The earphone 14 includes a computer 36, a microphone 38, a speaker 40, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 38, speaker 40, and camera 42 are also connected to the bus 52.
[0025] The microphone 38 receives voice signals from the user 20 and accepts instructions from the user 20. The microphone 38 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 40 outputs audio according to the instructions from the processor 46. Hereafter, the microphone 38 may be simply referred to as the microphone 38.
[0026] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0027] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0028] Figure 2 shows an example of the main functions of the data processing device 12 and the earphone 14.
[0029] As shown in Figure 2, in the data processing device 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "data processing program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0030] The storage 32 stores the data generation model 58. The data generation model 58 is used by the specific processing unit 290.
[0031] (Earphones 14)
[0032] In the earphone 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0033] The earphone 14 can be interpreted as a canal-type earphone that is fitted into the ear canal of user 20 (see Figure 1), as shown in Figure 3A. However, the earphone 14 is not limited to a canal type; it may also be an inner-ear type earphone that is inserted into the inner ear of user 20, or a headphone type earphone that covers the entire ear of user 20. Each of the two earphones 14 is equipped with a microphone 38, a speaker 40, and a camera 42. The sound and images collected by the two earphones 14 fitted into user 20's ears may be recorded as a life log in the database 24.
[0034] The life log can be interpreted as a history of the user 20's actions in daily life, and may include sounds and images associated with the user 20, specifically sounds collected by the microphone 38 and images taken by the camera 42 during daily life. The life log may record sounds and images associated with the user 20, along with the date, time, and location in which they were acquired.
[0035] The sounds collected by the microphone 38 may include the voice of the person the user 20 is talking to, and sounds that occur around the user 20 while walking or cycling (such as the sound of cars driving, birds chirping, the babbling of a stream, and the sound of trees swaying in the wind).
[0036] As shown in Figure 3C, the camera 42 may capture images of the scenery within its field of view that is in front of the user 20, or it may capture images of scenery within its field of view that is not in front of the user 20, for example, to the side, behind, below, or above the user 20. The images captured by the camera 42 may include images of the person the user 20 is talking to, the scenery around the user 20 when they are walking or cycling, and images of the pet the user 20 is walking with.
[0037] Since each of the two earphones 14 is equipped with a camera 42, the two earphones 14 worn on the user's ears 20 are positioned at a specific distance apart, one on the left ear and the other on the right ear, as shown in Figure 3B. Therefore, compared to cases where two cameras are arranged side by side in a single housing, such as in a video camera, the spacing between the two cameras 42 can be increased, making 3D sensing easier. 3D sensing can be interpreted as measuring three-dimensional shapes.
[0038] Furthermore, when the two earphones 14 are placed in the user 20's ears, the two cameras 42 are positioned close to the user 20's left and right eyes, allowing images (captured images) that are nearly identical to those seen with the naked eye to be recorded as a life log in the database 24. Consequently, in specific processing, it becomes easier to reproduce information corresponding to inquiries from the user 20, that is, information corresponding to the content of the user 20's speech.
[0039] While the two earphones 14 are attached to the user 20, all or part of the images captured by the camera 42 may be recorded in the database 24 as a life log. Specifically, when the two earphones 14 are attached to the user 20, the recording of images captured by the camera 42 to the database 24 may begin, and when the two earphones 14 are removed from the user 20, the recording of those images to the database 24 may end.
[0040] While the two earphones 14 are worn by the user 20, all or part of the sound collected by the microphone 38 may be recorded as a lifelog in the database 24. Specifically, when the two earphones 14 are worn by the user 20, the recording of the sound collected by the microphone 38 to the database 24 may begin, and when the two earphones 14 are removed from the user 20, the recording of the sound to the database 24 may end.
[0041] Next, we will describe the processing of the specific processing unit 290 when the data processing device 12 receives an utterance from the user 20 wearing the earphones 14 regarding the user 20's memories or actions, and performs specific processing to propose information corresponding to the content of the user 20's utterance to the user 20.
[0042] (Specific processing) In the first embodiment, the identification process uses a data generation model that takes user data as input and generates predetermined inference results corresponding to the input user data. Specifically, in the identification process, when utterances related to the user 20's memories or actions are received as user data from a user 20 wearing earphones 14, the system refers to the database 24 and executes a process to propose information corresponding to the content of the utterances to the user 20. Specifically, after a life log is recorded in the database 24, if the user 20 wearing earphones 14 makes an utterance related to the user 20's memories or actions, the system may, as part of the identification process, refer to the database 24 and execute a process to propose information corresponding to the content of the utterances to the user 20.
[0043] (Example of specific processing) If the user wearing the earphones requests a message that will trigger the recall of a specific memory, the specific processing unit 290 may propose one or more messages selected based on the life log to the user who made the request, as information corresponding to the content of the utterance (request).
[0044] For example, if user 20, wearing earphones 14, tries to recall their memory and asks, "What did I say to person A around [date] at [time]?", the identification processing unit 290, as part of its identification process, inputs this message as a prompt to the data generation model 58. The identification processing unit 290 may refer to the life log in database 24 and, based on the output obtained from the data generation model 58, generate a message such as, "I think you said, 'I found a nice restaurant, let's make a reservation.'" This message may be interpreted as an example of information corresponding to the content of user 20's utterance.
[0045] For example, if user 20 wearing earphones 14 tries to recall their memory and asks, "Who was I talking to around [date] at [time]?", the identification processing unit 290, as part of its identification process, inputs this message as a prompt to the data generation model 58. The identification processing unit 290 may refer to the lifelog in database 24 and, based on the output obtained from the data generation model 58, generate a message such as, "It seems you were talking with two friends at that time, probably B and C." This message can be interpreted as an example of information corresponding to the content of user 20's utterance.
[0046] For example, if user 20, wearing earphones 14, tries to recall their emotions and says, "How did I feel when I was talking to person A around [date] at [time]?", the identification processing unit 290, as part of its identification process, inputs this message as a prompt to the data generation model 58. The identification processing unit 290 may refer to the life log in database 24 and, based on the output obtained from the data generation model 58, generate a message such as, "At that time, you were laughing a lot, so it seems you had a good impression of your friend and were very happy." This message may be interpreted as an example of information corresponding to the content of user 20's utterance.
[0047] (Example of specific processing, part 2) If a user 20 wearing earphones 14 mutters a specific matter as part of their utterance, the specific processing unit 290 may suggest to the user 20 who requested the message, based on their life log, recommended actions for the user 20 regarding that matter, as information corresponding to the content of their utterance (muttering).
[0048] For example, when user 20 wearing earphones 14 is shopping at a specific retail store and says, "What should I buy?", the specific processing unit 290 inputs this message as a prompt to the data generation model 58 as a specific processing step. The specific processing unit 290 may refer to the life log in the database 24 and, based on the output obtained by the data generation model 58, generate a message such as, "A few months ago, you purchased product A at this store and commented that it wasn't very tasty, so how about purchasing recently released products B and C this time?" This message may be interpreted as an example of information corresponding to the content of user 20's utterance.
[0049] (Third example of specific processing) As shown in Figure 3D, when user 20, wearing earphones 14, is operating a PC and says, "What was the name of product A that I searched for the day before yesterday?", the identification processing unit 290 inputs this message as a prompt to the data generation model 58 as part of its identification processing. The data generation model 58 refers to the life log in the database 24 and analyzes the video of the PC screen when user 20 was operating it in the past to generate a specific output. Based on the output obtained by the data generation model 58, the identification processing unit 290 may generate a message such as "Product A is ○○○". This message may be interpreted as an example of information corresponding to the content of user 20's utterance.
[0050] (Fourth example of specific processing) As shown in Figure 3E, if user 20, wearing earphones 14, says "There was a place nearby with a great view, but I wonder where it is?" while cycling, the identification processing unit 290 inputs this message as a prompt to the data generation model 58 as part of its identification process. The data generation model 58 refers to the life log in database 24 and analyzes places previously visited by user 20 and the route to those places to generate a specific output. Based on the output obtained by the data generation model 58, the identification processing unit 290 may generate a message such as "I think it's Cape XX, about 500m from here." This message can be interpreted as an example of information corresponding to the content of user 20's utterance.
[0051] (Example 5 of specific processing) As shown in Figure 3F, when user 20, wearing earphones 14, meets Mr. X at company A, the company he is visiting, and says, "Can you tell me this person's name?", the identification processing unit 290 inputs this message as a prompt to the data generation model 58 as part of the identification process. The data generation model 58 refers to the life log in database 24 and generates specific output from the history of people that user 20 met when he visited company A. Based on the output obtained from the data generation model 58, the identification processing unit 290 may generate a message such as, "I think his name is ○○." This message may be interpreted as an example of information corresponding to the content of user 20's utterance.
[0052] As shown in Figure 4, the specific processing unit 290 includes an input unit 291, a processing unit 292, and an output unit 293.
[0053] The input unit 291 acquires user input received through the earphone 14. Specifically, it acquires the user's voice received through the earphone 14.
[0054] The processing unit 292 performs specific processing using the data generation model 58. Specifically, it inputs voice from the user into the data generation model 58 and obtains a generation result. More specifically, when it receives an utterance from the user 20 wearing the earphones 14 regarding the user 20's memories or actions, it performs a specific processing step of proposing information corresponding to the content of the utterance to the user 20.
[0055] The output unit 293 transmits the result of the specific processing to the earphone 14. In the earphone 14, the control unit 46A causes the speaker 40 to output the result of the specific processing. The microphone 38 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0056] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0057] Next, the operation of the data processing system 10 will be explained.
[0058] An example of the flow of a specific processing method will be explained with reference to Figure 5. Note that the flow of a specific processing method shown in Figure 5 is an example of a "data processing method" related to the technology disclosed herein.
[0059] In step S300, the data processing device 12 receives user data, including sound and images collected by the two earphones 14.
[0060] In step S302, if the data processing device 12 receives an utterance from the user wearing the earphones 14 regarding the user's memories or actions, it executes a specific process to propose information to the user 20 that corresponds to the content of the utterance, based on the user's life log.
[0061] In step S303, the data processing device 12 executes a process to play back the result of a specific process from the speaker 40.
[0062] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0063] In the first embodiment described above, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which may be performed by multiple computers, including computer 22.
[0064] In the first embodiment described above, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0065] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0066] (Second Embodiment) A second embodiment is described below. In the second embodiment, the main focus is on specific processing such as dialogue performed between a user 20 and the earphones 14 worn by the user 20. However, in terms of positioning, while the first embodiment is specific processing between one user 20 and earphones 14, the second embodiment is characterized in that the specific processing is performed in the data processing system 10A (see Figure 6) under the circumstances in which two (or more) users 20A and 20B are each wearing a pair of earphones (14A and 14B).
[0067] In the second embodiment, the description of the configuration of the data processing device 12, which is identical to that of the first embodiment (in particular, the data processing device 12 that can communicate with the earphones 14A and 14B), will be omitted.
[0068] As shown in Figure 6, in the second embodiment, the data processing system 10A includes a data processing device 12 and earphones 14A worn by user 20A and user 20B. The earphones 14A, 14A indicate that each (individually) described as a pair of earphones 14 in the first embodiment operates (is controlled) independently.
[0069] In other words, as shown in Figure 8, the pair of earphones 14A and 14B according to the second embodiment have the same structure and are intended to be worn by two users 20A and 20B, respectively, as shown in Figure 9.
[0070] Earphones 14A and 14B can be interpreted as canal-type earphones that are fitted into the ear canals of users 20A and 20B, respectively. However, earphones 14A and 14B are not limited to canal-type earphones; they may also be inner-ear type earphones that are inserted into the inner ear of user 20, or headphone-type earphones that cover the entire ear of user 20.
[0071] In the second embodiment, since there were two users 20A and 20B, the pair of earphones 14A and 14B from the first embodiment were used. However, additional earphones with the same function may be added depending on the number of users.
[0072] As shown in Figure 6, earphones 14A and 14B are equipped with a computer 36, a microphone 38, a speaker 40, a camera 42, a biosensor 43, and a communication interface 44, respectively. The computer 36 is equipped with a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to the bus 52. The microphone 38, speaker 40, and camera 42 are also connected to the bus 52.
[0073] Since the functions of each device are the same as those of the first embodiment described using Figure 1, a detailed explanation is omitted here.
[0074] As shown in Figure 7, in the data processing device 12, specific processing is performed by the processor 28. The storage 32 stores a specific processing program 56A. The specific processing program 56A is an example of a "data processing program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56A from the storage 32 and executes the read specific processing program 56A on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290A according to the specific processing program 56A executed on the RAM 30.
[0075] Furthermore, as shown in Figure 7, the earphones 14A and 14B are processed for reception and output by the processor 46. The storage 50 stores the reception and output program 60. The processor 46 reads the reception and output program 60 from the storage 50 and executes the read reception and output program 60 on the RAM 48. The reception and output processing is realized by the processor 46 operating as a control unit 46A according to the reception and output program 60 executed on the RAM 48.
[0076] Earphones 14A and 14B are each capable of communicating with the data processing device 12, and data transmission and reception are performed via this data processing device 22.
[0077] As shown in Figure 10, the specific processing unit 290A of the data processing device 12 (see Figure 7) comprises an input unit 291A, a processing unit 292A, and an output unit 293A2.
[0078] (Input section 291A) The earphone input section 291A acquires the following data. (Data 1) Biometric information: The biosensors 43 installed in the earphones 14A and 14B acquire heart rate, skin electrical response, body temperature, etc. The biosensors 43 attached to the earphones 14A and 14B may also be linked with a separate device to acquire brain waves, sweat volume, etc. (Data 2) Audio data: Microphones 38 installed in earphones 14A and 14B acquire the content of speech, tone of voice, and speaking speed of users 20A and 20B. (Data 3) Video data: Cameras 42 installed in earphones 14A and 14B capture the facial expressions and eye movements of users 20A and 20B.
[0079] (Processing Unit 292A) The processing unit 292A performs the following processing using the data generation model. (Process 1) Data preprocessing: Denoise removal and data synchronization are performed. (Process 2) Analysis of emotional state: Analyze biometric information, voice, and video data to estimate the user's emotional state. At this time, the user's emotional state is analyzed by inputting prompts into the data generation model that instruct it to estimate the user's emotional state for each of the biometric information, voice, and video data. (Process 3) Integration of emotional data: Integrate each data point and quantify the overall emotional state. (Process 4) Data encryption and transmission: Encrypt and send emotional data to the recipient user. (Process 5) Receiving and analyzing the other party's data: Decode and analyze the data from the other party to identify their emotional state. (Process 6) Feedback generation: Generate appropriate feedback according to the other party's emotional state.
[0080] Alternatively, the user's emotional state may be analyzed by integrating (process 2) and (process 3) and inputting a prompt into the data generation model that instructs it to integrate biometric information, audio data, and video data to analyze the user's emotional state in a multidimensional manner.
[0081] (Output section 293A) The output unit 293A performs feedback by the following means. (Method 1) Vibration feedback: Control a vibration motor to transmit emotional states through touch. (Method 2) Voice message: The speaker 50 transmits the other party's emotional state via voice. (Method 3) Visual feedback: Express emotions by changing the color and flashing patterns of LED lights.
[0082] At this time, by inputting prompts into the data generation model that instruct it to generate vibration patterns, voice messages, and visual feedback according to the emotional state of the other user, at least one of these is generated.
[0083] The operation of the second embodiment will be described below. Figures 11(A) and 11(B) are control flowcharts illustrating the operation of the emotion-sharing communication function according to the second embodiment.
[0084] Figure 11(A) shows the data collection process using either earphone 14A or 14B (here, earphone 14A). In step 301A, for example, data collection is performed using earphone 14A. Data collection refers to acquiring biometric information, voice, and video data. For example, let's assume that earphone 14A is the one collecting the data.
[0085] In the next step, 302A, emotion analysis is performed. Emotion analysis involves the processing unit 292A of the specific processing unit 290A analyzing the data and estimating the emotional state. This is performed on the earphone 14A side.
[0086] In the next step, 303A, data transmission takes place. Data transmission involves encrypting emotional data and sending it to the recipient user. For example, it is sent from earphone 14A to earphone 14B.
[0087] In accordance with Figure 11(B), the processing on the earphone 14B side that receives the data transmission in step 303A will be explained.
[0088] Step 304A performs data reception and analysis. Data reception and analysis involves receiving and analyzing the data that was transmitted in step 303A. For example, data transmitted from earphone 14A is received by earphone 14B and collected.
[0089] In the next step, 305A, feedback generation is performed. Feedback generation is the process of generating feedback that is tailored to the other party's emotional state. For example, earphone 14A generates feedback.
[0090] In the next step, 306A, feedback output is performed. Feedback output is the process of providing feedback through vibration, sound, or visual means. For example, feedback is provided from earphone 14B to earphone 14A.
[0091] The processes in steps 301A to 303A in Figure 11(A) and steps 304A to 306A in Figure 11(B) are repeatedly executed at regular intervals. The assignment of earphones 14A and 14B to Figures 11(A) and 11(B) is determined by the correspondence between them and any of the voice commands or other questions made by users 20A or 20B.
[0092] For example, when user 20A (earphone 14A) says, "How is XX (the other person) doing (emotional state)?", user 20B's earphone 14B executes the process shown in Figure 11(A), and user 20A's earphone 14A executes the process shown in Figure 11(B).
[0093] In the second embodiment, data exchange between the pair of earphones 14A and 14B is performed via the data processing device 22. However, by giving each earphone 14A and 14B the function of the specific processing unit 290A of the data processing device 22, it is possible to perform direct data exchange between the earphones 14A and 14B. Also, in the second embodiment, the pair of earphones 14A and 14B (so-called stereo earphones) are worn by separate users 20A and 20B. However, two sets of earphones (stereo earphones) may be worn by separate users 20A and 20B.
[0094] (Example of the second embodiment) For example, if user 20A is relaxed and user 20B is stressed, user 20A's earphone 14A will emit gentle vibrations or warm voice messages to reassure user 20B.
[0095] Additionally, user 20B's earphone 14B receives user 20A's relaxed state and outputs gentle feedback.
[0096] According to the second embodiment described above, users 20A and 20B can share emotions in real time, making it possible to reduce the psychological distance even over long distances. Furthermore, by combining biometric information analysis with AI technology, it is possible to estimate the emotional state of the users with high accuracy and provide appropriate feedback.
[0097] The technology disclosed herein can be widely used in fields such as communication equipment, wearable devices, healthcare, and mental care.
[0098] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0099] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0100] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0101] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0102] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0103] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0104] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0105] Furthermore, the following additional information is disclosed regarding the above explanation.
[0106] (Note 1) An input unit that receives user data collected by earphones including a microphone, speaker, biosensor, and camera, A processing unit that performs specific processing using a data generation model that analyzes the user's emotional state based on the aforementioned user data, An output unit that transmits the result of the specified processing to the other user and outputs feedback corresponding to the other user's emotional state, The device includes a communication unit that performs wireless communication with the earphones of the other user, The processing unit further generates at least one of the following depending on the emotional state of the other user: a vibration pattern, an audio message, and visual feedback. The output unit is a data processing device that outputs at least one of the generated vibration pattern, voice message, and visual feedback.
[0107] (Note 2) The processing unit generates at least one of the vibration pattern, voice message, and visual feedback by inputting a prompt to a data generation model instructing it to generate at least one of the vibration pattern, voice message, and visual feedback according to the emotional state of the other party user, as described in Appendix 1.
[0108] (Note 3) The user data includes biometric information, audio data, and video data. The processing unit analyzes the user's emotional state by inputting a prompt to the data generation model instructing it to integrate the biometric information, the audio data, and the video data to analyze the user's emotional state in a multidimensional manner, as described in Appendix 1.
[0109] (Note 4) Computers, The data processing device described in any one of the appendices 1 to 3 is to be operated as such. Data processing program. [Explanation of Symbols]
[0110] 10 Data Processing Systems 12 Data Processing Devices 14A, 14B earphones 20A, 20B users 38 Microphones 40 speakers 42 cameras 43. Biosensors 290A Specific Processing Unit 291A Input Section 292A Processing Unit 293A output section< / url:>
Claims
1. An input unit that receives user data collected by earphones including a microphone, speaker, biosensor, and camera, A processing unit that performs specific processing using a data generation model that analyzes the user's emotional state based on the aforementioned user data, An output unit that transmits the result of the specified processing to the other user and outputs feedback corresponding to the other user's emotional state, The device includes a communication unit that performs wireless communication with the earphones of the other user, The processing unit further generates at least one of the following depending on the emotional state of the other user: vibration pattern, voice message, and visual feedback. The output unit is a data processing device that outputs at least one of the generated vibration pattern, voice message, and visual feedback.
2. The data processing apparatus according to claim 1, wherein the processing unit generates at least one of the vibration pattern, voice message, and visual feedback by inputting a prompt to a data generation model instructing it to generate at least one of the vibration pattern, voice message, and visual feedback according to the emotional state of the other party user.
3. The user data includes biometric information, audio data, and video data. The data processing device according to claim 1, wherein the processing unit analyzes the user's emotional state by inputting a prompt to the data generation model instructing it to integrate the biometric information, the audio data, and the video data to analyze the user's emotional state in a multidimensional manner.
4. Computers, The data processing device is operated according to any one of claims 1 to 3. Data processing program.