Data processing device, data processing method, data processing program, necklace-type terminal, data processing system, and earphone

The data processing device using earphones with microphones and cameras addresses the lack of relevance in conventional responses by suggesting information based on user life logs, enhancing personalization and accuracy.

WO2026009725A1PCT designated stage Publication Date: 2026-01-08SOFTBANK GROUP CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/022204
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-01-07
Filing Date
2025-06-19
Publication Date
2026-01-08

AI Technical Summary

Technical Problem

Conventional technologies fail to incorporate actions taken by users in their daily lives when generating responses to their utterances, lacking relevance in suggested information.

Method used

A data processing device and method that utilizes earphones with microphones and cameras to collect user data, employing a data generation model to suggest information based on the user's life log, including sounds and images, and provide responses relevant to their memories or behaviors.

Benefits of technology

Enhances the relevance of suggested information by incorporating user actions and experiences, providing personalized and contextually accurate responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025022204_08012026_PF_FP_ABST
    Figure JP2025022204_08012026_PF_FP_ABST
Patent Text Reader

Abstract

This data processing device comprises: an input unit that includes a microphone, a speaker, and a camera and inputs user data containing sounds and images collected through two earphones worn on the ears of a user; a processing unit that performs a specific process using a data generation model for generating a prescribed inference result corresponding to the user data; and an output unit that causes the result of the specific process to be reproduced through the speaker. The input unit inputs the sounds detected at the microphone and the images captured by the camera as the user data. Upon reception of an utterance related to the memory or the behavior of the user from the user wearing the earphones, the processing unit performs, as the specific process, a process for proposing information corresponding to the content of the utterance to the user on the basis of a life log of the user in which sounds and images associated with the user are recorded.
Need to check novelty before this filing date? Find Prior Art

Description

Data processing device, data processing method, data processing program, necklace-type terminal, data processing system, and earphones

[0001] The technology of the present disclosure relates to a data processing device, a data processing method, a data processing program, a necklace-type terminal, a data processing system, and earphones.

[0002] JP 2022-180282 A discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

[0003] However, in conventional technology, when generating utterances in response to words spoken by a user, the actions taken by the user in their daily lives are not reflected, so there is room for improvement in suggesting information that corresponds to the content of the user's utterance.

[0004] A first aspect of the technology of the present disclosure is a data processing device comprising: an input unit that inputs user data including sounds and images collected by two earphones that include a microphone, a speaker, and a camera and are worn on the user's ears; a processing unit that performs specific processing using a data generation model that generates a predetermined inference result according to the user data; and an output unit that plays the results of the specific processing from the speaker, wherein the input unit inputs sounds detected by the microphone and images captured by the camera as the user data, and when the processing unit receives an utterance from the user wearing the earphones regarding the user's memory or behavior, it performs the specific processing by suggesting information corresponding to the content of the utterance to the user based on the user's life log in which the sounds and images linked to the user are recorded.

[0005] A second aspect of the technology of the present disclosure is a data processing method in which a computer inputs user data including sounds and images collected by two earphones that include a microphone, a speaker, and a camera and are worn on the user's ears, and executes a specific process using a data generation model that generates a predetermined inference result corresponding to the user data.The data processing method includes inputting sounds detected by the microphone and images captured by the camera as the user data, and when an utterance regarding the user's memory or behavior is received from the user wearing the earphones, executing as the specific process a process of suggesting information corresponding to the content of the utterance to the user based on the user's life log in which the sounds and images linked to the user are recorded, and executing a process in which the computer plays back the results of the specific process from the speaker.

[0006] A third aspect of the technology of the present disclosure is a data processing program that causes a computer to execute a specific process using a data generation model that inputs user data including sound and images collected by two earphones that include a microphone, a speaker, and a camera and are worn on the user's ears, and generates a predetermined inference result corresponding to the user data.The data processing program inputs sound detected by the microphone and images captured by the camera as the user data, and when an utterance regarding the user's memory or behavior is received from the user wearing the earphones, the specific process is executed to suggest information corresponding to the content of the utterance to the user based on the user's life log in which the sound and the images linked to the user are recorded, and causes the computer to execute a process of playing back the results of the specific process from the speaker.

[0007] A fourth aspect of the technology of the present disclosure is a necklace-type terminal including a camera that captures images of the wearer's surroundings, a sensor that detects biometric data of the wearer, a microphone, and a collection unit that collects the outputs of the camera, the sensor, and the microphone.

[0008] A fifth aspect of the technology of the present disclosure is a data processing device that includes an input unit that includes a microphone, a speaker, and a camera and inputs audio data and image data collected by two earphones worn on a user's ears, a processing unit that performs a specific processing using a data generation model that generates a predetermined inference result according to the audio data and the image data, and an output unit that plays the result of the specific processing from the speaker, wherein the processing unit identifies an object that a user wearing the earphones is paying attention to by analyzing the image data, and executes the specific processing by emphasizing the sound from the object in the audio data.

[0009] A sixth aspect of the technology of the present disclosure is a data processing method in which a computer executes a specific process using a data generation model that inputs audio data and image data collected by two earphones that include a microphone, a speaker, and a camera and are worn on a user's ears, and generates a predetermined inference result according to the audio data and the image data.The data processing method inputs audio data detected by the microphone and image data captured by the camera, identifies an object that a user wearing the earphones is paying attention to by analyzing the image data, executes a process to emphasize the sound from the object in the audio data as the specific process, and plays back the result of the specific process from the speaker.

[0010] A seventh aspect of the technology of the present disclosure is a data processing program that causes a computer to execute a specific process using a data generation model that inputs audio data and image data collected by two earphones that include a microphone, a speaker, and a camera and are worn on a user's ears, and generates a predetermined inference result according to the audio data and the image data. The data processing program causes a computer to execute the following steps: inputting audio data detected by the microphone and image data captured by the camera; identifying an object that a user wearing the earphones is paying attention to by analyzing the image data, and executing a process to emphasize the audio from the object in the audio data as the specific process; and playing back the results of the specific process from the speaker.

[0011] An eighth aspect of the technology of the present disclosure is a data processing device that includes a processor, and that collects user data output from at least one of a microphone, a camera, and a sensor provided in at least one of two earphones worn on the left and right ears of a user, respectively, obtains two pieces of audio data having different contents based on the user data, and transmits one of the two pieces of audio data to one of the two earphones and transmits the other of the two pieces of audio data to the other of the two earphones.

[0012] A ninth aspect of the technology of the present disclosure is a data processing method, comprising: a computer collecting user data output from at least one of a microphone, a camera, and a sensor provided in at least one of two earphones worn on a user's left and right ears, respectively; obtaining two pieces of audio data having different contents based on the user data; and transmitting one of the two pieces of audio data to one of the two earphones and transmitting the other of the two pieces of audio data to the other of the two earphones.

[0013] A tenth aspect of the technology of the present disclosure is a data processing program that causes a computer to collect user data output from at least one of a microphone, a camera, and a sensor provided in at least one of two earphones worn on a user's left and right ears, obtain two pieces of audio data having different contents based on the user data, and transmit one of the two pieces of audio data to one of the two earphones and transmit the other of the two pieces of audio data to the other of the two earphones.

[0014] An eleventh aspect of the technology of the present disclosure is a data processing device comprising: an input unit that acquires user data including sounds picked up by the microphones and images taken by the cameras included in two earphones worn on the user's left and right ears, each earphone including a microphone, a speaker, and a camera; a processing unit that generates the story by inputting a prompt that instructs the data generation model to generate a story according to the user data into the data generation model; and an output unit that plays the generated story from the speakers, wherein the input unit sequentially acquires the user data from the two earphones, and the processing unit repeatedly generates continuations of the story by inputting a prompt that instructs the data generation model to generate a continuation of the story according to the user data most recently acquired by the input unit, thereby generating a series of stories.

[0015] A twelfth aspect of the technology of the present disclosure is a data processing method executed by a computer, which includes acquiring user data including sound picked up by a microphone and images taken by a camera contained in two earphones worn on the user's left and right ears, each earphone including a microphone, a speaker, and a camera, generating a story by inputting a prompt into a data generation model instructing the model to generate a story according to the user data, and playing the generated story from the speaker, wherein acquiring the user data includes sequentially acquiring the user data from the two earphones, and generating the story includes inputting a prompt into the data generation model instructing the model to generate a continuation of the story according to the user data acquired immediately before, thereby repeatedly generating continuations of the story, thereby generating a series of stories.

[0016] A thirteenth aspect of the technology of the present disclosure is a data processing program that causes a computer to execute the following steps: acquiring user data including sound picked up by a microphone and images taken by a camera contained in two earphones worn on the user's left and right ears, each earphone including a microphone, a speaker, and a camera; generating a story by inputting a prompt into a data generation model that instructs the model to generate a story according to the user data; and playing the generated story from the speaker; acquiring the user data includes sequentially acquiring the user data from the two earphones; generating the story includes inputting a prompt into the data generation model that instructs the model to generate a continuation of the story according to the user data acquired immediately before, thereby repeatedly generating continuations of the story, thereby generating a series of stories.

[0017] A fourteenth aspect of the technology disclosed herein is a data processing device comprising: an input unit that acquires user data; a processing unit that performs a specific processing using a data generation model that generates a predetermined inference result according to the user data; and an output unit that uses the result of the specific processing to play audio from speakers of two earphones, at least one of which has a vibration generating unit and is worn on the left and right ears of the user, wherein the input unit acquires, as the user data, learning content including learning content information that can be played back audio about learning content performed by the user; the processing unit performs the specific processing by inputting, to the data generation model, a prompt based on the user data, instructing the data generation model to generate vibration information that vibrates the vibration generating unit in synchronization with the timing of the audio playback of important learning points in the learning content indicated by the learning content information; and the output unit plays back the learning content indicated by the learning content information from the speaker of the earphones and vibrates the vibration generating unit using the vibration information.

[0018] A fifteenth aspect of the technology disclosed herein is a data processing method that acquires user data, performs a specification process using a data generation model that generates a predetermined inference result according to the user data, and uses the result of the specification process to play audio from speakers of two earphones, at least one of which has a vibration generating unit and is worn on the left and right ears of the user, one each. The data processing method is performed by a computer to acquire, as the user data, learning content including learning content information that can play audio of learning content performed by the user, and based on the user data, input a prompt to the data generation model that instructs the model to generate vibration information that vibrates the vibration generating unit in synchronization with the timing of audio playback of important learning points in the learning content indicated by the learning content information, thereby generating the vibration information as the specification process, and playing the learning content indicated by the learning content information from the speaker of the earphones and vibrating the vibration generating unit using the vibration information.

[0019] A sixteenth aspect of the technology of the present disclosure is a data processing program that causes a computer to execute a process of acquiring user data, performing a specific processing using a data generation model that generates a predetermined inference result according to the user data, and using the result of the specific processing to play audio from speakers of two earphones, at least one of which has a vibration generating unit and is worn on the left and right ears of the user, wherein learning content including learning content information that can play audio of learning content performed by the user is acquired as the user data, and based on the user data, a prompt is input to the data generation model that instructs the computer to generate vibration information that vibrates the vibration generating unit in synchronization with the timing of playing audio of important learning points in the learning content indicated by the learning content information, thereby performing the process of generating the vibration information as the specific processing, and playing the learning content indicated by the learning content information from the speaker of the earphones and vibrating the vibration generating unit using the vibration information.

[0020] A seventeenth aspect of the technology of the present disclosure is a data processing device comprising: an input unit that inputs user data including biometric information collected by two earphones that include a microphone, a speaker, a camera, a biometric information sensor, and a vibration applying unit and are worn on the user's ears; a processing unit that performs a specific processing using a data generation model that generates a predetermined inference result according to the user data; and an output unit that plays the result of the specific processing from the speaker and causes the vibration applying unit to apply vibration, wherein the input unit inputs the biometric information detected by the biometric information sensor as the user data, and the processing unit performs the specific processing by evaluating a state of the user based on the biometric information and inputting a prompt to the data generation model that instructs the data generation model to generate music data and vibration pattern data according to the evaluated state of the user.

[0021] An eighteenth aspect of the technology of the present disclosure is a data processing method in which a computer executes a specific process using a data generation model that inputs user data including biometric information collected by two earphones that include a microphone, a speaker, a camera, a biometric information sensor, and a vibration applying unit and are worn on the user's ears, and generates a predetermined inference result corresponding to the user data. The data processing method executes, as the specific process, a process in which biometric information detected by the biometric information sensor is input as the user data, a process in which the data generation model evaluates the user's condition based on the biometric information, and a prompt instructing the model to generate music data and vibration pattern data corresponding to the evaluated user's condition is input, and the result of the specific process is played back from the speaker and vibration is applied to the vibration applying unit.

[0022] A nineteenth aspect of the technology of the present disclosure is a data processing program that causes a computer to execute a specific process using a data generation model that inputs user data including biometric information collected by two earphones that include a microphone, a speaker, a camera, a biometric information sensor, and a vibration imparting unit and are worn on the user's ears, and generates a predetermined inference result corresponding to the user data. The data processing program inputs biometric information detected by the biometric information sensor as the user data, evaluates the user's condition based on the biometric information, and inputs a prompt to the data generation model instructing the model to generate music data and vibration pattern data corresponding to the evaluated user's condition, thereby executing the specific process to generate the music data and vibration pattern data as a result of the specific process, and causes the computer to execute a process of playing the result of the specific process from the speaker and imparting vibration to the vibration imparting unit.

[0023] A twentieth aspect of the technology of the present disclosure is an earphone comprising: a housing to be worn on the ear of a second transmitting user who communicates with a first receiving user; a data collection unit that collects operation data generated by detecting at least one of a swipe operation, a tap operation, and a long press operation on the housing; a processing unit that performs a specification process using a data generation model that generates a predetermined estimation result according to the operation data, and generates vibration data that causes an earphone worn on the ear of the first receiving user to generate a vibration according to the result of the specification process; and an output unit that transmits the vibration data to the earphone worn on the ear of the first receiving user, wherein the processing unit executes, as the specification process, a process of estimating at least one of an intention and an emotion of the second transmitting user based on the operation data, thereby generating a vibration that reproduces at least one of the intention and the emotion according to the result of the specification process in the earphone worn on the ear of the first receiving user.

[0024] A twenty-first aspect of the technology of the present disclosure is earphones comprising: a vibration unit that vibrates a housing attached to the ear of a first receiving user; an input unit that receives operation data generated by detecting at least one of a swipe operation, a tap operation, and a long press operation performed by a second sending user communicating with the first receiving user on earphones attached to the ear of the second sending user; a processing unit that performs identification processing using a data generation model that generates a predetermined estimation result according to the input operation data; and a control unit that causes the vibration unit to generate vibrations according to the result of the identification processing, wherein the processing unit executes, as the identification processing, a process of estimating at least one of an intention and an emotion of the second sending user based on the operation data, and the control unit causes the vibration unit to generate vibrations that reproduce at least one of the intention and the emotion according to the result of the identification processing.

[0025] A 22nd aspect of the technology of the present disclosure is a data processing device comprising: an input unit that receives operation data generated by detecting at least one of a swipe operation, a tap operation, and a long press operation performed by a second transmitting user communicating with a first receiving user wearing earphones on the earphones worn on the ears of the second transmitting user; a processing unit that performs identification processing using a data generation model that generates a predetermined estimation result according to the input operation data; a control unit that generates vibration data that causes the earphones worn on the ears of the first receiving user to generate vibrations according to the result of the identification processing; and an output unit that transmits the vibration data to the earphones worn on the ears of the first receiving user, wherein the processing unit executes, as the identification processing, a process of estimating at least one of an intention and an emotion of the second transmitting user based on the operation data, and the control unit generates the vibration data that reproduces at least one of the intention and the emotion according to the result of the identification processing.

[0026] A 23rd aspect of the technology of the present disclosure is a data processing method that performs processing including: performing, by at least one processor, a determination process using a data generation model that generates a predetermined estimation result according to operation data generated by detecting at least one of a swipe operation, a tap operation, and a long press operation on earphones worn in the ears of a second sending-side user who is communicating with a first receiving-side user wearing earphones, by the second sending-side user; generating vibration data that causes the earphones worn in the ears of the first receiving-side user to generate vibrations according to the result of the determination process; transmitting the vibration data to the earphones worn in the ears of the first receiving-side user; performing, as the determination process, a process that determines at least one of an intention and an emotion of the second sending-side user based on the operation data; and generating the vibration data that reproduces at least one of the intention and the emotion according to the result of the determination process.

[0027] A 24th aspect of the technology of the present disclosure is a data processing program that causes at least one processor to execute a process including: performing a determination process using a data generation model that generates a predetermined estimation result according to operation data generated by detecting at least one of a swipe operation, a tap operation, and a long press operation on earphones worn on the ears of a second sending-side user who is communicating with a first receiving-side user wearing earphones, by the second sending-side user; generating vibration data that causes the earphones worn on the ears of the first receiving-side user to generate vibrations according to the result of the determination process; transmitting the vibration data to the earphones worn on the ears of the first receiving-side user; executing, as the determination process, a process that determines at least one of an intention and an emotion of the second sending-side user based on the operation data; and generating the vibration data that reproduces at least one of the intention and the emotion according to the result of the determination process.

[0028] A 25th aspect of the technology of the present disclosure is a data processing device including an acquisition unit that acquires video data including video of a user's surroundings, an extraction unit that extracts visual information intended for an unspecified number of people from the video data, a database that stores user profile data that indicates characteristics or attributes of the user, a selection unit that selects the visual information based on the user profile data, and a message generation unit that generates a message including the content of the selected visual information using a data generation model.

[0029] A 26th aspect of the technology of the present disclosure is a data processing system including the data processing device described in the 25th aspect and earphones worn by the user, wherein the earphones include a camera that captures images of the user's surroundings and generates the video data, and a speaker that outputs the message, and the data processing device is a data processing system that includes a communication unit that transmits audio data including the content of the message to the earphones.

[0030] A 27th aspect of the technology of the present disclosure is a data processing method in which a computer acquires video data including images of a user's surroundings, extracts visual information intended for an unspecified number of people from the video data, selects the visual information based on user profile data that indicates the characteristics or attributes of the user, and uses a data generation model to generate a message that includes the content of the selected visual information.

[0031] A 28th aspect of the technology of the present disclosure is a program for causing a computer to execute a process of acquiring video data including images of a user's surroundings, extracting visual information intended for an unspecified number of people from the video data, selecting the visual information based on user profile data indicating the characteristics or attributes of the user, and using a data generation model, generating a message including the content of the selected visual information.

[0032] A twenty-ninth aspect of the technology of the present disclosure provides a device including an input unit that inputs user data including sound, images, and biometric data collected by two earphones that include a microphone, a speaker, a camera, and a biometric sensor and are worn on the ears of a user; a processing unit that performs specific processing using a data generation model that generates a predetermined inference result according to the user data; and an output unit that reproduces the result of the specific processing from the speaker, wherein the earphones are canal-type earphones or inner-ear-type earphones, and only the user wearing the earphones can hear the sound output from the earphones, and the earphones have a vibration function and are capable of outputting haptic feedback, and the input unit is a data processing device that inputs the sound detected by the microphone, the image captured by the camera, and the biometric data acquired by the biometric sensor as the user data, and when the processing unit receives an utterance from the user wearing the earphones regarding the user's memory or behavior, suggests information corresponding to the content of the utterance to the user based on the user's life log in which the sound and the image associated with the user are recorded, and performs the specific processing of acquiring the information corresponding to the content of the utterance and the tactile feedback by inputting a prompt that instructs the earphones to generate the tactile feedback into a data generation model.

[0033] A thirtieth aspect of the technology of the present disclosure is a data processing method in which a computer executes a specific process using a data generation model that inputs user data including sound, images, and biometric data collected by two earphones that include a microphone, a speaker, a camera, and a biometric sensor and are worn on the ears of a user, and generates a predetermined inference result according to the user data, wherein the earphones are canal-type earphones or inner-ear-type earphones, and only the user wearing the earphones can hear the sound output from the earphones, the earphones have a vibration function and are capable of outputting haptic feedback, and the sound detected by the microphone, the image captured by the camera, and the image acquired by the earphones are input. and the biometric data obtained by inputting ...

[0034] A thirty-first aspect of the technology of the present disclosure is a data processing program that causes a computer to execute a specific process using a data generation model that inputs user data including sound, images, and biometric data collected by two earphones that include a microphone, a speaker, a camera, and a biometric sensor and are worn on the ears of a user, and generates a predetermined inference result according to the user data, wherein the earphones are canal-type earphones or inner-ear-type earphones, and only the user wearing the earphones can hear the sound output from the earphones, the earphones have a vibration function and are capable of outputting haptic feedback, and the sound detected by the microphone, the image captured by the camera, and the image acquired by the earphones are input. and the biometric data obtained by inputting ...

[0035] A thirty-second aspect of the technology of the present disclosure is a data processing device comprising: an input unit that inputs user data collected by earphones including a microphone, a speaker, a biometric sensor, and a camera; a processing unit that performs specific processing using a data generation model that analyzes the emotional state of the user based on the user data; an output unit that transmits the result of the specific processing to a partner user and outputs feedback according to the emotional state of the partner user; and a communication unit that wirelessly communicates with the partner user's earphones, wherein the processing unit further generates at least one of a vibration pattern, a voice message, and visual feedback according to the emotional state of the partner user, and the output unit outputs the generated at least one of the vibration pattern, the voice message, and the visual feedback.

[0036] A thirty-third aspect of the technique of the present disclosure is a data processing program that causes a computer to operate as the data processing device according to the thirty-second aspect.

[0037] FIG. 1 is a conceptual diagram showing an example of the configuration of a data processing system. FIG. 2 is a conceptual diagram showing an example of the main functions of a data processing device and earphones. FIG. 3A is a diagram showing an example of the configuration of earphones. FIG. 3B is a diagram showing a state in which a user is wearing the earphones. FIG. 3C is a diagram for explaining the angle of view of a camera 42. FIG. 3D is a diagram showing a state in which a user is wearing the earphones. FIG. 3E is a diagram showing a state in which a user is wearing the earphones. FIG. 3F is a diagram showing a state in which a user is wearing the earphones. FIG. 4 is a schematic diagram showing the functional configuration of a specific processing unit of a data processing device. FIG. 5 is a schematic diagram showing an example of the operation flow of specific processing by the data processing device. FIG. 6 is a conceptual diagram showing an example of the configuration of a data processing system. FIG. 7 is a conceptual diagram showing an example of the main functions of a data processing device and a necklace-type terminal. FIG. 8 is a side view showing the configuration of the necklace-type terminal. FIG. 9 is a top view showing the configuration of the necklace-type terminal. FIG. 10 is a schematic diagram showing the functional configuration of the control unit of the necklace-type terminal. FIG. 11 is a schematic diagram showing the functional configuration of the specific processing unit of the data processing device. FIG. 12 is a schematic diagram showing an example of the operation flow of specific processing by the data processing device. FIG. 13 schematically illustrates an example of the operational flow of specific processing by a data processing device according to the third embodiment. FIG. 14 is a conceptual diagram illustrating an example of the configuration of a data processing system according to the fourth embodiment. FIG. 15 is a conceptual diagram illustrating an example of the operational flow of specific processing by a data processing device according to the fourth embodiment. FIG. 16 is a conceptual diagram illustrating an example of the configuration of a data processing system. FIG. 17 is a conceptual diagram illustrating an example of the operational flow of specific processing by a data processing device according to the fifth embodiment. FIG. 18 is a conceptual diagram illustrating an example of the configuration of a data processing system according to the sixth embodiment. FIG. 19 is a conceptual diagram illustrating an example of the operational flow of specific processing by a data processing device according to the sixth embodiment. FIG. 20 is a conceptual diagram illustrating an example of the configuration of a data processing system. FIG. 21 is a conceptual diagram illustrating an example of the operational flow of specific processing by a data processing device according to the seventh embodiment. FIG. 22 is a conceptual diagram illustrating an example of the configuration of a data processing system according to an 8-1 embodiment of the present disclosure. FIG. 23 is a conceptual diagram illustrating an example of the main functions of a data processing device and earphones according to an 8-1 embodiment. FIG. 24 schematically illustrates the functional configuration of a specific processing unit of a data processing device according to an 8-1 embodiment.FIG. 25 schematically illustrates an example of the operational flow of specific processing by a data processing device according to embodiment 8-1. FIG. 26 is a conceptual diagram illustrating an example of the configuration of a data processing system according to embodiment 8-2 of the present disclosure. FIG. 27 is a conceptual diagram illustrating an example of the main functions of a data processing device and earphones according to embodiment 8-2. FIG. 28 is a conceptual diagram illustrating an example of the functional configuration of a specific processing unit of a data processing device according to embodiment 8-2. FIG. 29 is a conceptual diagram illustrating an example of the operational flow of specific processing by a data processing device according to embodiment 8-2. FIG. 30 is a conceptual diagram illustrating an example of the configuration of a data processing system according to embodiment 8-3 of the present disclosure. FIG. 31 is a conceptual diagram illustrating an example of the main functions of a data processing device and earphones according to embodiment 8-3. FIG. 32 is a conceptual diagram illustrating an example of the functional configuration of a specific processing unit of a data processing device according to embodiment 8-3. FIG. 33 is a conceptual diagram illustrating an example of the operational flow of specific processing by a data processing device according to embodiment 8-3. FIG. 34 is a diagram illustrating the flow of various data transmitted and received between the earphones and the data processing device. FIG. 35 is a diagram illustrating an example of the functional configuration of a specific processing unit of a data processing device. FIG. 36 is a diagram schematically illustrating an example of the operational flow of identification processing by a data processing device. FIG. 37 is a diagram illustrating an example of the configuration of earphones. FIG. 38 is a conceptual diagram illustrating an example of the configuration of a data processing system. FIG. 39 is a conceptual diagram illustrating an example of the configuration of a data processing system according to the 11th embodiment. FIG. 40 is a conceptual diagram illustrating an example of the main functions of a data processing device and earphones according to the 11th embodiment. FIG. 41 is a diagram illustrating an example of the configuration of earphones according to the 11th embodiment. FIG. 42 is a diagram illustrating a state in which a user according to the 11th embodiment is wearing the earphones. FIG. 43 schematically illustrates the functional configuration of the identification processing unit of a data processing device according to the 11th embodiment. FIG. 44 is an operational flowchart of identification processing by a data processing device according to the 11th embodiment, where (A) illustrates data collection processing and (B) illustrates feedback output processing.

[0038] First Embodiment Hereinafter, an example of an embodiment of a data processing device, a data processing method, and a program according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0039] First, the terms used in the following description will be explained.

[0040] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a central processing unit (CPU), a graphics processing unit (GPU), a general-purpose computing on graphics processing unit (GPGPU), or an accelerated processing unit (APU).

[0041] In the following embodiments, a coded random access memory (RAM) is a memory in which information is temporarily stored and is used as a working memory by the processor.

[0042] In the following embodiments, the coded storage refers to one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0043] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), and Bluetooth (registered trademark).

[0044] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed by connecting them with "and / or."

[0045] FIG. 1 shows an example of the configuration of a data processing system 10 according to the embodiment.

[0046] 1, a data processing system 10 includes a data processing device 12 and earphones 14. An example of the data processing device 12 is a server. In this embodiment, the data processing device 12 is an example of a "data processing device" according to the technology of the present disclosure.

[0047] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a wide area network (WAN) and / or a local area network (LAN).

[0048] The earphones 14 include a computer 36, a microphone 38, a speaker 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 38, the speaker 40, and the camera 42 are also connected to the bus 52.

[0049] The microphone 38 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 38 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 40 outputs audio in accordance with instructions from the processor 46. Hereinafter, the microphone 38 may be simply referred to as the mic 38.

[0050] The camera 42 is a small digital camera equipped with an optical system including a lens, an aperture, and a shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of the user 20 (for example, an imaging range defined by an angle of view equivalent to the width of the field of vision of a typical healthy person).

[0051] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0052] FIG. 2 shows an example of the main functions of the data processing device 12 and the earphone 14.

[0053] As shown in FIG. 2 , in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32. The specific process program 56 is an example of a "data processing program" according to the technology of the present disclosure. The processor 28 reads the specific process program 56 from the storage 32 and executes the read specific process program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific process program 56 executed on the RAM 30.

[0054] The storage 32 stores a data generation model 58. The data generation model 58 is used by the specific processing unit 290.

[0055] (Earphones 14) In the earphones 14, the processor 46 performs reception output processing. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0056] 3A , the earphones 14 may be interpreted as canal-type earphones that are fitted into the ear canals of the user 20. Note that the earphones 14 are not limited to canal-type earphones, and may be inner-ear-type earphones that are fitted into the inner ears of the user 20, or headphone-type earphones that cover the entire ears of the user 20. Each of the two earphones 14 is provided with a microphone 38, a speaker 40, and a camera 42. Sounds and images collected by the two earphones 14 fitted into the ears of the user 20 may be recorded in the database 24 as a life log.

[0057] The life log may be interpreted as a history of actions taken by the user 20 in daily life, and may include sounds and images associated with the user 20, specifically, sounds collected by the microphone 38 in daily life and images taken by the camera 42. The life log may record sounds and images associated with the user 20 in association with the date, time, and place at which they were acquired.

[0058] The sounds collected by the microphone 38 may include the voice of the person with whom the user 20 is talking, sounds occurring around the user 20 while walking or cycling (such as the sound of cars passing by, birds chirping, the sound of a river flowing, and the sound of trees rustling in the wind).

[0059] 3C , the camera 42 may capture an image of scenery within an angle of view that captures what is in front of the user 20, or may capture an image of scenery within an angle of view that captures what is not in front of the user 20, for example, what is to the side, behind, below, or above the user 20. The image captured by the camera 42 may include the image of someone with whom the user 20 is talking, the image of the scenery around the user 20 when walking or cycling, the image of a pet walking with the user 20, etc.

[0060] Because each of the two earphones 14 is provided with a camera 42, the two earphones 14 worn on the ears of the user 20 are positioned a specific distance apart on the left and right ears, as shown in Figure 3B. Therefore, compared to when two cameras are arranged side by side in a single housing, such as a video camera, the distance between the two cameras 42 can be made wider, making 3D sensing easier. 3D sensing may be interpreted as measuring a three-dimensional shape.

[0061] Furthermore, when the two earphones 14 are attached to the ears of the user 20, the two cameras 42 are positioned close to the left and right eyes of the user 20, so that images (photographed images) that are substantially the same as those seen with the naked eye can be recorded as a life log in the database 24. Therefore, in the identification process, it becomes easier to reproduce information corresponding to an inquiry from the user 20, i.e., information corresponding to the content of the user 20's utterances.

[0062] While the two earphones 14 are attached to the user 20, all or part of the images captured by the camera 42 may be recorded as a life log in the database 24. Specifically, when the two earphones 14 are attached to the user 20, recording of the images captured by the camera 42 in the database 24 may start, and when the two earphones 14 are removed from the user 20, recording of the images in the database 24 may end.

[0063] While the two earphones 14 are worn by the user 20, all or part of the sounds collected by the microphone 38 may be recorded as a life log in the database 24. Specifically, when the two earphones 14 are worn by the user 20, recording of the sounds collected by the microphone 38 in the database 24 may start, and when the two earphones 14 are removed from the user 20, recording of the sounds in the database 24 may end.

[0064] Next, we will explain the processing of the specific processing unit 290 when the data processing device 12 performs specific processing to suggest information corresponding to the content of the user 20's utterance when it receives an utterance from the user 20 wearing the earphones 14 regarding the user's 20's memory or behavior.

[0065] (Identification Process) In the identification process of this embodiment, user data is input and an identification process is performed using a data generation model that generates a predetermined inference result according to the input user data. Specifically, in the identification process, when an utterance related to the memory or behavior of the user 20 is received as user data from the user 20 wearing the earphones 14, a process of suggesting information corresponding to the content of the utterance to the user 20 by referring to the database 24 is executed. Specifically, after a life log is recorded in the database 24, when the user 20 wearing the earphones 14 makes an utterance related to the memory or behavior of the user 20, the identification process may be a process of suggesting information corresponding to the content of the utterance to the user 20 by referring to the database 24.

[0066] (First Example of Identification Processing) When a user wearing earphones requests, as the content of their utterance, a message that will trigger a specific memory, the identification processing unit 290 may suggest one or more messages selected based on the life log to the user who requested the message as information corresponding to the content (request) of the utterance. For example, when a user wearing earphones requests, as the content of their utterance, a message that will trigger a specific memory, the identification processing unit 290 instructs the data generation model 58 to suggest, in a prompt, one or more messages selected based on the life log to the user.

[0067] For example, if the user 20 wearing the earphones 14 tries to recall his or her memory and utters, "What did I say to A on a certain date at around ____ time?", the identification processing unit 290 inputs the message as a prompt into the data generation model 58 as an identification process. The identification processing unit 290 may refer to the life log in the database 24 and generate a message such as, "I think he or she said, 'I found a nice restaurant, so let's make a reservation.'" based on the output obtained by the data generation model 58. The message may be interpreted as an example of information corresponding to the content of the user 20's utterance.

[0068] For example, if the user 20 wearing the earphones 14 tries to recall his or her memory and utters, "Who were you talking to at around XX on XX o'clock?", the identification processing unit 290 inputs the message as a prompt into the data generation model 58 as an identification process. The identification processing unit 290 may refer to the life log in the database 24 and generate a message such as, "It seems that at that time, two friends, probably Mr. B and Mr. C, were having a conversation," based on the output obtained by the data generation model 58. The message may be interpreted as an example of information corresponding to the content of the user 20's utterance.

[0069] For example, if the user 20 wearing the earphones 14 tries to recall his or her own feelings and utters, "How did I feel when I was talking to Mr. A on a certain date at around ____ time?", the identification processing unit 290 inputs the message as a prompt into the data generation model 58 as an identification process. The identification processing unit 290 may refer to the life log in the database 24 and generate a message such as, "You were laughing a lot at that time, so you seemed to like your friend and be very happy." based on the output obtained by the data generation model 58. The message may be interpreted as an example of information corresponding to the content of the user 20's utterance.

[0070] (Second Example of Identification Processing) When the user 20 wearing the earphones 14 tweets a specific matter as the content of the utterance, the identification processing unit 290 may suggest, to the user 20 who has requested the message, a recommended action of the user 20 with respect to the matter based on the life log, as information corresponding to the content of the utterance (tweet). For example, when the user 20 wearing the earphones 14 tweets a specific matter as the content of the utterance, the identification processing unit 290 instructs the data generation model 58 to suggest, in a prompt, to the user 20 a recommended action of the user 20 with respect to the matter based on the life log.

[0071] For example, if the user 20 wearing the earphones 14 utters, "What should I buy?" while shopping at a particular retail store, the identification processing unit 290, as an identification process, inputs the message as a prompt into the data generation model 58. The identification processing unit 290 may generate a message such as, "A few months ago, after purchasing product A at this store, you commented that it wasn't very tasty, so how about purchasing product B or product C, which were recently released, this time," based on the output obtained by the data generation model 58 by referring to the life log in the database 24. The message may be interpreted as an example of information corresponding to the content of the user 20's utterance.

[0072] (Third Example of Identification Processing) As shown in FIG. 3D , when a user 20 wearing earphones 14 utters, "What was the name of product A I searched for the day before yesterday?" while operating a personal computer, the identification processing unit 290 inputs the message as a prompt into the data generation model 58 as identification processing. The data generation model 58 generates a specific output by referencing the life log in the database 24 and analyzing images of the personal computer screen when the user 20 was operating the computer in the past. The identification processing unit 290 may generate a message such as "Product A is XXX" based on the output obtained by the data generation model 58. The message may be interpreted as an example of information corresponding to the content of the user 20's utterance.

[0073] (Fourth Example of Identification Processing) As shown in FIG. 3E , if the user 20 wearing the earphones 14 utters, "There's a place nearby with a spectacular view, but where is it?" while cycling, the identification processing unit 290 inputs the message as a prompt to the data generation model 58 as identification processing. The data generation model 58 generates a specific output by referencing the life log in the database 24 and analyzing places the user 20 has previously visited and the route to those places. Based on the output obtained by the data generation model 58, the identification processing unit 290 may generate a message such as, "I think Cape XX is 500 meters from here." This message may be interpreted as an example of information corresponding to the content of the user 20's utterance.

[0074] (Fifth Example of Identification Processing) As shown in FIG. 3F , when the user 20 wearing the earphones 14 meets Mr. X from Company A while visiting and utters, "What is his name?", the identification processing unit 290 inputs the message as a prompt into the data generation model 58 as identification processing. The data generation model 58 references the life log in the database 24 and generates a specific output based on the history of people the user 20 met while visiting Company A. The identification processing unit 290 may generate a message such as, "I think his name is XX" based on the output obtained by the data generation model 58. The message may be interpreted as an example of information corresponding to the content of the user 20's utterance.

[0075] As shown in FIG. 4, the specific processing unit 290 includes an input unit 291, a processing unit 292, and an output unit 293.

[0076] The input unit 291 acquires a user input received by the earphone 14. Specifically, the input unit 291 acquires the user's voice received by the earphone 14.

[0077] The processing unit 292 performs identification processing using the data generation model 58. Specifically, the processing unit 292 inputs a voice input by the user into the data generation model 58 and obtains a generation result. More specifically, when an utterance related to the memory or behavior of the user 20 is received from the user 20 wearing the earphones 14, the processing unit 292 performs the identification processing by suggesting information corresponding to the content of the utterance to the user 20.

[0078] The output unit 293 transmits the result of the specific processing to the earphone 14. In the earphone 14, the control unit 46A causes the speaker 40 to output the result of the specific processing. The microphone 38 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0079] The data generation model 58 is a so-called generative AI (artificial intelligence). Examples of the data generation model 58 include generative AI such as ChatGPT (Internet search <URL: https: / / openai.com / blog / chatgpt>) and Gemini (Internet search <URL: https: / / gemini.google.com / ?hl=ja>). The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0080] Next, the operation of the data processing system 10 will be described.

[0081] An example of the flow of the identification process will be described with reference to Fig. 5. The flow of the identification process shown in Fig. 5 is an example of a "data processing method" according to the technology of the present disclosure.

[0082] In step S300, the data processing device 12 receives user data including sounds and images collected by the two earphones 14.

[0083] In step S302, when the data processing device 12 receives a speech from a user wearing the earphones 14 regarding the memory or behavior of the user 20, the data processing device 12 executes a specific process to suggest information corresponding to the content of the speech to the user 20 based on the user's 20 life log.

[0084] In step S303, the data processing device 12 executes a process of reproducing the result of the specific process from the speaker 40.

[0085] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[0086] In the above embodiment, an example was given in which a specific process is performed by one computer 22, but the technology disclosed herein is not limited to this, and distributed processing of the specific process may be performed by multiple computers, including computer 22.

[0087] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[0088] In addition, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0089] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[0090] The hardware resource for executing a specific process can be any of the following types of processors. Examples of processors include a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), and ASICs (Application Specific Integrated Circuits), which are processors with a circuit configuration specifically designed to execute a specific process. Each of these processors has built-in or connected memory, and each processor executes a specific process by using the memory.

[0091] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[0092] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[0093] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[0094] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[0095] In addition, the following supplementary notes are provided in relation to the above description.

[0096] (Supplementary Note 1) A data processing device comprising: an input unit that inputs user data including sounds and images collected by two earphones that include a microphone, a speaker, and a camera and are worn on the user's ears; a processing unit that performs specific processing using a data generation model that generates a predetermined inference result according to the user data; and an output unit that plays the result of the specific processing from the speaker, wherein the input unit inputs sounds detected by the microphone and images captured by the camera as the user data, and when the processing unit receives an utterance from the user wearing the earphones regarding the user's memory or behavior, it performs, as the specific processing, a process of suggesting information corresponding to the content of the utterance to the user based on the user's life log in which the sounds and images associated with the user are recorded.

[0097] (Supplementary Note 2) The data processing device according to Supplementary Note 1, wherein when the user wearing the earphones requests a message that will trigger a specific memory as the content of the utterance, the processing unit suggests to the user one or more messages selected based on the life log as information corresponding to the content of the utterance.

[0098] (Supplementary Note 3) The data processing device according to Supplementary Note 1 or 2, wherein when the user wearing the earphones tweets a specific matter as the content of the speech, the processing unit suggests to the user recommended actions for the matter based on the life log as information corresponding to the content of the speech.

[0099] (Supplementary Note 4) A data processing method in which a computer executes a specific processing using a data generation model that inputs user data including sounds and images collected by two earphones that include a microphone, a speaker, and a camera and are worn on the user's ears, and generates a predetermined inference result according to the user data, wherein the computer inputs sounds detected by the microphone and images captured by the camera as the user data, and when an utterance regarding the user's memory or behavior is received from the user wearing the earphones, the specific processing is executed to suggest information corresponding to the content of the utterance to the user based on the user's life log in which the sounds and images linked to the user are recorded, and the result of the specific processing is played back from the speaker.

[0100] (Supplementary Note 5) A data processing program that causes a computer to execute a specific process using a data generation model that inputs user data including sounds and images collected by two earphones that include a microphone, a speaker, and a camera and are worn on the user's ears, and generates a predetermined inference result according to the user data, wherein the data processing program inputs sounds detected by the microphone and images captured by the camera as the user data, and when an utterance regarding the user's memory or behavior is received from the user wearing the earphones, the specific process executes a process of suggesting information corresponding to the content of the utterance to the user based on the user's life log in which the sounds and images linked to the user are recorded, and plays back the results of the specific process from the speaker.

[0101] Second Embodiment An example of an embodiment of a data processing device, a data processing method, and a program according to the technology of the present disclosure will be described below with reference to the accompanying drawings. Note that the same parts as those in the first embodiment are designated by the same reference numerals, and detailed description thereof will be omitted.

[0102] FIG. 6 shows an example of the configuration of a data processing system 10 according to the embodiment.

[0103] As shown in Fig. 6, the data processing system 10 includes a necklace-type terminal 14A instead of the earphones 14 in Fig. 1. An example of the data processing device 12 is a server. In this embodiment, the data processing device 12 is an example of a "data processing device" according to the technology of the present disclosure, and the necklace-type terminal 14A is an example of a "necklace-type terminal" according to the technology of the present disclosure.

[0104] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a wide area network (WAN) and / or a local area network (LAN).

[0105] The necklace-type terminal 14A includes a computer 36, a microphone 38, a sensor 39, a speaker 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 38, the speaker 40, and the camera 42 are also connected to the bus 52.

[0106] The user 20 who wears the necklace-type terminal 14A may be, for example, a patient whose health condition is to be diagnosed, or may be a regular user.

[0107] The microphone 38 picks up the voice uttered by the user 20 who is wearing the necklace-type terminal 14A, as well as sounds around the user 20. The microphone 38 also receives instructions and the like from the user 20 by receiving the voice uttered by the user 20. The microphone 38 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 40 outputs audio in accordance with instructions from the processor 46. The speaker 40 is, for example, a directional speaker, and outputs audio toward the ears of the user 20.

[0108] The sensor 39 is a sensor that detects biological data of the user 20 who is wearing the necklace-type terminal. For example, the sensor 39 is a heart rate sensor or a blood oxygen sensor.

[0109] The camera 42 is a small digital camera equipped with an optical system including a lens, an aperture, and a shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of the user 20 (for example, an imaging range defined by an angle of view equivalent to the width of the field of vision of a typical healthy person).

[0110] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0111] FIG. 7 shows an example of the main functions of the data processing device 12 and the necklace-type terminal 14A.

[0112] 7 , in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32. The processor 28 reads the specific process program 56 from the storage 32 and executes the read specific process program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific process program 56 executed on the RAM 30.

[0113] The storage 32 stores a data generation model 58. The data generation model 58 is used by the specification processing unit 290. The storage 32 also includes a data accumulation unit .

[0114] In necklace-type terminal 14A, data collection processing is performed by processor 46. A data collection program 60 is stored in storage 50. Processor 46 reads data collection program 60 from storage 50 and executes the read data collection program 60 on RAM 48. The data collection processing is realized by processor 46 operating as control unit 46A in accordance with the data collection program 60 executed on RAM 48.

[0115] As shown in FIGS. 8 and 9 , the necklace type terminal 14A includes multiple microphones 38, multiple sensors 39, multiple speakers 40, and multiple cameras 42. FIGS. 8 and 9 show an example in which two microphones 38 are arranged so as to be located in front of the user 20 when the user 20 wears the necklace type terminal 14A. Also shown is an example in which two sensors 39 are arranged so as to be located on the right and left sides of the user 20 when the user 20 wears the necklace type terminal 14A. An example is shown in which two speakers 40 are arranged so as to be located on the right rear and left rear sides of the user 20 when the user 20 wears the necklace type terminal 14A. An example is shown in which two cameras 42 are arranged so as to be located on the right front and left front sides of the user 20 when the user 20 wears the necklace type terminal 14A. An example is shown in which two sensors 39 are arranged inside the necklace type terminal 14A so as to come into contact with the neck of the user 20 when the user 20 wears the necklace type terminal 14A.

[0116] Next, the processing of the control unit 46A when the necklace-type terminal 14A performs a data collection process for collecting data will be described.

[0117] In the data collection process of this embodiment, biometric data of the user is collected in real time. Furthermore, not only biometric data but also all situational data surrounding the user is collected. This makes it possible to detect early signs of, for example, Alzheimer's disease and dementia. It also makes it possible to monitor the user's health condition (e.g., heart disease).

[0118] As shown in FIG. 10, the control unit 46A includes a data collection unit 100 and a communication unit 102.

[0119] The data collection unit 100 collects the output of each of the microphone 38 , the sensor 39 , and the camera 42 .

[0120] The communication unit 102 transmits the outputs of the microphone 38 , the sensor 39 , and the camera 42 collected by the data collection unit 100 to the data processing device 12 .

[0121] Next, the processing of the identification processing unit 290 when the data processing device 12 performs identification processing to acquire a response corresponding to a user utterance will be described.

[0122] In the identification process of this embodiment, a response corresponding to a user utterance picked up by the microphone 38 of the necklace-type terminal 14A is acquired using the data generation model 58.

[0123] As shown in FIG. 11, the specific processing unit 290 includes an input unit 292, a processing unit 294, and an output unit 296.

[0124] The input unit 292 stores the outputs of the microphone 38, the sensor 39, and the camera 42 received from the necklace-type terminal 14A in the data storage unit 54.

[0125] The input unit 292 acquires the user's utterance received by the necklace type terminal 14 A. Specifically, the input unit 292 acquires the user's utterance picked up by the microphone 38 of the necklace type terminal 14 A.

[0126] The processing unit 294 performs a specific process using the data generation model 58. Specifically, a prompt including a user utterance is input to the data generation model 58 to obtain a generation result. At this time, the prompt may further include outputs from the sensor 39 and the camera 42 collected by the data collection unit 100.

[0127] The output unit 296 transmits the result of the identification process to the necklace-type terminal 14A. In the necklace-type terminal 14A, the control unit 46A causes the speaker 40 to output the result of the identification process. In this way, a response corresponding to the user utterance picked up by the microphone 38 is output to the user 20 by the speaker 40. The microphone 38 further acquires the user utterance in response to the result of the identification process. The control unit 46A transmits audio data indicating the user utterance acquired by the microphone 38 to the data processing device 12. In the data processing device 12, the identification processing unit 290 acquires the user utterance.

[0128] The data generation model 58 is a so-called generative AI (artificial intelligence). Examples of the data generation model 58 include generative AI such as ChatGPT (Internet search <URL: https: / / openai.com / blog / chatgpt>) and Gemini (Internet search <URL: https: / / gemini.google.com / ?hl=ja>). The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0129] The outputs of the microphone 38, the sensor 39, and the camera 42 stored in the data storage unit 54 are used, for example, to diagnose the health condition of the user 20. In this case, the outputs of the microphone 38, the sensor 39, and the camera 42 stored in the data storage unit 54 may be transmitted to a terminal on the medical institution side. Alternatively, the data processing device 12 may analyze the outputs of the microphone 38, the sensor 39, and the camera 42 stored in the data storage unit 54 to diagnose the health condition of the user 20.

[0130] Next, the operation of the data processing system 10 will be described.

[0131] First, an example of the flow of the data collection process will be described.

[0132] When the user 20 is wearing the necklace-type terminal 14A, the data collection unit 100 sequentially collects the outputs of the microphone 38, the sensor 39, and the camera 42. The communication unit 102 sequentially transmits the outputs of the microphone 38, the sensor 39, and the camera 42 collected by the data collection unit 100 to the data processing device 12.

[0133] Next, an example of the flow of the identification process will be described with reference to Fig. 12. Here, it is assumed that the input unit 292 of the data processing device 12 sequentially acquires the outputs of the microphone 38, the sensor 39, and the camera 42 received from the necklace-type terminal 14A and stores them in the data accumulation unit 54.

[0134] In step S300, the processing unit 294 determines whether a predetermined trigger condition is satisfied. Specifically, the trigger condition may be that a specific word (e.g., the name of an agent installed in the necklace-type terminal 14A) or phrase (e.g., “Hi! XXX” (XXX is the name of the agent)) is included in the user utterance picked up by the microphone 38.

[0135] If the trigger condition is met in step S300 (step S300; Yes), the data processing system 10 proceeds to step S301. On the other hand, if the trigger condition is not met in step S300 (step S300; No), the data processing system 10 ends the identification process.

[0136] In step S301, the processing unit 294 generates a prompt by adding an instruction sentence for obtaining a result of a specific process to text representing a user utterance picked up by the microphone 38.

[0137] For example, a prompt such as "The user is saying the following: XXX. Please respond as an agent" (XXX is the user utterance) can be generated. Alternatively, the outputs of the sensor 39 and the camera 42 can be added to the prompt to generate a prompt such as "This is biometric data representing the user's heart rate and video data representing the user's surroundings. The user is also saying the following: XXX. Please respond as an agent" (XXX is the user utterance).

[0138] In step S303, the processing unit 294 inputs the generated prompt to the data generation model 58, and obtains the result of the specific process based on the output of the data generation model 58.

[0139] In step S304, the output unit 296 outputs the result of the identification process to the necklace-type terminal 14A, and the identification process ends.

[0140] Third Embodiment An example of an embodiment of a data processing device, a data processing method, and a program according to the technology of the present disclosure will be described below with reference to the accompanying drawings. Note that the same parts as those in the first embodiment are designated by the same reference numerals, and detailed description thereof will be omitted.

[0141] In this embodiment, the lens closest to the subject among the lenses constituting the camera 42 may be an ultra-wide-angle lens such as a fisheye lens. Furthermore, if the camera 42 has an ultra-wide-angle lens, the camera 42 can also capture an image of the eyes of the user 20.

[0142] In the third embodiment, the camera 42 included in the earphone 14 has an ultra-wide-angle lens and is capable of capturing an image of the user's eyes.

[0143] Next, the specification process performed by the specification processing unit 290 of the data processing device 12 in the third embodiment will be described.

[0144] In the identification process in the third embodiment, voice data and image data collected by the earphones 14 are input, and identification process is performed using a data generation model that generates data for generating a predetermined inference result according to the input voice data and image data. Specifically, in the identification process, the image data is analyzed to identify an object that the user is paying attention to, and a process is performed to emphasize the sound from the object in the voice data.

[0145] The identification processing unit 290 extracts the user's pupil area from the two image data acquired by the cameras 42 of the left and right earphones 14 and derives the user's gaze. The gaze can be derived using three-dimensional gaze measurement technology. The identification processing unit 290 may also derive the user's gaze movement and blink frequency. The derived user's gaze is expressed by pixel positions in the image data.

[0146] After deriving the user's gaze, the identification processing unit 290 inputs the derived result of the user's gaze, the audio data, the image data, and a prompt to the data generation model 58, saying, "From the image data, identify an object in the direction of the user's gaze, and from the audio data, generate audio data in which the audio emitted by the object is emphasized." The data generation model 58 identifies an object in the user's gaze direction in the image represented by the image data. In this case, the data generation model 58 extracts all objects included in the image and identifies, among all the objects, the object in the user's gaze direction as the object.

[0147] The specific processing unit 290 uses the data generation model 58 to analyze the voice data and separate the voice uttered by the target from the voice data. One method is to use AI model-based methods such as those proposed in "https: / / crystal-method.com / solutions / sound-source-separation-system / " or "https: / / crystal-method.com / solutions / sound-source-separation-system / ." The former is a method that treats sounds other than the specific person's voice as noise and provides the specific person's voice to an AI model to separate the specific person's voice from a sound in which multiple people are speaking. The latter is a method that separates a specific sound from a sound containing a mixture of various sounds by training an AI model to specific sounds such as human voices, machine sounds, and instrument sounds.

[0148] The identification processing unit 290 uses the data generation model 58 to emphasize the separated audio, thereby enhancing the target audio. For example, audio data in which the separated audio is emphasized is generated by increasing the gain of the separated audio, decreasing the gain of audio other than the separated audio, or increasing the gain of the separated audio and decreasing the gain of audio other than the separated audio. Note that decreasing the gain of audio other than the separated audio essentially performs a process to reduce noise. The identification processing unit 290 may also optimize the sound quality of the emphasized audio so that it is easier to hear.

[0149] When emphasizing the voice, the specific processing unit 290 may adjust the parameters depending on the user's situation. For example, the specific processing unit 290 may additionally input a prompt to the data generation model 58, such as "determine whether the user is concentrating or distracted based on the user's eye movements or the number of blinks, and if the user appears distracted, increase the gain to emphasize the voice."

[0150] The specific processing unit 290 outputs the audio data in which the separated audio has been enhanced to the earphone 14. The user can hear the enhanced audio from the earphone 14.

[0151] Next, the operation of the data processing system 10 in the third embodiment will be described.

[0152] An example of the flow of the identification process will be described with reference to Fig. 13. Note that the flow of the identification process shown in Fig. 13 is an example of a "data processing method" according to the technology of the present disclosure.

[0153] In step S400, the data processing device 12 receives user data including sounds and images collected by the two earphones 14.

[0154] In step S402, the data processing device 12 analyzes the input image and identifies an object that the user is paying attention to.

[0155] In step S404, the data processing device 12 emphasizes the sound emitted by the identified target.

[0156] In step S406, the data processing device 12 executes a process of reproducing the emphasized sound from the speaker 40.

[0157] In the third embodiment, a database containing the user's gaze patterns, voice history, user preferences, etc. may be generated and stored in the data processing device 12. By referencing such a database and inputting additional prompts to the data generation model 58 to instruct the data generation model 58 to optimize the voice to suit the user's preferences, the user can be provided with voice that meets the user's individual needs.

[0158] If necessary, the specific processing unit 290 may output the state of gaze tracking and voice emphasis to the earphone 14 by voice guidance. Alternatively, the specific processing unit 290 may notify the state of gaze tracking and voice emphasis to a mobile device other than the earphone 14 owned by the user by an app. This allows the user to understand the current settings and states.

[0159] Furthermore, the specific processing unit 290 may acquire environmental information and motion information of the user, and may input additional prompts to the data generation model 58 to instruct real-time adjustment of specific processing parameters (for example, a parameter relating to the degree to which the gain of the separated audio is increased, or a parameter relating to the degree to which the gain of audio other than the separated audio is decreased) based on the environmental information and motion information. This allows the accuracy and performance of the audio enhancement process to be continuously improved.

[0160] Data generation model 58 may also learn the user's past gaze patterns and speech history to optimize speech enhancements to suit the user's preferences.

[0161] In the third embodiment, the camera 42 has an ultra-wide-angle lens, but this is not limiting. A subject located at the center of an image of the front of the user captured by the camera 42 may be extracted as a target.

[0162] The following supplementary notes are also disclosed in relation to the above description. (Supplementary Item 1) A data processing device comprising: an input unit including a microphone, a speaker, and a camera, which inputs audio data and image data collected by two earphones worn on the user's ears; a processing unit which performs identification processing using a data generation model which generates a predetermined inference result according to the audio data and the image data; and an output unit which plays back the result of the identification processing from the speaker, wherein the processing unit identifies an object that a user wearing the earphones is paying attention to by analyzing the image data, and performs the identification processing by emphasizing audio from the object in the audio data. (Supplementary Item 2) The data processing device according to Supplementary Item 1, wherein the processing unit acquires environmental information and motion information of the user, and uses the data generation model to adjust parameters of the identification processing in real time based on the environmental information and the motion information. (Supplementary Item 3) The data processing device according to Supplementary Item 1 or 2, wherein the processing unit uses the data generation model to optimize audio to suit the user's preferences based on the user's past gaze pattern and audio history. (Supplementary Item 4) A data processing method in which a computer executes a specific process using a data generation model that inputs audio data and image data collected by two earphones that include a microphone, a speaker, and a camera and are worn on a user's ears, and generates a predetermined inference result according to the audio data and the image data, the data processing method comprising: inputting audio data detected by the microphone and image data captured by the camera; identifying an object that a user wearing the earphones is paying attention to by analyzing the image data; executing a process to emphasize the sound from the object in the audio data as the specific process; and playing back the result of the specific process from the speaker.(Supplementary Item 5) A data processing program that includes a microphone, a speaker, and a camera, and causes a computer to execute a specific process using a data generation model that inputs voice data and image data collected by two earphones worn on a user's ears and generates a predetermined inference result according to the voice data and the image data, the data processing program causing a computer to execute the following steps: inputting voice data detected by the microphone and image data captured by the camera; identifying an object that a user wearing the earphones is paying attention to by analyzing the image data, and executing a process to emphasize the voice from the object in the voice data as the specific process; and playing back the results of the specific process from the speaker.

[0163] Fourth Embodiment An example of an embodiment of a data processing device, a data processing method, and a program according to the technology of the present disclosure will be described below with reference to the accompanying drawings. Note that the same parts as those in the first embodiment are designated by the same reference numerals, and detailed description thereof will be omitted.

[0164] FIG. 14 shows an example of the configuration of a data processing system 10 according to a fourth embodiment. In the first embodiment, the data processing system 10 is applied to single audio playback, which plays back one audio signal. However, single audio playback has the problem that the user can only perform one task at a time because the same audio signal is provided from both the left and right earphones. To solve this problem, the fourth embodiment describes the application of the data processing system 10 to dual audio playback, which plays back two different audio signals from two earphones.

[0165] The data processing system 10 according to the fourth embodiment includes a data processing device 12, a left earphone 14L, and a right earphone 14R. The earphones 14L and 14R are connected to the data processing device 12 via a network 54 so as to be able to communicate with each other.

[0166] In the fourth embodiment, two different sounds are reproduced from the left earphone 14L and the right earphone 14R. Therefore, it is important to avoid mixing of the left and right sounds. Therefore, the earphones 14L and 14R are preferably in-ear type earphones rather than bone conduction type or open-ear type earphones, and canal type earphones that have a tight fit are particularly preferable.

[0167] The left earphone 14L and the right earphone 14R may each have the same configuration as the earphone 14 according to the first embodiment. That is, the left earphone 14L may include a computer 36L, a microphone 38L, a speaker 40L, a camera 42L, and a communication I / F 44L. The computer 36L may also include a processor 46L, a RAM 48L, and a storage 50L. In addition, the left earphone 14L may also include a sensor 45L. These components may be connected to a bus 52L and be able to communicate with each other.

[0168] Similarly, the right earphone 14R may include a computer 36R, a microphone 38R, a speaker 40R, a camera 42R, and a communication I / F 44R. The computer 36R may also include a processor 46R, a RAM 48R, and a storage 50R. In addition, the right earphone 14R may also include a sensor 45R. These components may be connected to a bus 52R and be able to communicate with each other.

[0169] The sensors 45L and 45R (collectively referred to as "sensors 45") measure biological data of the user wearing the earphones 14L and 14R. Examples of the biological data may include body temperature and brain waves.

[0170] In such a data processing system 10, the data processing device 12 may be configured to be able to transmit and receive different data streams to and from the earphone 14L and the earphone 14R. In this case, the data processing device 12 may transmit different audio data in the left and right data streams while employing a left-right independent transmission method between the earphone 14L and the earphone 14R.

[0171] This allows the earphone 14L and the earphone 14R to play different sounds. Note that playing different sounds here means that the content itself that is the source of the playback plays different sounds, and may be interpreted as having a completely different meaning from playing the same content at different positions, such as in stereo playback.

[0172] 15 shows an example of an operational flow of a specific process performed by the data processing device 12 according to the fourth embodiment. This flow may be started by the user pressing a button to start dual audio playback or by speaking a command to that effect.

[0173] In step S401, the processor 28 collects user data. For example, the processor 28 may collect user data output from at least one of the microphone 38L, the camera 42L, the sensor 45L, the microphone 38R, the camera 42R, and the sensor 45R. At this time, the processor 28 may collect at least the user's biometric data as the user data. In this way, for example, the processor 28 may collect user data output from at least one of the microphone, the camera, and the sensor provided in at least one of two earphones worn on the user's left and right ears, respectively.

[0174] In step S402L, processor 28 generates a prompt for the left ear, and in step S402R, processor 28 generates a prompt for the right ear.

[0175] As an example, the processor 28 detects an utterance instructing the playback of two pieces of content from the output of at least one of the microphones 38L and 38R collected in step S401. For example, it is assumed that the processor 28 detects that the user uttered, "Play music and my schedule simultaneously?" The processor 28 also detects the user's behavior from the output of at least one of the cameras 42L and 42R collected in step S401. For example, it is assumed that the processor 28 detects that the user is walking along the seashore. The processor 28 also detects the user's state from the output of at least one of the sensors 45L and 45R collected in step S401. For example, it is assumed that the processor 28 detects that the user's stress level is high.

[0176] In this case, the processor 28 may generate a prompt for the left ear, such as "Please access the calendar app and read out today's schedule," based on the utterance instructing the playback of the two pieces of content. The processor 28 may also generate a prompt for the right ear, such as "Please play soothing ocean-themed music," based on the utterance instructing the playback of the two pieces of content, the user's actions, and the user's state. In this manner, the processor 28 may generate two prompts based on user data, for example.

[0177] In step S403L, the processor 28 acquires sound data for the left ear, and in step S403R, the processor 28 acquires sound data for the right ear.

[0178] For example, the processor 28 may input the prompt for the left ear generated in step S402L to the data generation model 58. In response to this, the processor 28 may acquire, from the data generation model 58, speech data for the left ear, such as "Today's schedule consists of three items: a meeting in the first conference room from 10:00 to 11:00, lunch with Muhammad from 12:00 to 13:00, and submitting a report to Yang by 17:00."

[0179] The processor 28 may also input the prompt for the right ear generated in step S402R to the data generation model 58. In response to this, the processor 28 may obtain music data of soothing ocean-themed music from the data generation model 58 as audio data for the right ear. As a result, the processor 28 may obtain two pieces of audio data using the two prompts. In this way, for example, the processor 28 can obtain two pieces of audio data with different content based on the user data.

[0180] In step S404L, the processor 28 transmits the audio data for the left ear to the left earphone 14L, and in step S404R, the processor 28 transmits the audio data for the right ear to the right earphone 14R.

[0181] For example, the processor 28 may transmit the audio data for the left ear acquired in step S403L to the communication I / F 44L via the communication I / F 26 and the network 54. The processor 28 may also transmit the audio data for the right ear acquired in step S403R to the communication I / F 44R via the communication I / F 26 and the network 54. In this manner, for example, the processor 28 may transmit one of the two pieces of audio data to one of the two earphones and transmit the other of the two pieces of audio data to the other of the two earphones.

[0182] As a result, the speaker 40L in the left earphone 14L plays a voice saying, "Today's schedule consists of three items: a meeting in Conference Room 1 from 10:00 to 11:00, lunch with Muhammad from 12:00 to 13:00, and submitting a report to Yang by 17:00." Also, the speaker 40R in the right earphone 14R plays soothing ocean-themed music data.

[0183] Here, the audio data played back by the right earphone 14R is for an entertainment-related task, while the audio data played back by the left earphone 14L is for a business-related task. Therefore, the audio data played back by the left earphone 14L can be considered to be more important than the audio data played back by the right earphone 14R. Therefore, the processor 28 then adjusts the balance between the left and right volumes.

[0184] In step S405L, the processor 28 issues a command to adjust the volume for the left ear, and in step S405R, the processor 28 issues a command to adjust the volume for the right ear.

[0185] For example, the processor 28 issues a command to increase the volume for the left ear so that the volume of the audio data played in the left earphone 14L is louder than the volume of the audio data played in the right earphone 14R. The processor 28 also issues a command to decrease the volume for the right ear so that the volume of the audio data played in the right earphone 14R is quieter than the volume of the audio data played in the left earphone 14L. This optimizes the volume for the left and right earphones. In this way, the processor 28 may issue a command to adjust the volume of at least one of the two earphones in accordance with the two pieces of audio data.

[0186] Then, processor 28 ends this flow. Note that the processes of steps S402L to S405L and steps S402R to S405R are executed independently of each other. Therefore, a dual AI assistant system may be provided in which separate AI assistants are simultaneously activated by the left and right earphones.

[0187] Although the above description has been given with an example of music and schedule reading, the technology of the present disclosure can also be applied to various combinations, such as music and email reading, news reading and schedule reading, and music and navigation guidance.

[0188] In this way, the data processing device 12 according to the fourth embodiment collects user data output from at least one of the microphone, camera, and sensor provided in at least one of the two earphones worn on the user's left and right ears, acquires two pieces of audio data having different contents based on the user data, and transmits one of the two pieces of audio data to one of the two earphones and the other of the two pieces of audio data to the other of the two earphones. Thus, according to the data processing device 12 according to the second embodiment, audio that allows the user to perform two tasks at once can be provided from the left and right earphones, thereby contributing to effective use of time.

[0189] In this case, the data processing device 12 may generate two prompts based on the user data and obtain two pieces of voice data using the two prompts, respectively, thereby allowing the data processing device 12 to operate the AI ​​assistants in parallel.

[0190] The data processing device 12 may also issue a command to adjust the volume of at least one of the two earphones in accordance with the two pieces of audio data, thereby enabling the data processing device 12 to optimize the left and right volumes in accordance with the importance of the content.

[0191] Furthermore, the data processing device 12 may collect at least the user's biometric data as the user data, thereby making it possible to acquire and provide voice data that also takes the biometric data into consideration using only the output from the earphones, without providing a separate means for collecting the biometric data.

[0192] The technology according to the fourth embodiment can be modified or applied in various ways. It is generally known that the left brain of the human brain is responsible for language, and the right brain is responsible for images. The structure of the nervous system from the human ear to the brain is crossed, with information received from the left ear traveling to the right brain, and information received from the right ear traveling to the left brain.

[0193] Therefore, when generating prompts for the left ear and prompts for the right ear, processor 28 may take such characteristics of the human body into consideration. For example, a prompt to play music may be generated as a prompt for the left ear to transmit music from the left ear to the right brain. Also, a prompt to read out a schedule or email may be generated as a prompt for the right ear to transmit text from the right ear to the left brain.

[0194] Fifth Embodiment An example of an embodiment of a data processing device, a data processing method, and a program according to the technology of the present disclosure will be described below with reference to the accompanying drawings. Note that the same parts as those in the first embodiment are designated by the same reference numerals, and detailed description thereof will be omitted.

[0195] As shown in FIG. 16, the earphone 14 according to this embodiment differs from the first embodiment in that it includes a sensor group 39 and a vibration applying unit 41 .

[0196] The sensor group 39 includes an environmental sensor that detects environmental data such as temperature, humidity, and air pressure, a biometric sensor that detects the user's biometric data, a motion sensor (e.g., an acceleration sensor, a gyroscope) that detects motion data such as the user's movements and posture, and a GPS sensor that detects location information.

[0197] The speaker 40 outputs sound in accordance with instructions from the processor 46. Hereinafter, the microphone 38 may be simply referred to as the mic 38.

[0198] The vibration applying unit 41 applies vibration to the user 20. For example, the vibration applying unit 41 is configured using an actuator.

[0199] As shown in FIG. 4, the specific processing section 290 of the data processing device 12 includes an input section 291, a processing section 292, and an output section 293.

[0200] The input unit 291 sequentially acquires user data including sounds picked up by the microphone 38 included in the earphone 14, images taken by the camera 42, and environmental data, biometric data, movement data, and location information detected by the sensor group 39.

[0201] Specifically, the input unit 291 acquires the user's voice picked up by the microphone 38 included in the earphone 14. The input unit 291 acquires images captured by the camera 42 included in the earphone 14. The input unit 291 acquires environmental data detected by the environmental sensor of the sensor group 39 included in the earphone 14, biological data detected by the biological sensor, motion data detected by the motion sensor, and location information detected by the GPS sensor.

[0202] The processing unit 292 generates a story according to the acquired user data, and also inputs a prompt to the data generation model 58 instructing it to generate vibration timing to vibrate in accordance with the playback of the story, thereby performing the process of generating the story and vibration timing as a specific process.

[0203] Specifically, the processing unit 292 generates the continuation of the story according to the user data most recently acquired by the input unit 291, and inputs a prompt to the data generation model 58 instructing it to generate vibration timings to vibrate in accordance with the playback of the story, thereby repeatedly generating the continuation of the story and vibration timings, and generates a series of stories and vibration timings corresponding to the series of stories.

[0204] In addition, the processing unit 292 may generate a story based on the user data and life log, and may also generate the story and vibration timing by inputting a prompt to the data generation model 58 that instructs the data generation model 58 to generate a vibration timing for vibrating in accordance with the playback of the story.

[0205] As a result of the specific processing, the output unit 293 transmits audio data reading the generated story and the generated vibration timing to the earphone 14. In the earphone 14, the control unit 46A causes the speaker 40 to play the audio data reading the generated story, and controls the vibration applying unit 41 to apply vibration to the user in accordance with the vibration timing. Next, the operation of the data processing system 10 will be described.

[0206] An example of the flow of the identification process will be described with reference to Fig. 17. Note that the flow of the identification process shown in Fig. 17 is an example of a "data processing method" according to the technology of the present disclosure.

[0207] In step S401, the input unit 291 inputs user data including sounds, images, environmental data, biological data, movement data, and position information collected by the two earphones 14.

[0208] In step S402, the processing unit 292 generates a story according to the acquired user data, and also generates a prompt instructing to generate vibration timing to vibrate in accordance with the reproduction of the story. Specifically, it generates a prompt stating, "The input data are sounds picked up around the user, images representing the user's surroundings, user environmental data, user biometric data, user movement data, and user position information. Please generate a story that matches this user data. Also, please generate vibration timing to vibrate in accordance with the reproduction of the generated story."

[0209] In step S403, the processing unit 292 inputs the acquired user data and the prompt generated in step S402 into the data generation model 58, thereby generating a story and vibration timing.

[0210] In step S404, the output unit 293 transmits the audio data reading the generated story and the generated vibration timing to the earphone 14. As a result, in the earphone 14, the control unit 46A causes the speaker 40 to play the audio data reading the generated story, and controls the vibration applying unit 41 to apply vibration to the user in accordance with the vibration timing.

[0211] In step S405, the input unit 291 inputs user data including sounds, images, environmental data, biological data, movement data, and location information collected by the two earphones 14.

[0212] In step S406, the processing unit 292 generates a story according to the acquired user data, and also generates a prompt instructing to generate vibration timing to vibrate in accordance with the reproduction of the story. Specifically, it generates a prompt stating, "The input data are sounds picked up around the user, images representing the user's surroundings, user environmental data, user biometric data, user movement data, and user position information. Please generate a continuation of the story that matches this user data. Also, please generate vibration timing to vibrate in accordance with the reproduction of the generated story."

[0213] In step S407, the processing unit 292 inputs the acquired user data and the prompt generated in step S406 into the data generation model 58, thereby generating the continuation of the story and vibration timing.

[0214] In step S408, the output unit 293 transmits the generated audio data for reading out the continuation of the story and the generated vibration timing to the earphone 14. As a result, in the earphone 14, the control unit 46A causes the speaker 40 to play the audio data for reading out the continuation of the story, and controls the vibration applying unit 41 to apply vibration to the user in accordance with the vibration timing.

[0215] In step S409, the processing unit 292 determines whether or not to end the process. For example, if the story has ended or if the user has input an instruction to end the generation of the story, the processing unit 292 determines to end the process and ends the specific process. On the other hand, if it determines not to end the process, the process returns to step S405.

[0216] Through the above processing, a story that responds to changes in the user's surrounding environment can be generated in real time and provided to the user, thereby providing the user with a new entertainment experience.

[0217] For example, if a user is walking through a forest, the camera recognizes the image of trees and determines that the scene is a "forest scene" through environmental analysis. Motion analysis detects that the user is walking. The processing unit generates a fantasy adventure story based on this information and plays the story with the user as the protagonist through audio. When encountering an enemy, a vibration function is used to create a sense of tension, and the story branches depending on the user's interactions.

[0218] In addition, the following supplementary notes are provided in relation to the above description.

[0219] (Supplementary Note 1) A data processing device comprising: an input unit that acquires user data including sounds picked up by microphones and images taken by cameras included in two earphones worn on the left and right ears of a user, each earphone including a microphone, a speaker, and a camera; a processing unit that generates a story by inputting a prompt that instructs the data generation model to generate a story according to the user data into a data generation model; and an output unit that plays the generated story from the speakers, wherein the input unit sequentially acquires the user data from the two earphones, and the processing unit repeatedly generates continuations of the story by inputting a prompt that instructs the data generation model to generate a continuation of the story according to the user data most recently acquired by the input unit, thereby generating a series of stories.

[0220] (Supplementary Note 2) The data processing device according to Supplementary Note 1, wherein the earphone further includes a sensor that detects biometric data or movement data of the user, and the input unit acquires the user data that further includes the biometric data or the movement data detected by the sensor.

[0221] (Supplementary Note 3) The data processing device according to Supplementary Note 1 or 2, wherein the earphones further include a vibration imparting unit that imparts vibration to the user, the processing unit generates a story according to the user data, and generates the story and vibration timing by inputting a prompt to the data generation model that instructs the data generation model to generate a vibration timing for vibrating in accordance with the playback of the story, and the output unit plays the generated story from the speaker, and causes the vibration imparting unit to impart vibration in accordance with the generated vibration timing.

[0222] (Appendix 4) The data processing device according to any one of Appendices 1 to 3, wherein the processing unit further acquires the user's life log in which the sounds and images associated with the user are recorded, and generates the story by inputting a prompt to the data generation model that instructs the data generation model to generate a story based on the user data and the life log.

[0223] (Supplementary Note 5) A data processing method executed by a computer, comprising: acquiring user data including sounds picked up by a microphone and images taken by a camera contained in two earphones worn on the user's left and right ears, each earphone including a microphone, a speaker, and a camera; generating the story by inputting a prompt to a data generation model instructing the model to generate a story according to the user data; and playing the generated story from the speakers; wherein acquiring the user data includes sequentially acquiring the user data from the two earphones; and generating the story includes inputting a prompt to the data generation model instructing the model to generate a continuation of the story according to the user data acquired immediately before, thereby repeatedly generating continuations of the story, thereby generating a series of stories.

[0224] (Supplementary Note 6) A data processing program that causes a computer to execute the steps of: acquiring user data including sounds picked up by a microphone and images taken by a camera contained in two earphones that are worn on the user's left and right ears, each earphone including a microphone, a speaker, and a camera; generating a story by inputting a prompt that instructs the data generation model to generate a story according to the user data; and playing the generated story from the speakers; wherein acquiring the user data includes sequentially acquiring the user data from the two earphones; and generating the story includes inputting a prompt that instructs the data generation model to generate a continuation of the story according to the user data acquired immediately before, thereby repeatedly generating continuations of the story, thereby generating a series of stories.

[0225] Sixth Embodiment An example of an embodiment of a data processing device, a data processing method, and a program according to the technology of the present disclosure will be described below with reference to the accompanying drawings. Note that the same parts as those in the first embodiment are designated by the same reference numerals, and detailed description thereof will be omitted.

[0226] In this embodiment, an example of an embodiment in which the data processing system 10 aims to enable the user 20 of the earphone 14 to efficiently study using study content will be described.

[0227] Previously, learning using earphones while on the move or in environments where visual information was unavailable was limited to one-way audio information provision, which posed challenges in terms of learning efficiency.In addition, there were few ways to emphasize important points, which sometimes led to insufficient comprehension and retention of the learning content.

[0228] The data processing system 10 according to the present embodiment aims to solve these problems and enable the user 20 to study efficiently using study content. In this embodiment, the study content includes information indicating audio of lectures at the school attended by the user 20 (corresponding to "study content information" described below), but is not limited to this form. For example, the study content may be an electronic book for study, or an electronic version of the contents of textbooks or teaching materials used in school lectures.

[0229] FIG. 18 shows an example of the configuration of a data processing system 10 according to this embodiment. As shown in FIG. 18, the data processing system 10 according to this embodiment differs from the data processing system 10 according to the first embodiment in that a sensor group 41A and a vibration generating unit 41B are added to the earphones 14. Note that in this embodiment, the sensor group 41A and the vibration generating unit 41B are provided in both of the two earphones 14 for the left and right ears, but this is not limited to this. For example, the sensor group 41A and the vibration generating unit 41B may be provided in only one of the two earphones 14.

[0230] The sensor group 41A in this embodiment includes a sensor that detects biometric data of a user 20 wearing earphones 14 in their ears, a sensor that detects movement data of the user 20, and a sensor that detects environmental data around the user 20.

[0231] Sensors that detect biological data include, for example, a heart rate sensor and a blood oxygen sensor. Sensors that detect movement data include, for example, an acceleration sensor and an angular acceleration sensor. Sensors that detect environmental data include, for example, a temperature sensor and a humidity sensor.

[0232] The earphones 14 according to the present embodiment are also equipped with a noise canceling function. The noise canceling function according to the present embodiment collects external sounds using the microphone 38 of the earphones 14, generates a sound that is out of phase with the collected sound using an internal digital circuit (not shown), and plays the generated sound together with the sound to be played back through the speaker 40 of the earphones 14. This significantly reduces ambient sounds, allowing users to hear almost only the sound to be played back. The noise canceling function according to the present embodiment is also adjustable in its intensity, i.e., the amount of ambient sound reduction. While the present embodiment is described as including noise canceling functions in both of the two earphones 14, this is not a limitation. For example, the noise canceling function may be implemented in only one of the two earphones 14.

[0233] Meanwhile, in the data processing system 10 according to this embodiment, a learning database 59, which is a database for learning for the users 20 of the data processing system 10, is constructed in the storage 32 of the data processing device 12. In the learning database 59 according to this embodiment, various learning contents to be used by the users 20 for learning, such as content related to history (hereinafter referred to as "history content"), content related to language listening (hereinafter referred to as "language content"), and content related to science (hereinafter referred to as "science content"), are registered for each user 20.

[0234] The study content according to this embodiment includes study content information that allows the target study content to be played back audibly. In this embodiment, audio information that allows the target study content to be played back audibly is used as the study content information, but this is not limited to this. For example, text information that allows the target study content to be displayed may also be used as the study content information.

[0235] In this way, in the data processing system 10 according to the present embodiment, the study content is registered in the data processing device 12, but this is not a limitation. For example, the study content may be acquired by downloading it from an external device connected to the network 54.

[0236] Furthermore, in the learning database 59 according to this embodiment, as information contained in the learning content, sound effect information is registered that can be played back when carrying out the learning content indicated by the learning content information, to produce sound effects that can improve the efficiency of the learning.

[0237] For example, sound effect information capable of reproducing sound effects such as music evoking the historical background of the time, the sounds of battle, and the hustle and bustle of the city is registered as sound effects that can improve learning efficiency when studying history content. Furthermore, sound effect information capable of reproducing environmental sounds and music from the region where the language is spoken is registered as sound effects that can improve learning efficiency when studying language content. Furthermore, sound effect information capable of reproducing sounds of space and sound effects of natural phenomena is registered as sound effects that can improve learning efficiency when studying science content.

[0238] The input unit 291 according to the present embodiment acquires user data. Specifically, the input unit 291 acquires, as user data, learning content including learning content information that allows the content of learning performed by the user 20 to be reproduced by audio.

[0239] Furthermore, the processing unit 292 according to this embodiment performs a specific process using the data generation model 58 that generates a predetermined inference result according to user data. Specifically, the processing unit 292 performs the specific process of generating vibration information by inputting a prompt to the data generation model 58, based on the user data, instructing the data generation model 58 to generate vibration information that vibrates the vibration generating unit 41B in synchronization with the timing of audio playback of an important part of the learning content indicated by the learning content information.

[0240] For example, along with the study content used by user 20 for study, a prompt such as "This is the study content used by the user for study. Please detect key points for study, such as keywords and sections, from the study content of this study content, and generate vibration information that vibrates the earphones in synchronization with the timing of audio playback of the detected points" is input to data generation model 58. As a result, vibration information is generated by data generation model 58.

[0241] Then, the output unit 293 according to the present embodiment uses the result of the identification process to play sounds from the speakers 40 of the two earphones 14. Specifically, the output unit 293 plays back the learning content indicated by the acquired learning content information from the speakers 40 of the earphones 14, and vibrates the vibration generating unit 41B using the generated vibration information.

[0242] In the present embodiment, the vibration information is instruction information that instructs the earphones 14 to vibrate the vibration generating unit 41B in synchronization with the timing of playback of the important parts when the content of the learning is being played back as audio through the speakers 40 of the earphones 14. The output unit 293 then transmits the generated instruction information to the earphones 14. Upon receiving the instruction information, the processor 46 of the earphones 14 according to the present embodiment vibrates the vibration generating unit 41B in accordance with the instruction information. However, this is not limited to this embodiment, and an embodiment may also be adopted in which the data generation model 58 generates information for vibrating the vibration generating unit 41B as the vibration information, and the data processing device 12 directly vibrates the vibration generating unit 41B of the earphones 14.

[0243] Here, the output unit 293 according to the present embodiment plays the learning content from the speaker 40 of one earphone 14, and plays sound effects related to the learning content from the speaker 40 of the other earphone 14. In the present embodiment, a case will be described in which the learning content is played from the speaker 40 of the earphone 14 worn on the right ear of the user 20, and the sound effects are played from the speaker 40 of the earphone 14 worn on the left ear of the user 20, but this is not limiting. Alternatively, the learning content may be played from the speaker 40 of the earphone 14 worn on the left ear of the user 20, and the sound effects may be played from the speaker 40 of the earphone 14 worn on the right ear of the user 20.

[0244] Furthermore, the input unit 291 according to this embodiment further acquires the above-described biometric data, movement data, and environmental data as user data from the sensor group 41A. Therefore, the processing unit 292 according to this embodiment inputs the learning content, biometric data, movement data, and environmental data acquired by the input unit 291 into the data generation model 58, thereby executing a process of generating vibration information as a specific process.

[0245] More specifically, the processing unit 292 according to this embodiment inputs a prompt including the learning content, biometric data, movement data, and environmental data acquired by the input unit 291 into the data generation model 58, which prompt instructs the generation of vibration information, and acquires the generated result, i.e., the vibration information.

[0246] For example, along with the learning content used by user 20 in learning, biometric data, movement data, and environmental data, a prompt such as "This is the learning content the user will learn, biometric data indicating the user's heart rate and blood oxygen level, movement data indicating the user's acceleration and angular acceleration, and environmental data indicating the user's surrounding environmental conditions. From the learning content of this learning content, please detect key points for learning, such as keywords and sections, and generate vibration information that vibrates the earphones in synchronization with the timing of audio playback of the detected points, according to the user's situation," is input to data generation model 58. As a result, data generation model 58 acquires vibration information that corresponds to the situation of user 20.

[0247] In this way, the processing unit 292 according to this embodiment uses, in addition to the learning content, biometric data, movement data, and environmental data related to the user 20, all of which are acquired by the sensor group 41A, but is not limited to this. For example, the processing unit 292 may use only the learning content, or may use the learning content in combination with one or two of the biometric data, movement data, and environmental data.

[0248] Furthermore, the input unit 291 according to this embodiment further acquires sounds around the user 20 from the microphone 38 as user data, and the processing unit 292 according to this embodiment performs a specific process of adjusting the intensity of the output target by the output unit 293, specifically, the playback volume of the voice and sound effects indicating the content of the learning, and the strength of the vibration applied to the vibration generating unit 41B, depending on the acquired sounds around the user 20.

[0249] Furthermore, the processing unit 292 according to this embodiment performs, as a specific process, a process of adjusting the intensity of the noise canceling function in accordance with the sounds around the user 20 acquired from the microphone 38. Specifically, the louder the sounds around the user 20, the greater the amount of noise reduction by the noise canceling function. This can further improve the efficiency with which the user 20 learns using the learning content.

[0250] Next, the operation of the data processing system 10 according to this embodiment will be described.

[0251] An example of the flow of the identification process will be described with reference to Figure 19. The flow of the identification process shown in Figure 19 is an example of a "data processing method" according to the technology of the present disclosure. In this example, to avoid confusion, the case where the learning content (hereinafter simply referred to as "learning content") to be studied by the user 20 (hereinafter simply referred to as "user") who is the target of the process is known in advance will be described.

[0252] In step S400, data processing device 12 acquires the study content by reading it from study database 59.

[0253] In step S402, the data processing device 12 receives and acquires the above-mentioned biological data, movement data, and environmental data from the earphones 14 worn by the user.

[0254] In step S404, the data processing device 12 generates the above-mentioned prompt using the acquired learning content, biometric data, movement data, and environmental data.

[0255] In step S406, the data processing device 12 inputs the generated prompt into the data generation model 58, thereby causing the data generation model 58 to generate vibration information.

[0256] In step S408, data processing device 12 starts playing audio indicating the learning content indicated by the learning content information included in the acquired learning content from speaker 40 of earphone 14 worn by the user on his right ear. Data processing device 12 also starts playing sound effects indicated by the sound effect information included in the acquired learning content from speaker 40 of earphone 14 worn by the user on his left ear. Furthermore, data processing device 12 transmits the generated vibration information to both earphones 14 worn by the user, thereby starting vibration of earphone 14 synchronized with the playback of important content in the learning content, in parallel with the playback of the learning content and sound effects.

[0257] In step S410, the data processing device 12 receives and acquires audio data representing sounds around the user from the earphones 14 worn by the user.

[0258] In step S412, the data processing device 12 determines whether the sounds around the user indicated by the acquired audio data require adjustment of the learning content and the playback volume of the sound effects, and if the determination is negative, the process proceeds to step S416, whereas if the determination is positive, the process proceeds to step S414.

[0259] In step S414, the data processing device 12 adjusts the volume of the voice and sound effects indicating the learning content according to the volume of the sounds around the user indicated by the acquired audio data, and also adjusts the strength of the vibration generated by the vibration generating unit 41B.

[0260] In step S416, the data processing device 12 determines whether the sounds around the user indicated by the acquired audio data require adjustment of the intensity of the noise canceling function, and if the determination is negative, proceeds to step S420, whereas if the determination is positive, proceeds to step S418.

[0261] In step S418, the data processing device 12 adjusts the intensity of the noise canceling function in accordance with the volume of the sound around the user indicated by the acquired audio data, and then proceeds to step S420.

[0262] The audio data acquired in the process of step S410 includes two pieces of audio data, one from the earphone 14 worn on the user's right ear and one from the earphone 14 worn on the user's left ear, and each piece of audio data is time-series data. Therefore, in this embodiment, the average value of the peak volumes of the two pieces of audio data within a predetermined period (in this embodiment, the most recent one second) is applied as the sound around the user to be applied in the processes of steps S412 to S418. However, this is not limited to this form, and for example, the peak volumes of the two pieces of audio data within a predetermined period may be applied.

[0263] In step S420, the data processing device 12 determines whether the timing has come to end this specific process (in this embodiment, this is the timing when the user 20 inputs an instruction to end the specific process, or the timing when all of the learning content has been played back using the learning content), and if the determination is negative, the process returns to step S410, whereas if the determination is positive, the process ends.

[0264] Through the above-described identification process, the earphones 14 worn by the user can be used to play the learning content and sound effects, and the earphones 14 can vibrate when important parts of the learning content are played. As a result, learning can be achieved by combining hearing and touch, allowing for efficient learning using the learning content even in situations where visual information is unavailable. Furthermore, the associated sound effects increase interest and attention in learning, thereby helping to maintain concentration. Furthermore, the vibration of the earphones 14 emphasizes important points, allowing for reliable recognition of important parts. For example, effective learning is possible even when studying on the move, such as during a commute to work or school, when hands or eyes are occupied.

[0265] The following scenarios can be considered as specific examples of the data processing system 10 according to this embodiment. First scenario: When learning history Situation: User learns a history lecture while commuting. Right ear: Lecture audio about historical events is played. Left ear: Music and sound effects that evoke the historical background of the time (sounds of battle, city noise, etc.) are played. Vibration: Vibration is performed when important dates or events are played.

[0266] Scenario 2: Language learning Situation: User is listening to learn a language while taking a walk. Right ear: Plays conversational audio in a foreign language. Left ear: Plays ambient sounds or music from the area where the language is spoken. Vibration: Vibrates when new words or important phrases are played.

[0267] Scenario 3: Learning science Situation: User is learning science lectures while jogging. Right ear: Audio explaining scientific theories and experiments is played. Left ear: Sounds of the universe and sound effects of natural phenomena are played. Vibration: Vibration is activated when key concepts and formulas are explained.

[0268] Although not mentioned in this embodiment, the reproduction of sound effects may be adjusted so as not to interfere with the audio indicating the learning content.

[0269] Furthermore, the learning content and the output settings by the output unit 293 may be adjusted according to the importance of the learning content, settings made by the user 20, and operations and reactions of the user 20.

[0270] In addition, the data generation model 58 may be used to provide additional explanations or question and answer sessions in real time depending on the user's 20 level of understanding, and the difficulty level of the learning content and the vibration pattern of the earphones 14 may be dynamically adjusted based on the user's learning history, level of understanding, progress, etc.

[0271] In addition, in this embodiment, the case where the learning content information and the sound effect information are registered in a single learning database 59 has been described, but the present invention is not limited to this, and the learning content information and the sound effect information may be registered in different databases. In this embodiment, the database for registering the learning content information and the database for registering the sound effect information may be constructed in different devices.

[0272] Alternatively, sound effect information related to the learning content indicated by the learning content information may be searched for and applied using the data generation model 58, without preparing the sound effect information in advance.

[0273] In this case, for example, along with the learning content used by user 20 for learning, a prompt such as "This is the learning content used by the user for learning. From the learning content of this learning content, detect important points for learning, such as keywords and sections, and generate vibration information that vibrates the earphones in synchronization with the timing of playing back audio of the detected points. Also, search for sound effect information that can play sound effects that will allow learning using this learning content to be efficient, and generate the results" is input to data generation model 58.

[0274] Furthermore, in this case, when biometric data, movement data, and environmental data are used, for example, the following prompt is input to data generation model 58 along with the learning content used by user 20 in learning, the biometric data, movement data, and environmental data: "This is the learning content the user will learn, biometric data indicating the user's heart rate and blood oxygen level, movement data indicating the user's acceleration and angular acceleration, and environmental data indicating the user's surrounding environmental conditions. From the learning content of this learning content, detect important learning points such as keywords and sections, and generate vibration information that vibrates the earphones in synchronization with the timing of playing back the detected points as audio. Also, taking into account this biometric data, movement data, and environmental data, search for sound effect information that indicates sound effects that are effective in allowing the user to learn the learning indicated by the learning content in a state optimal for the situation in which they are placed, and generate the results."

[0275] Furthermore, the heart rate, blood oxygen level, etc. obtained by a biosensor provided in the earphone 14 may be monitored to estimate the concentration level and fatigue level of the user 20 and encourage them to take a break at an appropriate time.

[0276] In addition, the following supplementary notes are provided in relation to the above description.

[0277] <Supplementary Note 1> A data processing device comprising: an input unit that acquires user data; a processing unit that performs a specific processing using a data generation model that generates a predetermined inference result according to the user data; and an output unit that uses a result of the specific processing to play audio from speakers of two earphones, at least one of which is provided with a vibration generating unit and which are worn on the left and right ears of the user, wherein the input unit acquires, as the user data, learning content including learning content information that can play audio of learning content performed by the user; the processing unit performs a process of generating the vibration information as the specific processing by inputting, to the data generation model, a prompt based on the user data, that instructs the data generation model to generate vibration information that vibrates the vibration generating unit in synchronization with timing of playing audio of important learning points in the learning content indicated by the learning content information; and the output unit plays the learning content indicated by the learning content information from the speakers of the earphones and vibrates the vibration generating unit using the vibration information. <Supplementary Note 2> The data processing device of Supplementary Note 1, wherein the study content further includes sound effect information capable of playing sound effects related to the study content, and the output unit plays the study content indicated by the study content information from the speaker of one of the earphones and plays sound effects indicated by the sound effect information from the speaker of the other of the earphones. <Supplementary Note 3> The data processing device of Supplementary Note 1 or Supplementary Note 2, wherein at least one of the two earphones includes a sensor that detects at least one of biometric data of the user, movement data of the user, and environmental data around the user, and the input unit further acquires at least one of the biometric data, the movement data, and the environmental data from the sensor as the user data. <Supplementary Note 4> The data processing device according to any one of Supplementary Note 1 to Supplementary Note 3, wherein at least one of the two earphones includes a microphone, the input unit further acquires ambient sounds of the user from the microphone as the user data, and the processing unit further performs, as the specific processing, a process of adjusting an intensity of an output target by the output unit in accordance with the ambient sounds.<Supplementary Note 5> The data processing device according to Supplementary Note 4, wherein at least one of the two earphones further has a noise canceling function, and the processing unit further performs, as the specific processing, a process of adjusting an intensity of the noise canceling function in accordance with the ambient sound. <Supplementary Note 6> A data processing method for acquiring user data, performing a specification process using a data generation model that generates a predetermined inference result according to the user data, and using the result of the specification process to reproduce audio from speakers of two earphones, at least one of which is provided with a vibration generating unit and which are worn on each of the user's ears, wherein the computer acquires, as the user data, learning content including learning content information that can reproduce audio of learning content performed by the user, inputs, based on the user data, a prompt to the data generation model that instructs the computer to generate vibration information that vibrates the vibration generating unit in synchronization with the timing of audio reproduction of important learning points in the learning content indicated by the learning content information, thereby performing the specification process to generate the vibration information, and reproduces the learning content indicated by the learning content information from the speakers of the earphones, and vibrates the vibration generating unit using the vibration information.<Supplementary Note 7> A data processing program that causes a computer to execute the following processes: acquire user data; perform a specific process using a data generation model that generates a predetermined inference result according to the user data; and use the result of the specific process to play audio from speakers of two earphones, at least one of which is provided with a vibration generating unit and which are worn on each of the user's ears, one each, wherein learning content performed by the user is acquired as the user data, including learning content information that allows learning content indicated by the learning content information to be played back by audio; perform the process of generating the vibration information as the specific process by inputting into the data generation model a prompt that instructs the data generation model to generate vibration information that vibrates the vibration generating unit in synchronization with the timing of playing back by audio important learning points in the learning content indicated by the learning content information; and play back the learning content indicated by the learning content information from the speakers of the earphones, and vibrate the vibration generating unit using the vibration information.

[0278] Seventh Embodiment An example of an embodiment of a data processing device, a data processing method, and a program according to the technology of the present disclosure will be described below with reference to the accompanying drawings. Note that the same parts as those in the first embodiment are designated by the same reference numerals, and detailed description thereof will be omitted.

[0279] As shown in FIG. 20, the earphone 14 according to this embodiment differs from the first embodiment in that it includes a biological information sensor 39 and a vibration applying unit 41 .

[0280] The biological information sensor 39 is a sensor for detecting biological information of the user, and includes a heart rate sensor for detecting a heart rate (HR: Heart Rate), a skin electrodermal response sensor for detecting a galvanic skin response (GSR: Galvanic Skin Response), and a blood oxygen saturation concentration sensor (SpO 2 The biological information sensor 39 is disposed inside the earphone 14 or in a portion that comes into contact with the skin of the user 20. The biological information sensor 39 also performs noise removal, smoothing, and the like on the detected biological information to generate data suitable for analysis.

[0281] The vibration applying unit 41 applies vibration to the user 20. For example, the vibration applying unit 41 is configured using an actuator.

[0282] As shown in FIG. 4, the specific processing section 290 of the data processing device 12 includes an input section 291, a processing section 292, and an output section 293.

[0283] The input unit 291 sequentially acquires user data including biometric information detected by the biometric information sensor 39 .

[0284] Specifically, the input unit 291 acquires biological information detected by the biological information sensor 39 included in the earphone 14 .

[0285] The processing unit 292 evaluates the user's condition based on the biometric information, and generates music data and vibration pattern data according to the evaluated user's condition as a result of the specific processing.

[0286] Specifically, the processing unit 292 evaluates the user's condition based on the biometric information acquired by the input unit 291, and generates music data and vibration pattern data according to the user's condition by inputting a prompt to the data generation model 58 to instruct the data generation model 58 to generate music data and vibration pattern data according to the evaluated user's condition. For example, the processing unit 292 inputs the biometric information acquired by the input unit 291 to the data generation model 58, and also inputs a prompt to the data generation model 58 saying, "This is the user's biometric information. Evaluate the user's condition based on this biometric information, and generate music data that is optimal for the evaluated user's condition, as well as generate vibration pattern data that matches the playback of the music data."

[0287] The processing unit 292 may generate music data and vibration pattern data according to the user's state by inputting a prompt to the data generation model 58 instructing the data generation model 58 to generate music data and vibration pattern data according to the user's state based on the biometric information acquired by the input unit 291, the user's life log in which past biometric information, music data, and vibration pattern data are recorded, and the user's music preference data. For example, the processing unit 292 inputs the biometric information acquired by the input unit 291, the user's life log, and the user's music preference data to the data generation model 58, and also inputs a prompt to the data generation model 58 saying, "These are the user's biometric information, the user's life log, and the user's music preference data. Please evaluate the user's state based on this biometric information, and take into consideration the user's life log and the user's music preference data, generate music data that is optimal for the evaluated user's state, and generate vibration pattern data that is synchronized with the playback of the music data." The life log and music preference data are stored in the database 24 and can be acquired by reading them from the database 24. The data generation model 58 includes a machine learning or deep learning algorithm and is a model trained to evaluate the mental and physical state of the user from biometric information.

[0288] As a result of the identification process, the output unit 293 transmits the generated music data and vibration pattern data to the earphones 14. In the earphones 14, the control unit 46A controls the speaker 40 to play the generated music data and controls the vibration applying unit 41 to apply vibrations to the user in accordance with the vibration pattern data. This allows the user 20 to achieve stress relief and relaxation effects.

[0289] The processing unit 292 may also perform an integrated analysis of the life log and music preference data to personalize the music data and vibration pattern data according to the characteristics of each individual user, thereby enabling the user 20 to have a personalized therapy experience.

[0290] The processing unit 292 may also be configured to continuously monitor biological information while providing music and vibrations, and perform feedback control to adjust the music data and vibration pattern data in real time in response to changes in the user's condition, thereby enabling the user 20 to more effectively experience the effects of stress reduction and relaxation.

[0291] Next, the operation of the data processing system 10 will be described.

[0292] An example of the flow of the identification process will be described with reference to Fig. 21. Note that the flow of the identification process shown in Fig. 21 is an example of a "data processing method" according to the technology of the present disclosure.

[0293] In step S400, the input unit 291 inputs user data including biometric information detected by the two earphones 14, a life log of the user 20 in which past biometric information, music data, and vibration pattern data are recorded, and music preference data of the user 20.

[0294] In step S401, the processing unit 292 evaluates the state of the user 20 based on the biometric information acquired by the input unit 291, the life log of the user 20 in which past biometric information, music data, and vibration pattern data are recorded, and the music preference data of the user 20, and inputs a prompt to the data generation model 58 to instruct the data generation model 58 to generate music data and vibration pattern data according to the evaluated state of the user 20. Specifically, it generates a prompt stating, "The input data is data including the user's current biometric information, a life log including past biometric information, music data, and vibration pattern data, and music preference data. Please generate music data and vibration pattern data that match these user data."

[0295] In step S402, the processing unit 292 inputs the acquired user data and the prompt generated in step S401 to the data generation model 58, thereby generating music data and vibration pattern data.

[0296] In step S403, the output unit 293 transmits the generated music data and vibration pattern data to the earphone 14. As a result, in the earphone 14, the control unit 46A controls the speaker 40 to play the generated music data and controls the vibration applying unit 41 to apply vibrations to the user in accordance with the vibration pattern data, thereby enabling the user 20 to obtain stress relief and relaxation effects.

[0297] In step S404, the processing unit 292 determines whether or not to end the process. For example, if the playback of the music data has ended or if the user has input an instruction to end the playback of the music data, the processing unit 292 determines to end the process and ends the specific process. On the other hand, if it determines not to end the process, the process returns to step S400.

[0298] Through the above processing, it is possible to generate music data and vibration pattern data according to the biometric information of the user 20. As a result, the user 20 can obtain stress relief and relaxation effects.

[0299] For example, if the user 20 is feeling stressed, an increase in heart rate and a change in skin electrical response are detected as biological information by the biological information sensor 39. The specific processing unit 290 evaluates the state of the user 20 as being stressed based on the biological information, and generates slow-tempo music and a gentle vibration pattern that have a high relaxation effect in accordance with the stressed state, and provides them to the user 20.

[0300] Furthermore, when the user 20 puts on the earphones 14 during a break from work, the biometric information sensor 39 detects an increase in the heart rate and GSR value. The specific processing unit 290 evaluates that the user 20 is feeling stressed, and generates slow-tempo, low-frequency music with a high relaxation effect and a gentle vibration pattern, and provides them to the user 20.

[0301] Furthermore, when the user 20 uses the earphones 14 while jogging, the biometric information sensor 39 detects an appropriate heart rate and GSR value. The specific processing unit 290 evaluates that the user 20 is exercising, and generates motivation-boosting up-tempo music and a rhythmic vibration pattern in accordance with the user's exercising state, and provides these to the user 20.

[0302] In addition, the following supplementary notes are provided in relation to the above description.

[0303] (Supplementary Note 1) A data processing device comprising: an input unit that inputs user data including biometric information collected by two earphones that include a microphone, a speaker, a camera, a biometric information sensor, and a vibration applying unit and are worn on the ears of a user; a processing unit that performs a specific processing using a data generation model that generates a predetermined inference result according to the user data; and an output unit that plays the result of the specific processing from the speaker and causes the vibration applying unit to apply vibration, wherein the input unit inputs the biometric information detected by the biometric information sensor as the user data, and the processing unit performs the specific processing by evaluating a state of the user based on the biometric information and inputting a prompt to the data generation model that instructs the data generation model to generate music data and vibration pattern data corresponding to the evaluated state of the user.

[0304] (Supplementary Note 2) The data processing device according to Supplementary Note 1, wherein the processing unit generates the music data and vibration pattern data as a result of the identification process by inputting a prompt to the data generation model, the prompt instructing the data generation model to generate music data and vibration pattern data according to a state of the user, based on the user data, a life log of the user in which the past biometric information, the music data, and the vibration pattern data are recorded, and music preference data of the user.

[0305] (Supplementary Note 3) The data processing device according to Supplementary Note 2, wherein the processing unit performs an integrated analysis of the life log and the music preference data, and personalizes the music data and the vibration pattern data in accordance with characteristics of an individual user.

[0306] (Supplementary Note 4) The data processing device according to any one of Supplementary Notes 1 to 3, wherein the processing unit continuously monitors the biological information while providing music and vibration, and performs feedback control to adjust the music data and the vibration pattern data in real time according to changes in the user's condition.

[0307] (Supplementary Note 5) The data processing device according to any one of Supplementary Notes 1 to 4, wherein the data generation model includes a machine learning or deep learning algorithm and is a model trained to evaluate the mental and physical state of the user from the biometric information.

[0308] (Supplementary Note 6) A data processing method in which a computer executes a specific processing using a data generation model that inputs user data including biometric information collected by two earphones that include a microphone, a speaker, a camera, a biometric information sensor, and a vibration applying unit and are worn on the user's ears, and generates a predetermined inference result according to the user data, wherein the specific processing is executed by inputting biometric information detected by the biometric information sensor as the user data, evaluating the user's condition based on the biometric information, and inputting a prompt to the data generation model that instructs the model to generate music data and vibration pattern data according to the evaluated user's condition, thereby generating the music data and vibration pattern data as a result of the specific processing, and playing the result of the specific processing from the speaker and causing the vibration applying unit to vibrate.

[0309] (Supplementary Note 7) A data processing program that causes a computer to execute a specific process using a data generation model that receives user data including biometric information collected by two earphones that include a microphone, a speaker, a camera, a biometric information sensor, and a vibration applying unit and are worn on the user's ears, and generates a predetermined inference result according to the user data, the data processing program comprising: inputting biometric information detected by the biometric information sensor as the user data; evaluating the user's condition based on the biometric information; and inputting a prompt to the data generation model that instructs the model to generate music data and vibration pattern data according to the evaluated user's condition, thereby executing, as the specific process, a process to generate the music data and vibration pattern data as a result of the specific process; and playing the result of the specific process from the speaker and causing the vibration applying unit to vibrate.

[0310] Eighth Embodiment An example of an embodiment of a data processing device, a data processing method, and a program according to the technology of the present disclosure will be described below with reference to the accompanying drawings. Note that the same parts as those in the first embodiment are designated by the same reference numerals, and detailed description thereof will be omitted.

[0311] The following describes embodiments 8-1 to 8-3, which allow distant parties to share emotions and feelings that cannot be conveyed through words or images alone.

[0312] [Embodiment 8-1] (Outline of Embodiment 8-1) Embodiment 8-1 describes a configuration example in which the transmitting side performs operation data, emotion determination, and vibration data generation and transmission, and the receiving side generates vibration. In embodiment 8-1, as shown in FIG. 22 , an example configuration is described in which earphones 14-2 worn by a specific receiving-side first user (user 20) input vibration data transmitted from earphones 14-2 worn by a transmitting-side second user (user 20A) other than the receiving-side first user, and vibrate the housing 14a of the earphones 14-2 worn by the receiving-side first user based on the vibration data. The vibration data may be interpreted as data that reproduces at least one of the intention and emotion of the transmitting-side second user as vibration in the earphones 14-2 worn by the receiving-side first user. Hereinafter, for simplicity of explanation, the earphones 14-2 may be simply referred to as earphones.

[0313] When the second transmitting user operates the earphones worn by the transmitting user with a finger in accordance with the strength and rhythm corresponding to the mood at the time, the message to be conveyed, etc., the earphones of the first receiving user vibrate in conjunction with the operation. For example, if the second transmitting user taps his / her earphones at a certain rhythm to inform the first receiving user that he / she will soon begin transmitting a specific message, the earphones of the first receiving user also vibrate at the same rhythm. For example, if the second transmitting user swipes his / her earphones to convey a specific message to the first receiving user (e.g., that he / she is tired or irritated), the earphones of the first receiving user also vibrate intermittently in accordance with the swipe. For example, if the second transmitting user presses and holds his / her earphones to inform the first receiving user that something urgent has happened, the earphones of the first receiving user also vibrate continuously while the press and hold is continued.

[0314] In this way, by having the earphones of the first receiving user vibrate in conjunction with the emotions, situation, etc. of the second transmitting user, emotions, feelings, etc. that cannot be conveyed through words alone can be shared between the first receiving user and the second transmitting user through vibrations, enabling more intimate interactions. Furthermore, even in situations where it is difficult to speak during a meeting or to operate the screen of a terminal device such as a smartphone, communication can be carried out between the first receiving user and the second transmitting user.

[0315] Furthermore, by conveying the emotions of the second transmitting user to the first receiving user through vibration, it is possible to communicate with the other party in real time even in situations where it is difficult to hear through voice, such as when there are many cars driving around the first receiving user or when the first receiving user is in a crowd.

[0316] Furthermore, by conveying the emotions of the second sending user to the first receiving user through vibration, the first receiving user can intuitively understand the emotions and intentions of the second sending user without losing concentration, compared to communicating by voice, even when the first receiving user is performing a specific task, such as desk work, attending a meeting, or driving a car.

[0317] Furthermore, by conveying the emotions of the second sending user to the first receiving user through vibration, even in situations where conversation is difficult, such as when the first receiving user is in a quiet space, there is no risk that people around the first receiving user will learn the emotions of the second sending user.

[0318] In particular, when the earphones are inner-ear type (open type), there is a risk that the sound will leak to the surroundings of the first receiving user. In contrast, according to the 8-1 embodiment, even if the earphones are open type, there is no risk that the emotions of the second transmitting user will be known by people around the first receiving user. The same applies to canal type earphones and inner-ear type earphones.

[0319] In the 8-1 embodiment, in addition to the emotion of the second transmitting user, information about the second transmitting user's body, such as body temperature, pulse rate, blood pressure fluctuations, etc., may be communicated by vibration. For example, the first receiving user can be notified by the vibration of the earphone that the pulse rate of the second transmitting user is increasing due to exercise.

[0320] In the 8-1 embodiment, the earphone of the first receiving user may directly communicate wirelessly with the earphone of the second transmitting user.

[0321] In the 8-1 embodiment, the case where the terminal device of the transmitting second user that transmits the operation data is an earphone is described, but the terminal device of the transmitting second user may also be a terminal device other than an earphone, such as a smartphone or a laptop computer.

[0322] 22 shows an example of the configuration of a data processing system 10-2 according to embodiment 8-1. The data processing system 10-2 may include a data processing device 12 and a plurality of earphones.

[0323] The user 20 wearing one of the earphones may be interpreted as a "first user on the receiving side" according to the technology of the present disclosure. The user 20A wearing the other earphone may be interpreted as a "second user on the transmitting side" according to the technology of the present disclosure.

[0324] Each earphone shown in FIG. 22 may include a processor 46-2, a RAM 48, a storage 50-2, a camera 42, a microphone 38, a touch sensor 39, a vibration unit 43, a speaker 41, and a communication I / F 44.

[0325] The touch sensor 39 may detect an operation on the earphones worn by the second transmitting user, such as at least one of a swipe operation, a tap operation, and a long press operation. The earphones of the second transmitting user may detect the operation using the touch sensor 39, generate operation data including information indicating the detected operation, and generate vibration data based on the generated operation data. The vibration data may then be transmitted to the earphones of the first receiving user. The processor 46-2 may perform a specification process using a data generation model that generates a predetermined estimation result according to the operation data, and the vibration data may be interpreted as data to generate vibrations according to the result of the specification process in the earphones worn in the ears of the first receiving user.

[0326] The earphone of the second transmitting user may transmit the vibration data wirelessly directly to the earphone of the first receiving user, i.e., without going through the data processing device 12 .

[0327] The earphones of the first receiving user may execute a specific process to estimate at least one of the intentions and emotions of the second transmitting user based on the received vibration data, and may generate vibrations according to the results of this specific process using a vibration unit 43 in the earphones worn by the first receiving user.

[0328] 23 shows an example of the main functions of the data processing device 12 and the earphones. In the earphones, a processor 46-2 performs data collection processing and identification processing.

[0329] The storage 50-2 stores a data collection program 60-2, a specific processing program 61-2, and a data generation model 62-2. The specific processing program 61-2 and the data generation model 62-2 are used for specific processing in the specific processing unit 46B-2.

[0330] The processor 46-2 reads the data collection program 60-2 from the storage 50-2 and executes the read data collection program 60-2 on the RAM 48. The data collection process is realized by the processor 46-2 operating as the control unit 46A-2 in accordance with the data collection program 60-2 executed on the RAM 48.

[0331] The processor 46-2 reads the specific processing program 61-2 from the storage 50-2 and executes the read specific processing program 61-2 on the RAM 48. The specific processing is realized by the processor 46-2 operating as a specific processing unit 46B-2 in accordance with the specific processing program 61-2 executed on the RAM 48.

[0332] The data generation model 62-2 is a so-called generative AI. Examples of the data generation model 62-2 include generative AI such as ChatGPT (registered trademark) (Internet search <URL: https: / / openai.com / blog / chatgpt>) and Gemini (registered trademark) (Internet search <URL: https: / / gemini.google.com / ?hl=ja>). The data generation model 62-2 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 62-2, and estimation data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 62-2 estimates the input estimation data in accordance with the instructions indicated by the prompt and outputs the estimation result in a data format such as voice data and text data. Here, estimation refers to, for example, analysis, classification, prediction, and / or summarization.

[0333] 24, the control unit 46A-2 may include a data collection unit 100-2. The specific processing unit 46B-2 may include an input unit 200-2, a processing unit 201-2, and an output unit 202-2.

[0334] The data collection unit 100-2 may collect the outputs of the microphone 38, the touch sensor 39, and the camera 42 shown in Fig. 22, and may further collect the operation data described above. As described above, the operation data may include information indicating the operation content detected by the touch sensor 39.

[0335] The operations include swiping, tapping, and long press operations on the earphones. A swipe operation may be interpreted as, for example, placing a finger on the touch sensor 39 and sliding the finger in any direction. A tap operation may be interpreted as, for example, touching the touch sensor 39 with a finger as if tapping it for a moment and then releasing it. A tap operation may include a double tap, which is a quick tap twice in succession. A long press operation may be interpreted as touching the touch sensor 39 with a finger and pressing it for a long time without releasing it.

[0336] For example, when user 20A, who is attending the same conference as user 20, comes up with a good idea and wants to let user 20 know about it, he or she performs a tap operation, for example, to generate operation data indicating that it is a tap operation, and vibration data corresponding to the operation data is sent to the earphone.

[0337] For example, in order to convey to user 20 that user 20A is in a bad mood and irritated during a meeting, user 20A performs a double tap operation, whereby operation data indicating that it is a double tap operation is generated, and vibration data corresponding to the operation data is transmitted to the earphone.

[0338] (Processing Unit 201-2) Data collected by the data collection unit 100-2 may be input to the processing unit 201-2 via the input unit 200-2. The processing unit 201-2 may perform a determination process using a data generation model 62-2 (generative AI model) that generates a predetermined estimation result according to the input operation data. The processing unit 201-2 may generate vibration data that generates vibrations in earphones worn on the ears of the first receiving user according to the results of the determination process. For example, the processing unit 201-2 may execute a determination process that estimates at least one of the intentions and emotions of the second transmitting user based on the operation data, thereby generating vibrations in the earphones worn on the ears of the first receiving user according to the results of the determination process.

[0339] Specifically, the processing unit 201-2 may use the data generation model 62-2 to estimate the emotions, intentions, etc. of the user 20A based on information contained in the operation data, such as the strength of the touch operation, the rhythm of the touch operation, and the duration of the touch operation.

[0340] For example, if the earphones worn by user 20A are tapped three times, processing unit 201-2 recognizes that a pre-defined situation has occurred in which user 20 wants user 20 to notice user 20A. In this case, processing unit 201-2 may generate vibration data indicating a specific vibration period, a specific vibration intensity, a specific number of vibrations, etc., and transmit the vibration data to output unit 202-2 in order to accurately convey user 20A's intention. Specifically, processing unit 201-2 may transmit vibration data that generates vibrations at the lowest vibration level three times at one-second intervals to output unit 202-2. This allows, for example, user 20A participating in a conference to quickly convey to other users 20 participating in the same conference room that user 20A wants to convey some kind of signal, instead of using a phone call or a short message.

[0341] For example, when user 20A performs a long press on the earphones worn by the user 20A, for example, by continuously pressing the housing 14a with a finger for five seconds, the processing unit 201-2 recognizes that a situation has occurred in which user 20A needs to convey an urgent matter or an important matter to a pre-defined user 20. In this case, the processing unit 201-2 may generate vibration data indicating a specific vibration period, a specific vibration intensity, a specific number of vibrations, etc., and transmit the vibration data to the output unit 202-2 in order to accurately convey user 20A's intention. Specifically, the processing unit 201-2 may transmit vibration data in which the highest vibration level is generated for three consecutive seconds, followed by ten consecutive vibrations separated by a one-second pause, to the output unit 202-2. This allows, for example, user 20A to quickly convey to the user 20 that he or she needs to convey an urgent matter, instead of by phone call or short message.

[0342] The processing unit 201-2 may perform the identification process in consideration of the operation history of the housing 14a worn by the second sending user on the ear of the second sending user.

[0343] While generating vibrations, the processing unit 201-2 may generate audio data to be played from earphones worn on the ears of the first receiving user, as audio guidance containing at least one of the intention and emotion of the second transmitting user, according to the result of the specific processing. This allows the first receiving user to more clearly understand the content of the vibrations.

[0344] (Output unit 202-2) The output unit 202-2 may transmit vibration data to the earphone 14-2 worn on the ear of the first receiving user. Specifically, the output unit 202-2 may transmit operation data directly to the earphone 14-2 worn on the ear of the first receiving user via wireless communication. This allows the vibration data to be transmitted in a shorter time than when the vibration data is transmitted to the earphone 14-2 worn on the ear of the first receiving user via the data processing device 12. Therefore, if the first receiving user is close enough to the second transmitting user to see them, communication can be established in a shorter time.

[0345] The earphone 14-2 of the first receiving user that has received the vibration data may generate vibrations according to the result of the identification process performed by the processing unit 201-2 in the vibration unit 43. Specifically, the earphone 14-2 of the first receiving user may generate vibration data that reproduces at least one of the intention and emotion of the user 20A based on the data generated according to the result of the identification process, and transmit the vibration data to the vibration unit 43.

[0346] For example, if the processing unit 201-2 transmits vibration data that generates vibrations with the lowest vibration level three times at one-second intervals to the earphone 14-2 of the first receiving user, the earphone 14-2 of the first receiving user may transmit a pulsed drive signal corresponding to the vibration data to the vibration unit 43 (e.g., a motor with an eccentric rotor). The drive signal is a signal that repeats high and low levels several dozen times within one second, generating a rectangular signal with a low voltage level three times at one-second intervals.

[0347] Furthermore, when the processing unit 201-2 transmits vibration data to the earphone 14-2 of the first receiving user, the vibration data being generated by generating the highest vibration level for three consecutive seconds, followed by ten consecutive vibrations with a one-second pause between them, the earphone 14-2 of the first receiving user may transmit a pulsed drive signal corresponding to the vibration data to the vibration unit 43 (for example, a motor with an eccentric rotor). The drive signal is a signal that repeats high and low levels several hundred times within three seconds, generating a rectangular signal with a high voltage level ten times at one-second intervals.

[0348] Next, an example of operation relating to specific processing by the data processing system 10-2 according to the 8-1 embodiment will be described with reference to FIG.

[0349] The earphone of the second user on the transmitting side may generate operation data in step S21, estimate emotions based on the operation data in step S22, generate vibration data in step S23, and transmit the vibration data to the earphone of the first user on the receiving side in step S24.

[0350] When the earphone of the receiving-side first user receives the vibration data in step S25, the earphone may cause the vibration unit 43 to generate vibrations according to the result of the identification process in step S26.

[0351] In the data processing system 10-2 according to the 8-1 embodiment, the results of the specific processing may be linked to a special application installed in the earphones used by the user 20. For example, by managing a history of the number of times the second user on the sending side touched the earphones and the details of the operations as a log in the application, even if the first user on the receiving side is unable to react to input operation data, the first user on the receiving side can check the history and react appropriately.

[0352] In addition, the operation data may include the intention of the sending second user, for example, being well or wanting to express gratitude, and may also include the sending second user's emotions, for example, joy, sadness, excitement, relief, etc.

[0353] As described above, the data processing system 10-2 according to the 8-1 embodiment enables earphones to communicate with each other directly or wirelessly via a data processing device, enabling real-time tactile sharing. Furthermore, because the earphones fit in the ears, tactile sensations due to vibrations can be felt naturally and directly. While conventional earphones are specialized for playing voice or music, the present disclosure combines vibration functionality with wireless communication to enable tactile sharing with remote locations. Furthermore, emotions and feelings that cannot be conveyed through words or images alone can be shared through tactile sensations, improving the quality of communication. Furthermore, intimate interactions can be realized, and tactile sharing enables intimate real-time communication with distant parties. Modern digital communication relies primarily on sight and hearing, but these means alone make it difficult to fully convey emotions and subtle nuances. Communication with distant parties, in particular, can lack intimacy and realism due to the lack of physical contact. The present disclosure provides a new means of communication that allows users to share tactile sensations with distant parties and communicate emotions on a deeper level. Furthermore, by linking the data processing device 12 and the earphones, it is possible to share complex tactile patterns among multiple users via the data processing device 12 as needed.

[0354] [Embodiment 8-2] (Outline of Embodiment 8-2) Embodiment 8-2 describes a configuration example in which operation data is generated and transmitted on the transmitting side, and emotion determination, vibration data generation, and vibration are performed on the receiving side. In embodiment 8-2, as shown in FIG. 27 , an example configuration is described in which earphones 14-3 worn by a specific receiving-side first user (user 20) input operation data transmitted from earphones 14-3 worn by a transmitting-side second user (user 20A) other than the receiving-side first user, and vibrate the housings 14a of the earphones 14-3 worn by the receiving-side first user based on the operation data. The operation data may include information indicating the detected operation. Hereinafter, to simplify the description, the earphones 14-3 may be simply referred to as earphones.

[0355] When the second transmitting user operates the earphones worn by the transmitting user with a finger in accordance with the strength and rhythm corresponding to the mood at the time, the message to be conveyed, etc., the earphones of the first receiving user vibrate in conjunction with the operation. For example, if the second transmitting user taps his / her earphones at a certain rhythm to inform the first receiving user that he / she will soon begin transmitting a specific message, the earphones of the first receiving user also vibrate at the same rhythm. For example, if the second transmitting user swipes his / her earphones to convey a specific message to the first receiving user (e.g., that he / she is tired or irritated), the earphones of the first receiving user also vibrate intermittently in accordance with the swipe. For example, if the second transmitting user presses and holds his / her earphones to inform the first receiving user that something urgent has happened, the earphones of the first receiving user also vibrate continuously while the press and hold is continued.

[0356] In this way, by having the earphones of the first receiving user vibrate in conjunction with the emotions, situation, etc. of the second transmitting user, emotions, feelings, etc. that cannot be conveyed through words alone can be shared between the first receiving user and the second transmitting user through vibrations, enabling more intimate interactions. Furthermore, even in situations where it is difficult to speak during a meeting or to operate the screen of a terminal device such as a smartphone, communication can be carried out between the first receiving user and the second transmitting user.

[0357] Furthermore, by conveying the emotions of the second transmitting user to the first receiving user through vibration, it is possible to communicate with the other party in real time even in situations where it is difficult to hear through voice, such as when there are many cars driving around the first receiving user or when the first receiving user is in a crowd.

[0358] Furthermore, by conveying the emotions of the second sending user to the first receiving user through vibration, the first receiving user can intuitively understand the emotions and intentions of the second sending user without losing concentration, compared to communicating by voice, even when the first receiving user is performing a specific task, such as desk work, attending a meeting, or driving a car.

[0359] Furthermore, by conveying the emotions of the second sending user to the first receiving user through vibration, even in situations where conversation is difficult, such as when the first receiving user is in a quiet space, there is no risk that people around the first receiving user will learn the emotions of the second sending user.

[0360] In particular, when the earphones are inner-ear type (open type), there is a risk that the sound will leak to the surroundings of the first receiving user. In contrast, according to the 8-2 embodiment, even if the earphones are open type, there is no risk that the emotions of the second transmitting user will be known by people around the first receiving user. The same applies to canal type earphones and inner-ear type earphones.

[0361] In the 8-2 embodiment, in addition to the emotion of the second transmitting user, information about the second transmitting user's body, such as body temperature, pulse rate, blood pressure fluctuations, etc., may be communicated by vibration. For example, the first receiving user can be notified by the vibration of the earphones that the pulse rate of the second transmitting user is increasing due to exercise.

[0362] In the 8-2 embodiment, the earphone of the first receiving user may directly communicate wirelessly with the earphone of the second transmitting user.

[0363] In the 8-2 embodiment, the case where the terminal device of the transmitting second user that transmits the operation data is an earphone is described, but the terminal device of the transmitting second user may also be a terminal device other than an earphone, such as a smartphone or a laptop computer.

[0364] 26 shows an example of the configuration of a data processing system 10-3 according to embodiment 8-2. The data processing system 10-3 may include a data processing device 12 and a plurality of earphones.

[0365] The user 20 wearing one of the earphones may be interpreted as the "first user on the receiving side" according to embodiment 8-2 of the present disclosure. The user 20A wearing the other earphone may be interpreted as the "second user on the transmitting side" according to embodiment 8-2 of the present disclosure.

[0366] Each earphone shown in FIG. 26 may include a processor 46-3, a RAM 48, a storage 50-3, a camera 42, a microphone 38, a touch sensor 39, a vibration unit 43, a speaker 41, and a communication I / F 44.

[0367] The touch sensor 39 may detect an operation on the earphones worn by the second transmitting user, such as at least one of a swipe operation, a tap operation, and a long press operation. The earphones of the second transmitting user may detect the operation using the touch sensor 39, generate operation data including information indicating the detected operation, and transmit the operation data to the earphones of the first receiving user. Specifically, the earphones of the second transmitting user may transmit the operation data directly to the earphones of the first receiving user via wireless communication, i.e., without going through the data processing device 12.

[0368] The earphones of the first receiving user that have received the operation data may generate vibration data based on the operation data. The vibration data may be interpreted as data that causes the earphones worn on the ears of the first receiving user to generate vibrations according to the results of the identification process, after the processor 46-3 performs identification processing using a data generation model that generates a predetermined estimation result according to the operation data.

[0369] The earphones of the first receiving user may execute a specific process to estimate at least one of the intentions and emotions of the second transmitting user based on the vibration data, and may generate vibrations according to the results of this specific process using a vibration unit 43 in the earphones worn by the first receiving user.

[0370] 27 shows an example of the main functions of the data processing device 12 and the earphones. In the earphones, a processor 46-3 performs data collection processing and identification processing.

[0371] The storage 50-3 stores a data collection program 60-3, a specific processing program 61-3, and a data generation model 62-3. The specific processing program 61-3 and the data generation model 62-3 are used for specific processing in the specific processing unit 46B-3.

[0372] The processor 46-3 reads the data collection program 60-3 from the storage 50-3 and executes the read data collection program 60-3 on the RAM 48. The data collection process is realized by the processor 46-3 operating as a control unit 46A-3 in accordance with the data collection program 60-3 executed on the RAM 48.

[0373] The processor 46-3 reads the specific processing program 61-3 from the storage 50-3 and executes the read specific processing program 61-3 on the RAM 48. The specific processing is realized by the processor 46-3 operating as a specific processing unit 46B-3 in accordance with the specific processing program 61-3 executed on the RAM 48.

[0374] The data generation model 62-3 is a so-called generative AI. Examples of the data generation model 62-3 include generative AI such as ChatGPT (registered trademark) (Internet search <URL: https: / / openai.com / blog / chatgpt>) and Gemini (registered trademark) (Internet search <URL: https: / / gemini.google.com / ?hl=ja>). The data generation model 62-3 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 62-3, and estimation data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 62-3 estimates the input estimation data in accordance with the instructions indicated by the prompt and outputs the estimation result in a data format such as voice data and text data. Here, estimation refers to, for example, analysis, classification, prediction, and / or summarization.

[0375] 28, the control unit 46A-3 may include a data collection unit 100-3. The specific processing unit 46B-3 may include an input unit 200-3, a processing unit 201-3, and a control unit 202-3.

[0376] The data collection unit 100-3 may collect the outputs of the microphone 38, the touch sensor 39, and the camera 42 shown in Fig. 28, and may further collect the operation data described above. As described above, the operation data may include information indicating the operation content detected by the touch sensor 39.

[0377] The operations include swiping, tapping, and long press operations on the earphones. A swipe operation may be interpreted as, for example, placing a finger on the touch sensor 39 and sliding the finger in any direction. A tap operation may be interpreted as, for example, touching the touch sensor 39 with a finger as if tapping it for a moment and then releasing it. A tap operation may include a double tap, which is a quick tap twice in succession. A long press operation may be interpreted as touching the touch sensor 39 with a finger and pressing it for a long time without releasing it.

[0378] For example, when user 20A, who is attending the same conference as user 20, comes up with a good idea and wants to let user 20 know about it, he or she performs, for example, a tap operation, thereby generating operation data indicating that it is a tap operation, and the operation data is sent to the earphones of user 20.

[0379] For example, in order to convey to user 20 that user 20A is in a bad mood and irritated during a meeting, user 20A performs a double tap operation, which generates operation data indicating that it is a double tap operation, and the operation data is transmitted to user 20's earphone.

[0380] (Processing Unit 201-3) The data collected by the data collection unit 100-3 may be input to the processing unit 201-3 via the input unit 200-3.

[0381] The data collection unit 100-3 or the input unit 200-3 may receive, via wireless communication, operation data transmitted from an earphone worn on the ear of the second transmitting user. Specifically, the input unit 200-3 may receive, via wireless communication, operation data directly from the earphone 14-3 worn on the ear of the first receiving user. This allows the transmission of the transmitted data in a shorter time than when the operation data is transmitted to the earphone 14-3 worn on the ear of the first receiving user via the data processing device 12. Therefore, if the first receiving user is located close enough to the second transmitting user to be able to see them, communication can be established in a shorter time.

[0382] The processing unit 201-3 may perform a determination process using a data generation model 62-3 (generative AI model) that generates a predetermined estimation result according to the input operation data. The processing unit 201-3 may generate vibration data that causes earphones worn on the ears of the first receiving user to generate vibrations according to the results of the determination process. For example, the processing unit 201-3 may execute a determination process that estimates at least one of the intentions and emotions of the second transmitting user based on the operation data, thereby causing the earphones worn on the ears of the first receiving user to generate vibrations that reproduce at least one of the intentions and emotions of the second transmitting user according to the results of the determination process.

[0383] Specifically, the processing unit 201-3 may use the data generation model 62-3 to estimate the emotions, intentions, etc. of the user 20A based on information contained in the operation data, such as the strength of the touch operation, the rhythm of the touch operation, and the duration of the touch operation.

[0384] For example, if the user 20A taps the earphones three times, the processing unit 201-3 recognizes that a pre-defined situation has occurred in which the user 20 wants the user 20 to notice the user 20A. In this case, the processing unit 201-3 may generate vibration data indicating a specific vibration period, a specific vibration intensity, a specific number of vibrations, etc., in order to accurately convey the user 20A's intention, and transmit the vibration data to the control unit 202-3. Specifically, the processing unit 201-3 may transmit vibration data that generates the lowest vibration level three times at one-second intervals to the control unit 202-3. The control unit 202-3 causes the vibration unit 43 to generate vibrations that reproduce at least one of the intention and emotion of the transmitting second user according to the result of the identification process, i.e., the vibration data. This allows, for example, the user 20A participating in a conference to quickly convey to the other users 20 participating in the same conference room that the user 20A wants to convey some kind of signal, instead of using a phone call or a short message.

[0385] For example, when user 20A presses and holds the earphones worn by the user 20A, for example, by continuously pressing the housing 14a with a finger for five seconds, the processing unit 201-3 recognizes that a situation has occurred in which user 20A needs to convey an urgent matter or an important matter to a predetermined user 20. In this case, the processing unit 201-3 may generate vibration data indicating a specific vibration period, a specific vibration intensity, a specific number of vibrations, etc., in order to accurately convey user 20A's intention, and transmit the vibration data to the control unit 202-3. Specifically, the processing unit 201-3 may transmit vibration data to the control unit 202-3 that generates vibrations at the highest vibration level for three consecutive seconds, followed by ten consecutive vibrations separated by a one-second pause. The control unit 202-3 controls the vibration unit 43 to generate vibrations that reproduce at least one of the intention and emotion of the second transmitting user, based on the result of the identification process, i.e., the vibration data. This allows the user 20A to quickly convey to the user 20, for example, that he / she has an urgent matter to convey, instead of making a phone call or sending a short message.

[0386] The processing unit 201-3 may perform the identification process by taking into account the operation history of the second sending user on the housing 14a worn on the second sending user's ear. For example, the processing unit 201-3 may generate history data that chronologically associates the operation content of the second sending user's touch operation at a specific time with the operation content of a touch operation performed at a specific time later than the specific time. This history data may be transmitted to the earphones of the first receiving user, and the history content may be played back on the first earphones. The history data may include message information such as, for example, "Mr. / Ms. XX sent you some kind of signal five minutes ago and one minute ago," or "Mr. / Ms. XX sent you a signal of agreement three minutes ago." By taking the operation history into account in this way, the first earphones can prevent the second sending user's signal from being overlooked and can also confirm the specific reaction to the operation.

[0387] While generating the vibration data, the processing unit 201-3 may generate audio data to be played from earphones worn on the ears of the first receiving user, as audio guidance, containing at least one of the intention and emotion of the second transmitting user, according to the result of the identification process, thereby enabling the first receiving user to more clearly understand the content of the vibration.

[0388] (Control unit 202-3) The control unit 202-3 may, in accordance with the result of the identification process, i.e., the vibration data, cause the vibration unit 43 to generate vibrations that reproduce at least one of the intention and emotion of the second transmitting user. Specifically, the earphone 14-3 of the first receiving user may, based on the vibration data generated in accordance with the result of the identification process, generate a drive signal that reproduces at least one of the intention and emotion of user 20A, and transmit the drive signal to the vibration unit 43.

[0389] For example, if the processing unit 201-3 generates vibration data that generates vibrations with the lowest vibration level three times at one-second intervals, the control unit 202-3 may transmit a pulsed drive signal corresponding to the vibration data to the vibration unit 43 (for example, a motor with an eccentric rotor). The drive signal is a signal that repeats high and low levels several tens of times within one second, and generates a rectangular signal with a low voltage level three times at one-second intervals.

[0390] Furthermore, when the processing unit 201-3 generates vibration data in which vibrations with the highest vibration level are generated for three consecutive seconds, followed by ten consecutive vibrations with a one-second pause between them, the control unit 202-3 may transmit a pulsed drive signal corresponding to the vibration data to the vibration unit 43 (for example, a motor with an eccentric rotor). The drive signal is a signal that repeats high and low levels several hundred times within three seconds, generating a rectangular signal with a high voltage level ten times at one-second intervals.

[0391] Next, an example of operation relating to specific processing by the data processing system 10-3 according to the 8-2 embodiment will be described with reference to FIG.

[0392] The earphones of the second transmitting user generate operation data in step S31 and transmit the operation data to the earphones of the first receiving user in step S32. The earphones of the second transmitting user receive the operation data in step S33, estimate an emotion based on the operation data in step S34, generate vibration data in step S35, and transmit a drive signal to the vibration unit 43 in step S36, thereby causing the vibration unit 43 to generate vibrations according to the results of the specific processing.

[0393] In the data processing system 10-3 according to the 8-2 embodiment, the results of the specific processing may be linked to a special application installed in the earphones used by the user 20. For example, by managing a history of the number of times the second user on the sending side touched the earphones and the details of the operations as a log in the application, even if the first user on the receiving side is unable to react to input operation data, the first user on the receiving side can check the history and react appropriately.

[0394] In addition, the operation data may include the intention of the sending second user, for example, being well or wanting to express gratitude, and may also include the sending second user's emotions, for example, joy, sadness, excitement, relief, etc.

[0395] As described above, the data processing system 10-3 according to the 8-2 embodiment enables earphones to communicate with each other directly or wirelessly via a data processing device, enabling real-time tactile sharing. Furthermore, because the earphones fit in the ears, tactile sensations due to vibrations can be felt naturally and directly. While conventional earphones are specialized for playing voice or music, the present disclosure combines vibration functionality with wireless communication to enable tactile sharing with remote locations. Furthermore, emotions and feelings that cannot be conveyed through words or images alone can be shared through tactile sensations, improving the quality of communication. Furthermore, intimate interactions can be realized, and tactile sharing enables intimate real-time communication with distant parties. Modern digital communication relies primarily on sight and hearing, but these means alone make it difficult to fully convey emotions and subtle nuances. Communication with distant parties, in particular, can lack intimacy and realism due to the lack of physical contact. The present disclosure provides a new means of communication that allows users to share tactile sensations with distant parties and communicate emotions on a deeper level. Furthermore, by linking the data processing device 12 and the earphones, it is possible to share complex tactile patterns among multiple users via the data processing device 12 as needed.

[0396] [Embodiment 8-3] (Outline of Embodiment 8-3) Embodiment 8-3 describes a configuration example in which operation data is generated and transmitted on the transmitting side, emotion determination is performed on the processing device, vibration data is generated and transmitted, and vibration occurs on the receiving side. In embodiment 8-3, as shown in FIG. 30, a data processing device 12-4 receives operation data transmitted from earphones 14-4 worn by a second transmitting user (user 20A) other than the first receiving user, and vibrates the housing 14a of the earphones 14-4 worn by the first receiving user based on the operation data. The operation data may include information indicating the detected operation. Hereinafter, to simplify the description, the earphones 14-4 may be simply referred to as earphones.

[0397] When the second transmitting user operates the earphones worn by the transmitting user with a finger in accordance with the strength and rhythm corresponding to the mood at the time, the message to be conveyed, etc., the earphones of the first receiving user vibrate in conjunction with the operation. For example, if the second transmitting user taps his / her earphones at a certain rhythm to inform the first receiving user that he / she will soon begin transmitting a specific message, the earphones of the first receiving user also vibrate at the same rhythm. For example, if the second transmitting user swipes his / her earphones to convey a specific message to the first receiving user (e.g., that he / she is tired or irritated), the earphones of the first receiving user also vibrate intermittently in accordance with the swipe. For example, if the second transmitting user presses and holds his / her earphones to inform the first receiving user that something urgent has happened, the earphones of the first receiving user also vibrate continuously while the press and hold is continued.

[0398] In this way, by having the earphones of the first receiving user vibrate in conjunction with the emotions, situation, etc. of the second transmitting user, emotions, feelings, etc. that cannot be conveyed through words alone can be shared between the first receiving user and the second transmitting user through vibrations, enabling more intimate interactions. Furthermore, even in situations where it is difficult to speak during a meeting or to operate the screen of a terminal device such as a smartphone, communication can be carried out between the first receiving user and the second transmitting user.

[0399] Furthermore, by conveying the emotions of the second transmitting user to the first receiving user through vibration, it is possible to communicate with the other party in real time even in situations where it is difficult to hear through voice, such as when there are many cars driving around the first receiving user or when the first receiving user is in a crowd.

[0400] Furthermore, by conveying the emotions of the second sending user to the first receiving user through vibration, the first receiving user can intuitively understand the emotions and intentions of the second sending user without losing concentration, compared to communicating by voice, even when the first receiving user is performing a specific task, such as desk work, attending a meeting, or driving a car.

[0401] Furthermore, by conveying the emotions of the second sending user to the first receiving user through vibration, even in situations where conversation is difficult, such as when the first receiving user is in a quiet space, there is no risk that people around the first receiving user will learn the emotions of the second sending user.

[0402] In particular, when the earphones are inner-ear type (open type), there is a risk that the sound will leak to the surroundings of the first receiving user. In contrast, according to the 8D embodiment, even if the earphones are open type, there is no risk that the emotions of the second transmitting user will be known by people around the first receiving user. The same applies to canal type earphones and inner-ear type earphones.

[0403] In the 8-3 embodiment, in addition to the emotion of the second transmitting user, information about the second transmitting user's body, such as body temperature, pulse rate, blood pressure fluctuations, etc., may be communicated by vibration. For example, the first receiving user can be notified by the vibration of the earphone that the pulse rate of the second transmitting user is increasing due to exercise.

[0404] In the 8-3 embodiment, the earphone of the first receiving user may directly communicate wirelessly with the earphone of the second transmitting user.

[0405] In the 8-3 embodiment, the case where the terminal device of the transmitting second user that transmits the operation data is an earphone is described, but the terminal device of the transmitting second user may also be a terminal device other than an earphone, such as a smartphone or a laptop computer.

[0406] 30 shows an example of the configuration of a data processing system 10-4 according to embodiment 8-3. The data processing system 10-4 may include a data processing device 12 and a plurality of earphones.

[0407] The user 20 wearing one of the earphones may be interpreted as the "first user on the receiving side" according to embodiment 8D of the present disclosure, and the user 20A wearing the other earphone may be interpreted as the "second user on the transmitting side" according to embodiment 8-3 of the present disclosure.

[0408] Each earphone shown in FIG. 30 may include a processor 46-4, a RAM 48, a storage 50-4, a camera 42, a microphone 38, a touch sensor 39, a vibration unit 43, a speaker 41, and a communication I / F 44.

[0409] The touch sensor 39 may detect an operation on the earphones worn by the second transmitting user, such as at least one of a swipe operation, a tap operation, and a long press operation. The earphones of the second transmitting user may detect the operation with the touch sensor 39, generate operation data including information indicating the detected operation, and transmit the operation data to the data processing system 10-4. Specifically, the earphones of the second transmitting user may transmit the operation data to the data processing system 10-4 via wireless communication.

[0410] The data processing system 10-4 that has received the operation data may generate vibration data based on the operation data. The processor 46-4 may perform a specific process on the vibration data using a data generation model that generates a predetermined estimation result according to the operation data, and the vibration data may be interpreted as data to generate vibrations according to the result of the specific process in earphones worn on the ears of the first user on the receiving side.

[0411] The data processing system 10-4 may execute a specific process to estimate at least one of the intentions and emotions of the second transmitting user based on the vibration data, and may generate vibrations according to the results of this specific process using a vibration unit 43 in the earphones worn by the first receiving user.

[0412] 31 shows an example of the main functions of the data processing device 12 and the earphones. In the earphones, a processor 46-4 performs data collection processing and identification processing.

[0413] The storage 50-4 stores a data collection program 60-4 and a specific processing program 61-4, which is used for specific processing in the specific processing unit 46B-4.

[0414] The processor 46-4 reads the data collection program 60-4 from the storage 50-4 and executes the read data collection program 60-4 on the RAM 48. The data collection process is realized by the processor 46-4 operating as the control unit 46A in accordance with the data collection program 60-4 executed on the RAM 48.

[0415] The processor 46-4 reads the specific processing program 61-4 from the storage 50-4 and executes the read specific processing program 61-4 on the RAM 48. The specific processing is realized by the processor 46-4 operating as a specific processing unit 46B-4 in accordance with the specific processing program 61-4 executed on the RAM 48.

[0416] In the data processing device 12-4, a processor 28-4 performs data collection processing and specific processing.

[0417] The storage 32-4 stores a data collection program 55-4, a specific processing program 56-4, and a data generation model 58-4. The specific processing program 56-4 and the data generation model 58-4 are used for specific processing in the specific processing unit 290-4.

[0418] The processor 28-4 reads the data collection program 55-4 from the storage 32-4 and executes the read data collection program 55-4 on the RAM 30. The data collection process is realized by the processor 28-4 operating as a control unit 291-4 in accordance with the data collection program 55-4 executed on the RAM 30.

[0419] The processor 28-4 reads the specific processing program 56-4 from the storage 32-4 and executes the read specific processing program 56-4 on the RAM 30. The specific processing is realized by the processor 28-4 operating as a specific processing unit 290-4 in accordance with the specific processing program 56-4 executed on the RAM 30.

[0420] The data generation model 58-4 is a so-called generative AI. Examples of the data generation model 58-4 include generative AI such as ChatGPT (registered trademark) (Internet search <URL: https: / / openai.com / blog / chatgpt>) and Gemini (registered trademark) (Internet search <URL: https: / / gemini.google.com / ?hl=ja>). The data generation model 58-4 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58-4, and estimation data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58-4 estimates the input estimation data in accordance with the instructions indicated by the prompt and outputs the estimation result in a data format such as voice data and text data. Here, estimation refers to, for example, analysis, classification, prediction, and / or summarization.

[0421] The data generation model 58-4 is a so-called generative AI. Examples of the data generation model 58-4 include generative AI such as ChatGPT (registered trademark) (Internet search <URL: https: / / openai.com / blog / chatgpt>) and Gemini (registered trademark) (Internet search <URL: https: / / gemini.google.com / ?hl=ja>). The data generation model 58-4 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58-4, and estimation data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58-4 estimates the input estimation data in accordance with the instructions indicated by the prompt and outputs the estimation result in a data format such as voice data and text data. Here, estimation refers to, for example, analysis, classification, prediction, and / or summarization.

[0422] 32, the control unit 291-4 may include a data collection unit 100-4. The specific processing unit 290-4 may include an input unit 200-4, a processing unit 201-4, and an output unit 202-4.

[0423] The data collection unit 100-4 may collect the outputs of the microphone 38, the touch sensor 39, and the camera 42 shown in Fig. 30, and may further collect the operation data described above. As described above, the operation data may include information indicating the operation content detected by the touch sensor 39.

[0424] The operations include swiping, tapping, and long press operations on the earphones. A swipe operation may be interpreted as, for example, placing a finger on the touch sensor 39 and sliding the finger in any direction. A tap operation may be interpreted as, for example, touching the touch sensor 39 with a finger as if tapping it for a moment and then releasing it. A tap operation may include a double tap, which is a quick tap twice in succession. A long press operation may be interpreted as touching the touch sensor 39 with a finger and pressing it for a long time without releasing it.

[0425] For example, when user 20A, who is attending the same meeting as user 20, comes up with a good idea and wants to let user 20 know about it, he or she performs a tap operation, for example, to generate operation data indicating that it is a tap operation, and the operation data is transmitted to data processing device 12-4.

[0426] For example, in order to convey to user 20 that user 20A is in a bad mood and irritated during a meeting, user 20A performs a double tap operation, which generates operation data indicating that it is a double tap operation, and the operation data is transmitted to data processing device 12-4.

[0427] (Processing Unit 201-4) The data collected by the data collection unit 100-4 may be input to the processing unit 201-4 via the input unit 200-4.

[0428] The data collection unit 100-4 or the input unit 200-4 may wirelessly receive the operation data transmitted from the earphones worn on the ears of the second transmitting user. Specifically, the input unit 200-4 may wirelessly receive the operation data directly from the earphones 14-4 worn on the ears of the first receiving user. As a result, even if the first receiving user is not close enough to see the second transmitting user and the wireless data from the second transmitting user does not reach the first receiving user, the operation data is received by the data processing device 12, allowing the first receiving user and the second transmitting user to communicate.

[0429] The processing unit 201-4 may perform a determination process using a data generation model 58-4 (generative AI model) that generates a predetermined estimation result according to the input operation data. The processing unit 201-4 may generate vibration data that causes earphones worn on the ears of the first receiving user to generate vibrations according to the results of the determination process. For example, the processing unit 201-4 may execute a determination process that estimates at least one of the intentions and emotions of the second transmitting user based on the operation data, thereby causing the earphones worn on the ears of the first receiving user to generate vibrations that reproduce at least one of the intentions and emotions of the second transmitting user according to the results of the determination process.

[0430] Specifically, the processing unit 201-4 may use the data generation model 58-4 to estimate the emotions, intentions, etc. of the user 20A based on information contained in the operation data, such as the strength of the touch operation, the rhythm of the touch operation, and the duration of the touch operation.

[0431] For example, when a tap operation on the earphones worn by the user 20A is performed three times, the processing unit 201-4 recognizes that a preset situation has occurred in which the user 20 wants the user 20 to notice the user 20A. In this case, the processing unit 201-4 may generate vibration data indicating a specific vibration period, a specific vibration intensity, a specific number of vibrations, etc., in order to accurately convey the intention of the user 20A, and transmit the vibration data to the output unit 202-4. Specifically, the processing unit 201-4 may transmit vibration data that generates vibrations at the lowest vibration level three times at one-second intervals to the output unit 202-4.

[0432] The output unit 202-4 may transmit the vibration data to an earphone worn on the ear of the receiving-side first user.

[0433] The data collection unit 70 of the earphone of the first receiving user that has received the vibration data transfers the collected vibration data to the control unit 46A. The control unit 46A controls the vibration unit 43 to generate vibrations that reproduce at least one of the intention and emotion of the second transmitting user in accordance with the vibration data. This allows, for example, user 20A participating in a conference to quickly convey to other users 20 participating in the same conference room that user 20A wants to convey some kind of signal, instead of by phone call or text message.

[0434] For example, when user 20A presses and holds the earphones worn by the user 20A, for example, by pressing the housing 14a with a finger for five seconds, the processing unit 201-4 recognizes that a situation has occurred in which user 20A needs to convey an urgent matter or an important matter to a pre-defined user 20. In this case, the processing unit 201-4 may generate vibration data indicating a specific vibration period, a specific vibration intensity, a specific number of vibrations, etc., in order to accurately convey user 20A's intention, and transmit the vibration data to the earphones of the first receiving user via the output unit 202-4. Specifically, the processing unit 201-4 may transmit vibration data that generates vibrations at the highest vibration level for three consecutive seconds, followed by ten consecutive vibrations separated by a one-second pause. The control unit 46A of the earphones of the first receiving user causes the vibration unit 43 to generate vibrations that reproduce at least one of the intention and emotion of the second transmitting user in accordance with the vibration data. This allows the user 20A to quickly convey to the user 20, for example, that he / she has an urgent matter to convey, instead of making a phone call or sending a short message.

[0435] The processing unit 201-4 may perform the identification process by taking into account the operation history of the second sending user on the housing 14a worn on the second sending user's ear. For example, the processing unit 201-4 may generate history data that chronologically associates the operation content of the second sending user's touch operation at a specific time with the operation content of a touch operation performed at a specific time later than the specific time. This history data may be transmitted to the earphones of the first receiving user, and the history content may be played back on the first earphones. The history data may include message information such as, for example, "Mr. / Ms. XX sent you some kind of signal five minutes ago and one minute ago," or "Mr. / Ms. XX sent you a signal of agreement three minutes ago." By taking the operation history into account in this way, the first earphones can prevent the second sending user's signal from being overlooked and can also confirm the specific reaction to the operation.

[0436] While generating the vibration data, the processing unit 201-4 may generate audio data to be played from earphones worn on the ears of the first receiving user, as audio guidance, containing at least one of the intention and emotion of the second transmitting user, according to the result of the identification process, thereby enabling the first receiving user to more clearly understand the content of the vibration.

[0437] The control unit 46A may, in accordance with the result of the identification process, i.e., the vibration data, cause the vibration unit 43 to generate vibrations that reproduce at least one of the intention and emotion of the second transmitting user. Specifically, the control unit 46A may generate a drive signal that reproduces at least one of the intention and emotion of the user 20A based on the vibration data generated in accordance with the result of the identification process, and transmit the drive signal to the vibration unit 43.

[0438] For example, when the processing unit 201-4 generates vibration data that generates vibrations with the lowest vibration level three times at one-second intervals, the control unit 46A may transmit a pulsed drive signal corresponding to the vibration data to the vibration unit 43 (for example, a motor having an eccentric rotor). The drive signal is a signal that repeats high and low levels several tens of times within one second, and generates a rectangular signal with a low voltage level three times at one-second intervals.

[0439] Furthermore, when the processing unit 201-4 generates vibration data in which vibrations with the highest vibration level are generated for three consecutive seconds, followed by ten consecutive vibrations with a one-second pause between them, the control unit 46A may transmit a pulsed drive signal corresponding to the vibration data to the vibration unit 43 (for example, a motor having an eccentric rotor). The drive signal is a signal that repeats high and low levels several hundred times within three seconds, for example, and generates a rectangular signal with a high voltage level ten times at one-second intervals.

[0440] Next, an example of operation relating to specific processing by the data processing system 10-4 according to the 8-3 embodiment will be described with reference to FIG.

[0441] The earphones of the second transmitting user generate operation data in step S41 and transmit the operation data to the data processing device 12-4 in step S42. The data processing device 12-4 receives the operation data in step S43, estimates the emotion based on the operation data in step S44, generates vibration data in step S45, and transmits the vibration data to the earphones of the first receiving user in step S46.

[0442] The earphones of the first receiving user may receive the vibration data in step S47, and in step S48, transmit a drive signal to the vibration unit 43, thereby causing the vibration unit 43 to generate vibrations according to the results of the specific processing.

[0443] In the data processing system 10-4 according to the 8-3 embodiment, the results of the specific processing may be linked to a special application installed in the earphones used by the user 20. For example, by managing a history of the number of times the second user on the sending side touched the earphones and the details of the operations as a log in the application, even if the first user on the receiving side is unable to react to input operation data, the first user on the receiving side can check the history and react appropriately.

[0444] In addition, the operation data may include the intention of the sending second user, for example, being well or wanting to express gratitude, and may also include the sending second user's emotions, for example, joy, sadness, excitement, relief, etc.

[0445] As described above, the data processing system 10-4 according to the 8-3 embodiment enables earphones to communicate with each other directly or wirelessly via a data processing device, enabling real-time tactile sharing. Furthermore, because the earphones fit in the ears, tactile sensations due to vibrations can be felt naturally and directly. While conventional earphones are specialized for playing voice or music, the present disclosure combines vibration functionality with wireless communication to enable tactile sharing with remote locations. Furthermore, emotions and feelings that cannot be conveyed through words or images alone can be shared through tactile sensations, improving the quality of communication. Furthermore, intimate interactions can be realized, and tactile sharing enables intimate real-time communication with distant parties. Modern digital communication relies primarily on sight and hearing, but these means alone make it difficult to fully convey emotions and subtle nuances. Communication with distant parties, in particular, can lack intimacy and realism due to the lack of physical contact. The present disclosure provides a new means of communication that allows users to share tactile sensations with distant parties and communicate emotions on a deeper level. Furthermore, by linking the data processing device 12-4 and the earphones, it is possible to share complex tactile patterns among multiple users via the data processing device 12-4 as needed.

[0446] In addition, the following supplementary notes are provided in relation to the above description.

[0447] (Supplementary Note 1) An earphone comprising: a housing worn on the ear of a second sending-side user who communicates with a first receiving-side user; a data collection unit that collects operation data generated by detecting at least one of a swipe operation, a tap operation, and a long press operation on the housing; a processing unit that performs a specification process using a data generation model that generates a predetermined estimation result according to the operation data, and generates vibration data for an earphone worn on the ear of the first receiving-side user to generate vibrations according to the result of the specification process; and an output unit that transmits the vibration data to the earphone worn on the ear of the first receiving-side user, wherein the processing unit executes, as the specification process, a process of estimating at least one of an intention and an emotion of the second sending-side user based on the operation data, thereby generating vibrations in the earphone worn on the ear of the first receiving-side user that reproduce at least one of the intention and the emotion according to the result of the specification process. (Supplementary Note 2) The earphone according to Supplementary Note 1, wherein the processing unit performs the identification process taking into consideration an operation history of a housing worn on the ear of the second transmitting user by the second transmitting user. (Supplementary Note 3) The earphone according to Supplementary Note 1, wherein the output unit transmits the operation data via wireless communication. (Supplementary Note 4) The earphone according to Supplementary Note 1, wherein the processing unit generates, while generating the vibration, audio data to be played from the earphone worn on the ear of the first receiving user, using at least one of the intention and the emotion as audio guidance according to a result of the identification process, and the output unit transmits the audio data via wireless communication.(Supplementary Note 5) Earphones comprising: a vibration unit that vibrates a housing worn on the ear of a receiving-side first user; an input unit that receives operation data generated by detecting at least one of a swipe operation, a tap operation, and a long press operation on earphones worn on the ear of the second sending-side user by a second sending-side user who communicates with the first receiving-side user; a processing unit that performs identification processing using a data generation model that generates a predetermined estimation result according to the input operation data; and a control unit that causes the vibration unit to generate vibrations according to the result of the identification processing, wherein the processing unit executes, as the identification processing, a process of estimating at least one of an intention and an emotion of the second sending-side user based on the operation data, and the control unit causes the vibration unit to generate vibrations that reproduce at least one of the intention and the emotion according to the result of the identification processing. (Supplementary Note 6) The earphones according to Supplementary Note 5, wherein the processing unit performs the identification processing in consideration of an operation history of the earphones worn on the ear of the second sending-side user by the second sending-side user. (Supplementary Note 7) The earphone according to Supplementary Note 5, wherein the input unit receives the operation data transmitted from an earphone worn on the ear of the transmitting-side second user via wireless communication. (Supplementary Note 8) The earphone according to Supplementary Note 5, wherein the control unit generates audio data for reproducing at least one of the intention and the emotion as audio guidance in accordance with a result of the identification process while causing the vibration unit to generate vibrations.(Supplementary Note 9) A data processing device comprising: an input unit that receives operation data generated by detecting at least one of a swipe operation, a tap operation, and a long press operation performed by a second transmitting-side user communicating with a first receiving-side user wearing earphones on the earphones worn on the ears of the second transmitting-side user; a processing unit that performs identification processing using a data generation model that generates a predetermined estimation result according to the input operation data; a control unit that generates vibration data that causes the earphones worn on the ears of the first receiving-side user to generate vibrations according to the result of the identification processing; and an output unit that transmits the vibration data to the earphones worn on the ears of the first receiving-side user, wherein the processing unit executes, as the identification processing, a process of estimating at least one of an intention and an emotion of the second transmitting-side user based on the operation data, and the control unit generates the vibration data that reproduces at least one of the intention and the emotion according to the result of the identification processing. (Supplementary Note 10) The data processing device according to Supplementary Note 9, wherein the processing unit performs the identification process in consideration of an operation history of an earphone worn on the ear of the second transmitting user by the second transmitting user. (Supplementary Note 11) The data processing device according to Supplementary Note 9, wherein the input unit receives the operation data transmitted from the earphone worn on the ear of the second transmitting user via wireless communication. (Supplementary Note 12) The data processing device according to Supplementary Note 9, wherein the control unit generates audio data for playing back at least one of the intention and the emotion as audio guidance according to a result of the identification process while causing vibrations in the earphone worn on the ear of the first receiving user, and the output unit transmits the audio guidance to the earphone worn on the ear of the first receiving user.(Supplementary Note 13) A data processing method comprising: at least one processor performing an identification process using a data generation model that generates a predetermined estimation result according to operation data generated by detecting at least one of a swipe operation, a tap operation, and a long press operation on an earphone worn on the ear of a second sending-side user who communicates with a first receiving-side user wearing earphones, by the second sending-side user; generating vibration data that causes the earphone worn on the ear of the first receiving-side user to generate a vibration according to the result of the identification process; transmitting the vibration data to the earphone worn on the ear of the first receiving-side user; performing, as the identification process, a process that estimates at least one of an intention and an emotion of the second sending-side user based on the operation data; and generating the vibration data that reproduces at least one of the intention and the emotion according to the result of the identification process. (Supplementary Note 14) A data processing program that causes at least one processor to execute processes including: performing a determination process using a data generation model that generates a predetermined estimation result according to operation data generated by detecting at least one of a swipe operation, a tap operation, and a long press operation on an earphone worn in the ear of a second sending-side user who communicates with a first receiving-side user wearing earphones, by the second sending-side user; generating vibration data that causes the earphone worn in the ear of the first receiving-side user to generate vibrations according to the result of the determination process; transmitting the vibration data to the earphone worn in the ear of the first receiving-side user; executing, as the determination process, a process that determines at least one of an intention and an emotion of the second sending-side user based on the operation data; and generating the vibration data that reproduces at least one of the intention and the emotion according to the result of the determination process.

[0448] Ninth Embodiment An example of an embodiment of a data processing device, a data processing method, and a program according to the technology of the present disclosure will be described below with reference to the accompanying drawings. Note that the same parts as those in the first embodiment are designated by the same reference numerals, and detailed description thereof will be omitted.

[0449] Modern urban environments are overflowing with visual information such as digital signage, advertisements, and neon signs, which can distract users and reduce their ability to concentrate. The data processing system according to this embodiment solves the problem of visual information overload. Details of the data processing system according to this embodiment are described below.

[0450] 34 is a diagram showing the flow of various data transmitted and received between the earphones 14 and the data processing device 12. The camera 42 of the earphones 14 captures images of the surroundings of the user 20 to generate video data. The images captured by the camera 42 include images of information transmission media that are present around the user 20 and that provide visual information to an unspecified number of people, such as digital signage, advertisements, billboards, and signs. The video data generated by the camera 42 is transmitted to the data processing device 12 via the communication I / F 44 of the earphones 14.

[0451] The processor 28 of the data processing device 12 operates as a specific processing unit 290A by executing the specific processing program 56 on the RAM 30. As shown in Fig. 35 , the specific processing unit 290A has an acquisition unit 501, an extraction unit 502, a selection unit 503, a message generation unit 504, and a communication unit 505.

[0452] The acquisition unit 501 acquires video data transmitted from the earphones 14. The extraction unit 502 identifies video of information transmission media such as digital signage, advertisements, billboards, and signs from the video data acquired by the acquisition unit 501. The extraction unit 502 extracts visual information intended for an unspecified number of people and provided by the identified information transmission media from the video of the identified information transmission media. Visual information is information represented by characters, symbols, figures, marks, etc. that can be recognized visually. Visual information may include, for example, sales information, store information, weather information, current events information, time information, traffic information, warning information, guidance information, etc. The extraction unit 502 extracts visual information from the video data using image recognition technology.

[0453] The selection unit 503 selects the visual information extracted by the extraction unit 502 based on user profile data indicating the characteristics or attributes of the user 20. The user profile data is recorded in the database 24. The user profile data includes the user 20's hobbies, interests, behavioral history, web browsing history, product purchase history, application usage history, schedule, age, gender, and affiliation. The user profile data may be acquired by the data processing device 12 linking with a user terminal such as a smartphone or personal computer used by the user 20. For example, web browsing, product purchase, and application usage history performed using the user terminal may be transmitted from the user terminal to the data processing device 12. Some or all of the user profile data may be provided to the data processing device 12 by an input operation on the user terminal. The processor 28 of the data processing device 12 records the acquired user profile data in the database 24.

[0454] The selection unit 503 identifies the requests, desires, interests, and concerns of the user 20 based on the user profile data. The selection unit 503 selects visual information corresponding to the requests, desires, interests, and concerns of the user 20 from the visual information extracted by the extraction unit 502. That is, from the visual information, information that the user 20 is interested in, information that is useful to the user 20, and information that the user 20 needs is selected. For example, if it is identified from the user profile data that the user 20 is interested in a specific event, visual information related to the event is selected.

[0455] The message generator 504 uses the data generation model 58 to generate a message including the content of the visual information selected by the selector 503. The message generator 504 inputs a prompt sentence to the data generation model 58, instructing the data generation model 58 to generate a voice that explains the content of the visual information selected by the selector 503. The data generation model 58 generates a voice that explains the content of the visual information selected based on the prompt sentence.

[0456] The communication unit 505 transmits the voice data of the message voice generated by the message generation unit 504 to the earphone 14 .

[0457] FIG. 36 is a flowchart showing an example of the flow of the specification process carried out by the specification processing unit 290A of the data processing device 12.

[0458] In step S310, the acquisition unit 501 acquires the video data transmitted from the earphones 14. The video data is generated by the camera 42 of the earphones 14 capturing an image of the surroundings of the user 20. The video data includes images of information transmission media for providing visual information to an unspecified number of people, such as digital signage, advertisements, billboards, and signs.

[0459] In step S311, the extraction unit 502 extracts visual information directed at an unspecified number of people from the video data acquired in step S310.

[0460] In step S312, the selection unit 503 selects the visual information extracted in step S311 based on the user profile data recorded in the database 24. The user profile data includes at least one of the hobbies, interests, behavioral history, web browsing history, application usage history, schedule, age, gender, and affiliation of the user 20.

[0461] In step S313, the message generation unit 504 generates a message including the content of the visual information selected in step S312, using the data generation model 58. Specifically, the message generation unit 504 inputs a prompt sentence to the data generation model 58, instructing the data generation model 58 to generate a voice that explains the content of the visual information selected by the selection unit 503. The data generation model 58 generates a voice that explains the content of the visual information selected based on the prompt sentence.

[0462] In step S314, the communication unit 505 transmits audio data of the message generated in step S313 to the earphone 14. The earphone 14 outputs an audio message including the content of the visual information selected based on the user profile data. The processes from step S310 to step S314 are performed in real time. That is, the visual information captured by the camera 42 of the earphone 14 is immediately selected by the data processing device 12, and audio data corresponding to the content is provided to the user 20.

[0463] As described above, the data processing device 12 according to this embodiment includes an acquisition unit 501 that acquires video data including images of the user 20's surroundings, an extraction unit 502 that extracts visual information intended for an unspecified number of people from the video data, a database 24 that stores user profile data indicating the characteristics or attributes of the user, a selection unit 503 that selects visual information based on the user profile data, and a message generation unit 504 that uses a data generation model 58 to generate a message including the content of the selected visual information.

[0464] According to the data processing system 10 of this embodiment, two cameras 42 attached to two earphones 14 worn by the user 20 are used to acquire video data captured at an angle of view that substantially matches the visual range of the user 20. The video data includes visual information provided via information transmission media such as digital signage, advertisements, billboards, and signs. The overwhelming amount of visual information surrounding the user 20 includes much information that the user 20 is not interested in or concerned with, and information that is unnecessary for the user 20. According to the data processing system 10 of this embodiment, information that the user 20 is interested in, information that is useful to the user 20, and information that the user 20 needs is selected from the visual information and provided to the user 20. In other words, information that the user 20 is not interested in or concerned with and information that is unnecessary for the user 20 is blocked. This solves the problem of excessive visual information, which distracts the user 20 and reduces their concentration.

[0465] 37 , the earphones 14 may have a vibration motor 43. When the volume of the environmental sound input to the microphone 38 of the earphones 14 exceeds a threshold, the communication unit 505 of the data processing device 12 may transmit a control signal to the earphones 14 to vibrate the vibration motor 43 together with the message audio. This makes it possible to notify the user 20 of the arrival of a message containing the selected visual information content, even when the user 20 is in a noisy environment.

[0466] 37 , the earphones 14 may also have a sensor 45 for acquiring biometric information of the user 20. The biometric information acquired by the sensor 45 may be, for example, blood pressure, heart rate, body temperature, or brain waves. The data processing device 12 may estimate the user's emotions by inputting the biometric information acquired by the sensor 45 into a data generation model 58. The selection unit 503 may select visual information based on the estimated emotions of the user 20. For example, if the estimated emotions of the user 20 are negative, visual information that promotes the negative emotions may be excluded by the selection process of the selection unit 503.

[0467] The following supplementary note is further disclosed in relation to the above description: (Supplementary note 1) A data processing device including: an acquisition unit that acquires video data including a video of a user's surroundings, an extraction unit that extracts visual information intended for an unspecified number of people from the video data, a database that stores user profile data indicating characteristics or attributes of the user, a selection unit that selects the visual information based on the user profile data, and a message generation unit that generates a message including the content of the selected visual information using a data generation model.

[0468] (Supplementary Note 2) The data processing device according to Supplementary Note 1, wherein the extraction unit extracts the visual information from an image of an information transmission medium included in the video data.

[0469] (Supplementary Note 3) The data processing device according to Supplementary Note 1 or Supplementary Note 2, wherein the user profile data includes at least one of the user's hobbies, interests, behavioral history, web browsing history, application usage history, schedule, age, gender, and affiliation.

[0470] (Supplementary Note 4) A data processing system including: a data processing device according to any one of Supplementary Note 1 to Supplementary Note 3; and earphones worn by the user, wherein the earphones include a camera that captures an image of the user's surroundings and generates the video data, and a speaker that outputs the message, and the data processing device includes a communication unit that transmits audio data including the content of the message to the earphones.

[0471] (Supplementary Note 5) The data processing system according to Supplementary Note 4, wherein the earphone includes a microphone and a vibration motor, and the communication unit transmits a control signal for vibrating the vibration motor to the earphone together with the audio data when the volume of the environmental sound input to the microphone exceeds a threshold.

[0472] (Supplementary Note 6) The data processing system according to Supplementary Note 4 or Supplementary Note 5, wherein the earphone includes a sensor for acquiring biometric information, and the selection unit selects the visual information based on an emotion of the user estimated based on the biometric information.

[0473] (Supplementary Note 7) A data processing method in which a computer acquires video data including images of a user's surroundings, extracts visual information intended for an unspecified number of people from the video data, selects the visual information based on user profile data that indicates the characteristics or attributes of the user, and generates a message including the content of the selected visual information using a data generation model.

[0474] (Supplementary Note 8) A program for causing a computer to execute the following processes: acquiring video data including images of the user's surroundings; extracting visual information intended for an unspecified number of people from the video data; selecting the visual information based on user profile data indicating the characteristics or attributes of the user; and generating a message including the content of the selected visual information using a data generation model.

[0475] Tenth Embodiment An example of an embodiment of a data processing device, a data processing method, and a program according to the technology of the present disclosure will be described below with reference to the accompanying drawings. Note that the same parts as those in the first embodiment are designated by the same reference numerals, and detailed description thereof will be omitted.

[0476] The present embodiment relates to earphones that analyze a user's visual information in real time and generate music and haptic feedback. In particular, the present embodiment relates to a technology that provides a synesthesia experience that combines visual, auditory, and haptic sensations.

[0477] Conventional earphones have focused on audio functions such as voice assistants, noise cancellation, and translation functions. However, these earphones have difficulty meeting users' needs for new experiences that incorporate visual information. For example, when viewing a painting in a museum or a natural landscape, users need a way to combine visual information with other senses to achieve a deeper sense of immersion and excitement.

[0478] However, conventional technologies have had difficulty analyzing visual information in real time and converting it into music or haptic feedback, which has led to the issue of not being able to provide users with a new sensory experience that combines vision, hearing, and touch.

[0479] Therefore, in this embodiment, visual information (color, shape, movement, etc.) acquired by a camera equipped in the earphone is analyzed in real time, and music or sound effects, etc. are generated using a data generation model 58. Also, in the 10B embodiment, a vibration function equipped in the earphone is utilized to provide haptic feedback synchronized with the generated music or sound effects, etc.

[0480] This embodiment allows users to have a new sensory experience that combines vision, hearing, and touch. For example, when appreciating a painting in a museum or a natural landscape, the visual information the user receives is fed back to them as music and touch, allowing the user to enjoy a deeper sense of immersion and excitement. This embodiment also provides a multimodal experience that was not possible with conventional earphones.

[0481] Hereinafter, this embodiment will be described in detail with reference to the drawings.

[0482] <System Configuration> Fig. 38 shows an example of the configuration of a data processing system 210 according to this embodiment. As shown in Fig. 38, the earphones of this embodiment are two earphones 214 that are worn in the ears of a user. The earphones 214 are canal-type earphones or inner-ear type earphones. Therefore, only the user wearing the earphones 214 can hear the sound output from the earphones 214.

[0483] The earphones 214 include a computer 36, a microphone 38, a speaker 40, a camera 42, a biosensor 43A, a vibration device 43B, and a communication I / F 44. The camera 42 is located at the same eye level as the user 20, and is therefore capable of acquiring information similar to the visual information that the user 20 obtains.

[0484] The biosensor 43A provided in the earphone 214 can acquire biometric data of the user from the user's ear. The biometric data may be, for example, the body temperature, brain waves, or heart rate of the user 20.

[0485] The vibration device 43B included in the earphone 214 vibrates in response to the received control signal, so that the earphone 214 has a vibration function and is capable of outputting haptic feedback.

[0486] Next, we will explain the processing of the specific processing unit 290 when the data processing device 12 of this embodiment performs specific processing to suggest information corresponding to the content of the user 20's utterance when it receives an utterance from the user 20 wearing the earphones 214 regarding the user's memory or behavior.

[0487] (Identification Process) In the identification process of this embodiment, user data is input and identification is performed using a data generation model that generates a predetermined inference result according to the input user data. In the identification process of the 10B embodiment, the user data includes sound detected by the microphone 38, an image captured by the camera 42, and biometric data acquired by the biometric sensor 43A. In the identification process, when an utterance related to the memory or behavior of the user 20 is received as user data from the user 20 wearing the earphones 214, a process of suggesting information corresponding to the content of the utterance to the user 20 by referring to the database 24 is executed. Specifically, after a life log is recorded in the database 24, when the user 20 wearing the earphones 214 makes an utterance related to the memory or behavior of the user 20, the identification process includes a process of suggesting information corresponding to the content of the utterance to the user 20 by referring to the database 24. Furthermore, in the 10B embodiment, the identification process includes outputting haptic feedback from the earphones 214.

[0488] The specification processing unit 290 of the data processing device 12 in this embodiment performs specification processing using the data generation model 58.

[0489] The input unit 291 in this embodiment acquires user data received by the earphone 214. Specifically, the input unit 291 acquires, as user data, the user's voice detected by the microphone 38 of the earphone 214, the image captured by the camera 42, and the biometric data acquired by the biometric sensor 43A.

[0490] The processing unit 292 in this embodiment performs identification processing using the data generation model 58. Specifically, when the processing unit 292 in this embodiment receives an utterance related to the memory or behavior of the user 20 from the user 20 wearing the earphones 214, the processing unit 292 suggests information corresponding to the content of the utterance to the user 20 and inputs a prompt to the data generation model 58 instructing the earphones 214 to generate haptic feedback, thereby performing the identification processing to acquire the information and haptic feedback corresponding to the content of the utterance.

[0491] More specifically, in this embodiment, the processing unit 292 obtains haptic feedback output from the data generation model 58 by inputting a prompt to the data generation model 58 instructing the data generation model 58 to generate haptic feedback through the earphones 214. For example, the prompt may include a message such as "Please generate haptic feedback (e.g., a vibration pattern) that is appropriate for the input user data (audio, image, or biological data)." This causes the data generation model 58 to output haptic feedback that matches the user data.

[0492] The processing unit 292 in this embodiment may obtain haptic feedback through a process different from the above process. For example, the processing unit 292 in this embodiment analyzes video data, which is a series of images acquired from the camera 42, in real time to extract color, shape, and movement features. Next, the processing unit 292 in the tenth embodiment inputs the extracted features into the data generation model 58 and generates music, sound effects, etc. according to the features and biometric data. Note that the processing unit 292 may input the image itself and biometric data into the data generation model 58 without extracting features from the image to generate music, sound effects, etc. Then, the processing unit 292 in this embodiment generates a vibration pattern corresponding to the haptic feedback based on the generated music, sound effects, etc.

[0493] The output unit 293 transmits the result of the specific processing to the earphone 214. In the earphone 214, the control unit 46A causes the speaker 40 to output the result of the specific processing. Specifically, the speaker 40 plays the generated music or sound effect. Furthermore, the vibration device 43B outputs haptic feedback using a vibration function. Specifically, the vibration device 43B outputs haptic feedback synchronized with the music or sound effect. The microphone 38 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0494] (First Example of Identification Processing) <Use in an Art Museum> When a user 20 wearing earphones 214 is viewing a painting in an art museum, the camera 42 captures an image of the painting, and the color, shape, etc. of the image are analyzed. The identification processing unit 290 in the 10B embodiment generates classical music or unique sound effects, etc., based on the information and outputs them from the speaker 40. Furthermore, the vibration device 43B provides haptic feedback in the form of minute vibrations in sync with the generated music. This allows the user 20 to feel the atmosphere of the artwork, including the painting, with their entire body.

[0495] (Second Example of Specific Processing) <Appreciating Natural Scenery> When the user 20 wearing the earphones 214 is strolling in nature, the camera 42 captures the colors and movements of the scenery. The specific processing unit 290 in this embodiment generates refreshing music or natural sounds, etc., and outputs them from the speaker 40. Furthermore, the vibration device 43B provides haptic feedback in the form of minute vibrations in sync with the generated music. This provides the user 20 with a rich experience through both hearing and touch.

[0496] Other system configurations and operations of the data processing system 210 according to this embodiment are the same as those of the first embodiment, and therefore detailed description thereof will be omitted.

[0497] As described above, the data processing system 210 according to this embodiment can be used not only in the entertainment field but also in a wide range of fields such as education, rehabilitation, tourism, etc. The data processing system 210 according to this embodiment allows users to experience visual information multisensorily, thereby creating new value.

[0498] In addition, the following supplementary notes are provided in relation to the above description.

[0499] (Supplementary Note 1) An apparatus comprising: an input unit that inputs user data including sound, images, and biometric data collected by two earphones that include a microphone, a speaker, a camera, and a biometric sensor and are worn on the user's ears; a processing unit that performs specific processing using a data generation model that generates a predetermined inference result according to the user data; and an output unit that reproduces the result of the specific processing from the speaker; wherein the earphones are canal-type earphones or inner-ear-type earphones, and only the user wearing the earphones can hear the sound output from the earphones; the earphones have a vibration function and are capable of outputting haptic feedback; the input unit inputs the sound detected by the microphone, the image captured by the camera, and the biometric data acquired by the biometric sensor as the user data; When the processing unit receives an utterance from the user wearing the earphones regarding the user's memory or behavior, it suggests information corresponding to the content of the utterance to the user based on the user's life log in which the sounds and images associated with the user are recorded, and performs the specific processing of obtaining the information corresponding to the content of the utterance and the tactile feedback by inputting a prompt to a data generation model instructing the earphones to generate the tactile feedback.

[0500] (Supplementary Note 2) The data processing device according to Supplementary Note 1, wherein, when the user wearing the earphones requests a message that will trigger a specific memory as the content of the utterance, the processing unit instructs the user in the prompt to suggest one or more messages selected based on the life log as information corresponding to the content of the utterance to the user.

[0501] (Supplementary Note 3) The data processing device according to Supplementary Note 1 or 2, wherein when the user wearing the earphones tweets a specific matter as the content of the speech, the processing unit instructs the user in the prompt to suggest to the user, as information corresponding to the content of the speech, an action of the user recommended for the matter based on the life log.

[0502] (Supplementary Note 4) A data processing method in which a computer executes a specific process using a data generation model that receives user data including sound, images, and biometric data collected by two earphones that include a microphone, a speaker, a camera, and a biometric sensor and are worn on the user's ears, and generates a predetermined inference result according to the user data, wherein the earphones are canal-type earphones or inner-ear type earphones, and only the user wearing the earphones can hear the sound output from the earphones, the earphones have a vibration function and are capable of outputting haptic feedback, and the sound detected by the microphone, the image captured by the camera, and the biometric data acquired by the earphones are input as the user data, A data processing method executed by the computer, in which, when an utterance relating to the user's memory or behavior is received from the user wearing the earphones, the computer suggests information corresponding to the content of the utterance to the user based on the user's life log in which the sounds and images associated with the user are recorded, and a prompt instructing the earphones to generate the haptic feedback is input into a data generation model, thereby executing the specific processing to obtain the information corresponding to the content of the utterance and the haptic feedback; and the computer plays back the results of the specific processing from the speaker and outputs haptic feedback using the vibration function.

[0503] (Supplementary Note 5) A data processing program that causes a computer to execute specific processing using a data generation model that receives user data including sound, images, and biometric data collected by two earphones that include a microphone, a speaker, a camera, and a biometric sensor and are worn on a user's ears, and generates a predetermined inference result according to the user data, wherein the earphones are canal-type earphones or inner-ear type earphones, and only the user wearing the earphones can hear the sound output from the earphones, the earphones have a vibration function and are capable of outputting haptic feedback, and receives as the user data the sound detected by the microphone, the image captured by the camera, and the biometric data acquired by the earphones, A data processing program that causes the computer to execute, when an utterance related to the user's memory or behavior is received from the user wearing the earphones, a process of suggesting information corresponding to the content of the utterance to the user based on the user's life log in which the sounds and images associated with the user are recorded, and acquiring the information corresponding to the content of the utterance and the haptic feedback by inputting a prompt that instructs the earphones to generate the haptic feedback into a data generation model, as the specific process; and a process of playing the results of the specific process from the speaker and outputting haptic feedback using the vibration function.

[0504] Eleventh Embodiment An example of a data processing device, a data processing method, and a program according to the technology of the present disclosure will be described below with reference to the accompanying drawings. Note that the same parts as those in the first embodiment are designated by the same reference numerals, and detailed description thereof will be omitted.

[0505] An eleventh embodiment will be described below. In the eleventh embodiment, specific processing such as dialogue executed between a user 20 and AI earphones 14 (hereinafter simply referred to as earphones 14) worn by the user 20 is the main focus, but while the first embodiment is specific processing between a single user 20 and the earphones 14, the eleventh embodiment is characterized in that specific processing is executed in a situation in which two (or more) users 20A, 20B wear a pair of earphones (14A, 14B) in a data processing system 10A (see FIG. 39 ).

[0506] In the eleventh embodiment, the same configuration as that of the first embodiment (particularly, the data processing device 12 capable of communicating with the earphones 14A and 14B) will not be described.

[0507] 39 , in the eleventh embodiment, a data processing system 10A includes a data processing device 12, an earphone 14A worn by a user 20A, and an earphone 14A worn by a user 20B. The earphones 14A, 14A are each (single item) described as a pair of earphones 14 in the eleventh embodiment, and each of them operates (is controlled) independently.

[0508] That is, as shown in FIG. 41, a pair of earphones 14A, 14B according to the eleventh embodiment have the same structure, and as shown in FIG. 42, it is assumed that they will be worn by two users 20A, 20B, respectively.

[0509] The earphones 14A and 14B may be interpreted as canal-type earphones that are fitted into the ear canals of the users 20A and 20B, respectively. However, the earphones 14A and 14B are not limited to canal-type earphones, and may be inner-ear-type earphones that are fitted by being inserted into the inner ear of the user 20, or headphone-type earphones that cover the entire ear of the user 20.

[0510] In the eleventh embodiment, since there are two users 20A and 20B, the pair of earphones 14A and 14B of the first embodiment are applied, but additional earphones with the same functions may be added according to the number of target users.

[0511] 39 , the earphones 14A and 14B each include a computer 36, a microphone 38, a speaker 40, a camera 42, a biometric sensor 43, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 38, the speaker 40, and the camera 42 are also connected to the bus 52.

[0512] The functions of the devices are the same as those of the first embodiment described with reference to FIG. 1, and therefore detailed description thereof will be omitted here.

[0513] As shown in Figure 40, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56A is stored in the storage 32. The specific process program 56A is an example of a "data processing program" according to the technology of the present disclosure. The processor 28 reads the specific process program 56A from the storage 32 and executes the read specific process program 56A on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290A in accordance with the specific process program 56A executed on the RAM 30.

[0514] 40 , in the earphones 14A and 14B, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0515] The earphones 14A and 14B are capable of communicating with the data processing device 12, and data transmission and reception between the earphones 14A and 14B are performed via the data processing device 22.

[0516] As shown in FIG. 43, a specific processing unit 290A of the data processing device 12 (see FIG. 40) includes an input unit 291A, a processing unit 292A, and an output unit 293A2.

[0517] (Input Unit 291A) The input unit 291A of the earphones acquires the following data. (Data 1) Biometric information: The biometric sensor 43 provided in the earphones 14A and 14B acquires heart rate, electrical skin response, body temperature, etc. Brain waves, sweat rate, etc. may be acquired in cooperation with a device separate from the biometric sensor 43 attached to the earphones 14A and 14B. (Data 2) Audio data: The microphone 38 provided in the earphones 14A and 14B acquires the content of speech, tone of voice, speaking speed, etc. of the users 20A and 20B. (Data 3) Video data: The camera 42 provided in the earphones 14A and 14B acquires the facial expressions and eye movements of the users 20A and 20B.

[0518] (Processing Unit 292A) The processing unit 292A performs the following processes using the data generation model. (Process 1) Data preprocessing: Performs noise removal and data synchronization. (Process 2) Emotional state analysis: Analyzes biometric information, audio, and video data to estimate the user's emotional state. At this time, the user's emotional state is analyzed by inputting prompts to the data generation model instructing the data generation model to estimate the user's emotional state for each of the biometric information, audio, and video data. (Process 3) Emotional data integration: Integrates each piece of data and quantifies the overall emotional state. (Process 4) Data encryption and transmission: Encrypts and transmits emotional data to the other user. (Process 5) Receiving and analyzing other party's data: Decrypts and analyzes data from the other party to identify the other party's emotional state. (Process 6) Feedback generation: Generates appropriate feedback according to the other party's emotional state.

[0519] In addition, the user's emotional state may be analyzed by integrating (Process 2) and (Process 3) and inputting a prompt to the data generation model that instructs the data generation model to integrate biometric information, audio data, and video data to perform a multidimensional analysis of the user's emotional state.

[0520] (Output unit 293A) The output unit 293A provides feedback using the following means: (Means 1) Vibration feedback: Controls a vibration motor to tactilely transmit the emotional state. (Means 2) Voice message: Transmits the other person's emotional state by voice from the speaker 50. (Means 3) Visual feedback: Changes the color and blinking pattern of an LED light to express emotions.

[0521] At this time, at least one of a vibration pattern, a voice message, and visual feedback is generated by inputting a prompt to the data generation model to instruct the generation of a vibration pattern, a voice message, and visual feedback according to the emotional state of the other user.

[0522] The operation of the eleventh embodiment will be described below. Figures 44(A) and 44(B) are control flowcharts showing the operation of the emotion sharing communication function according to the eleventh embodiment.

[0523] 44A shows the data collection process in either the earphone 14A or 14B (here, the earphone 14A). In step 301A, for example, data collection is performed in the earphone 14A. Data collection includes the acquisition of biometric information, audio, and video data. For example, it is assumed that the earphone 14A collects data.

[0524] In the next step 302A, emotion analysis is performed. The emotion analysis is performed by the earphone 14A, where the processing unit 292A of the specific processing unit 290A analyzes the data and estimates the emotional state.

[0525] In the next step 303A, data transmission is executed. Data transmission means that emotion data is encrypted and transmitted to the other user. For example, the data is transmitted from earphone 14A to earphone 14B.

[0526] The process on the earphone 14B side that receives the data transmission in step 303A will be described with reference to FIG. 44(B).

[0527] In step 304A, data reception and analysis is performed. Data reception and analysis refers to receiving and analyzing the data received from the data transmitted in step 303A. For example, data transmitted from earphone 14A is received by earphone 14B and collected.

[0528] In the next step 305A, feedback generation is performed. Feedback generation refers to generating feedback according to the emotional state of the other party. For example, the feedback is generated by the earphone 14A.

[0529] In the next step 306A, feedback output is performed. Feedback output is feedback by vibration, sound, or visual means, for example, from earphone 14B to earphone 14A.

[0530] The processes of steps 301A to 303A in Fig. 44A and steps 304A to 306A in Fig. 44B are repeatedly executed at regular intervals. The allocation of the earphones 14A and 14B to Fig. 44A and Fig. 44B is determined based on a correspondence relationship between the earphones 14A and 14B and a question, such as a voice question, uttered by either user 20A or 20B.

[0531] For example, when user 20A (earphone 14A) utters, "I wonder how Mr. / Ms. XX (the other person) is feeling (emotional state)?", the earphone 14B of user 20B executes the process of FIG. 44(A), and the earphone 14A of user 20A executes the process of FIG. 44(B).

[0532] In the eleventh embodiment, data exchange between the pair of earphones 14A, 14B is performed via the data processing device 22, but it is possible to exchange data directly between the earphones 14A, 14B by providing each of the earphones 14A, 14B with the function of the specific processing unit 290A of the data processing device 22. In the eleventh embodiment, the pair of earphones 14A, 14B (so-called stereo earphones) are worn by the respective users 20A, 20B, but two sets of earphones (stereo earphones) may also be worn by the respective users 20A, 20B.

[0533] Example of Eleventh Embodiment For example, if user 20A is relaxed and user 20B is feeling stressed, earphone 14A of user 20A outputs gentle vibrations and a warm voice message to reassure user 20B.

[0534] Additionally, the earphone 14B of the user 20B receives the relaxed state of the user 20A and outputs gentle feedback.

[0535] According to the eleventh embodiment, emotions can be shared between users 20A and 20B in real time, and the psychological distance can be reduced even when the users are far away. Furthermore, by combining analysis of biometric information with AI technology, the emotional state of the user can be estimated with high accuracy and appropriate feedback can be provided.

[0536] The technology of the present disclosure is widely applicable in fields such as communication devices, wearable devices, healthcare, and mental care.

[0537] In addition, the following supplementary notes are provided in relation to the above description.

[0538] (Supplementary Note 1) A data processing device comprising: an input unit that inputs user data collected by earphones including a microphone, a speaker, a biometric sensor, and a camera; a processing unit that performs specific processing using a data generation model that analyzes the emotional state of a user based on the user data; an output unit that transmits the result of the specific processing to a partner user and outputs feedback according to the emotional state of the partner user; and a communication unit that wirelessly communicates with the partner user's earphones, wherein the processing unit further generates at least one of a vibration pattern, a voice message, and visual feedback according to the emotional state of the partner user, and the output unit outputs the generated at least one of the vibration pattern, the voice message, and the visual feedback.

[0539] (Supplementary Note 2) The data processing device according to Supplementary Note 1, wherein the processing unit generates at least one of a vibration pattern, a voice message, and visual feedback by inputting a prompt to a data generation model instructing the processing unit to generate at least one of the vibration pattern, the voice message, and visual feedback according to the emotional state of the other user.

[0540] (Supplementary Note 3) The data processing device according to Supplementary Note 1, wherein the user data includes biometric information, audio data, and video data, and the processing unit analyzes the emotional state of the user by inputting a prompt to the data generation model that instructs the data generation model to integrate the biometric information, audio data, and video data to perform a multidimensional analysis of the emotional state of the user.

[0541] (Supplementary Note 4) A data processing program that causes a computer to operate as the data processing device according to any one of Supplementary Notes 1 to 3.

[0542] Disclosure of Japanese Patent Application No. 2024-107472 filed on July 3, 2024, Disclosure of Japanese Patent Application No. 2024-107124 filed on July 3, 2024, Disclosure of Japanese Patent Application No. 2024-203258 filed on November 21, 2024, Disclosure of Japanese Patent Application No. 2024-204256 filed on November 22, 2024, Disclosure of Japanese Patent Application No. 2024-204981 filed on November 25, 2024, Disclosure of Japanese Patent Application No. 2024-209 filed on December 3, 2024 The disclosures of Japanese Patent Application No. 2024-210587 filed on December 3, 2024, Japanese Patent Application No. 2024-210673 filed on December 3, 2024, Japanese Patent Application No. 2024-210674 filed on December 3, 2024, Japanese Patent Application No. 2024-232987 filed on December 27, 2024, and Japanese Patent Application No. 2025-002468 filed on January 7, 2025 are incorporated herein by reference in their entirety.

Claims

1. A data processing device comprising: an input unit that inputs user data including sounds and images collected by two earphones that include a microphone, a speaker, and a camera and are worn on the user's ears; a processing unit that performs specific processing using a data generation model that generates a predetermined inference result according to the user data; and an output unit that plays the results of the specific processing from the speaker, wherein the input unit inputs sounds detected by the microphone and images taken by the camera as the user data, and when the processing unit receives an utterance from the user wearing the earphones regarding the user's memory or behavior, it performs the specific processing by suggesting information to the user corresponding to the content of the utterance based on the user's life log in which the sounds and images linked to the user are recorded.

2. The data processing device of claim 1, wherein when the user wearing the earphones requests a message that will trigger a specific memory as the content of the utterance, the processing unit suggests to the user one or more messages selected based on the life log as information corresponding to the content of the utterance.

3. The data processing device of claim 1, wherein when the user wearing the earphones tweets a specific matter as the content of the speech, the processing unit suggests to the user recommended actions for the user regarding the matter based on the life log as information corresponding to the content of the speech.

4. A data processing method in which a computer executes a specific process using a data generation model that inputs user data including sounds and images collected by two earphones that include a microphone, a speaker, and a camera and are worn on the user's ears, and generates a predetermined inference result according to the user data, wherein the sound detected by the microphone and the image captured by the camera are input as the user data, and when an utterance regarding the user's memory or behavior is received from the user wearing the earphones, the specific process executes a process of suggesting information corresponding to the content of the utterance to the user based on the user's life log in which the sounds and images associated with the user are recorded, and the results of the specific process are played back from the speaker.

5. A data processing program that causes a computer to execute specific processing using a data generation model that inputs user data including sounds and images collected by two earphones that include a microphone, a speaker, and a camera and are worn on the user's ears, and generates a predetermined inference result according to the user data, wherein the program inputs sounds detected by the microphone and images taken by the camera as the user data, and when an utterance regarding the user's memory or behavior is received from the user wearing the earphones, the program executes as the specific processing a process of suggesting information corresponding to the content of the utterance to the user based on the user's life log in which the sounds and images associated with the user are recorded, and causes the computer to execute a process of playing back the results of the specific processing from the speakers.

6. A necklace-type terminal including: a camera that photographs the wearer's surroundings; a sensor that detects biometric data of the wearer; a microphone; and a collection unit that collects the outputs of the camera, the sensor, and the microphone.

7. The necklace-type terminal according to claim 6, further comprising a communication unit that transmits the outputs of the camera, the sensor, and the microphone collected by the collection unit to a data processing device.

8. The necklace-type terminal according to claim 7, further comprising a speaker that outputs a response corresponding to a user's utterance picked up by said microphone.

9. A data processing system comprising the necklace type terminal of claim 8 and a data processing device, wherein the data processing device comprises: an input unit that accepts user utterances picked up by the microphone; a processing unit that inputs prompts including the user utterances into a data generation model and obtains a response to the user utterances using the output of the data generation model; and an output unit that outputs the obtained response to the necklace type terminal.

10. The data processing system of claim 9, wherein the processing unit, when the user utterance satisfies a predetermined trigger condition, inputs a prompt including the user utterance into the data generation model and obtains a response to the user utterance using the output of the data generation model.

11. A data processing device comprising: an input unit including a microphone, a speaker, and a camera, which inputs voice data and image data collected by two earphones worn on the user's ears; a processing unit which performs a specific processing using a data generation model which generates a predetermined inference result according to the voice data and the image data; and an output unit which plays back the result of the specific processing from the speaker, wherein the processing unit identifies an object that a user wearing the earphones is paying attention to by analyzing the image data, and performs the specific processing by emphasizing the sound from the object in the voice data.

12. The data processing device according to claim 11, wherein the processing unit acquires environmental information and motion information of a user, and adjusts parameters of the specific process in real time based on the environmental information and the motion information using the data generation model.

13. The data processing device according to claim 11 or 12, wherein the processing unit uses the data generation model to optimize the voice to suit the user's preferences based on the user's past gaze patterns and voice history.

14. A data processing method in which a computer executes a specific processing using a data generation model that inputs audio data and image data collected by two earphones that include a microphone, a speaker, and a camera and are worn on the user's ears, and generates a predetermined inference result according to the audio data and the image data, the data processing method comprising: inputting audio data detected by the microphone and image data captured by the camera; identifying an object that the user wearing the earphones is paying attention to by analyzing the image data; executing, as the specific processing, a process that emphasizes the sound from the object in the audio data; and playing back the result of the specific processing from the speaker.

15. A data processing program that causes a computer to execute specific processing using a data generation model that inputs voice data and image data collected by two earphones that include a microphone, a speaker, and a camera and are worn on the user's ears, and generates a predetermined inference result according to the voice data and the image data, the data processing program causing a computer to execute the following steps: inputting voice data detected by the microphone and image data captured by the camera; identifying an object that a user wearing the earphones is paying attention to by analyzing the image data, and executing processing to emphasize the sound from the object in the voice data as the specific processing; and playing back the results of the specific processing from the speaker.

16. A data processing device comprising a processor that: collects user data output from at least one of a microphone, a camera, and a sensor provided in at least one of two earphones worn on a user's left and right ears; obtains two pieces of audio data having different contents based on the user data; and transmits one of the two pieces of audio data to one of the two earphones and the other of the two pieces of audio data to the other of the two earphones.

17. The data processing device according to claim 16, wherein the processor generates two prompts based on the user data, and obtains the two pieces of audio data using the two prompts.

18. The data processing device according to claim 16, wherein the processor issues a command to adjust the volume of at least one of the two earphones in response to the two pieces of audio data.

19. The data processing device according to claim 16, wherein the processor collects at least biometric data of the user as the user data.

20. A data processing method comprising: a computer collecting user data output from at least one of a microphone, a camera, and a sensor provided in at least one of two earphones worn on a user's left and right ears; obtaining two pieces of audio data having different contents based on the user data; and transmitting one of the two pieces of audio data to one of the two earphones and transmitting the other of the two pieces of audio data to the other of the two earphones.

21. A data processing program that causes a computer to collect user data output from at least one of a microphone, camera, and sensor provided in at least one of two earphones worn on a user's left and right ears, respectively; obtain two pieces of audio data having different contents based on the user data; and transmit one of the two pieces of audio data to one of the two earphones and the other of the two pieces of audio data to the other of the two earphones.

22. A data processing device comprising: an input unit that acquires user data including sounds picked up by the microphones and images taken by the cameras contained in two earphones worn on the user's left and right ears, each earphone including a microphone, speaker, and camera; a processing unit that generates the story by inputting a prompt that instructs the data generation model to generate a story according to the user data into the data generation model; and an output unit that plays the generated story from the speakers, wherein the input unit sequentially acquires the user data from the two earphones, and the processing unit repeatedly generates continuations of the story by inputting a prompt that instructs the data generation model to generate a continuation of the story according to the user data most recently acquired by the input unit, thereby generating a series of stories.

23. A data processing device according to claim 22, wherein the earphone further includes a sensor that detects biometric data or movement data of the user, and the input unit acquires the user data that further includes the biometric data or the movement data detected by the sensor.

24. A data processing device as described in claim 22, wherein the earphones further include a vibration imparting unit that imparts vibration to the user, the processing unit generates a story according to the user data and generates the story and vibration timing by inputting a prompt to the data generation model that instructs the data generation model to generate a vibration timing for vibrating in accordance with the playback of the story, and the output unit plays the generated story from the speaker and causes the vibration imparting unit to impart vibration in accordance with the generated vibration timing.

25. A data processing device as described in claim 22, wherein the processing unit further acquires the user's life log in which the sounds and images associated with the user are recorded, and generates the story by inputting a prompt to the data generation model instructing it to generate a story based on the user data and the life log.

26. A data processing method executed by a computer, comprising: acquiring user data including sounds picked up by a microphone and images taken by a camera contained in two earphones worn on the user's left and right ears, each earphone including a microphone, speaker, and camera; generating a story by inputting a prompt to a data generation model instructing the model to generate a story according to the user data; and playing the generated story from the speakers; wherein acquiring the user data includes sequentially acquiring the user data from the two earphones; and generating the story includes inputting a prompt to the data generation model instructing the model to generate a continuation of the story according to the user data acquired immediately before, thereby repeating the process of generating continuations of the story, thereby generating a series of stories.

27. A data processing program that causes a computer to execute the following steps: acquiring user data including sounds picked up by a microphone and images taken by a camera contained in two earphones worn on the user's left and right ears, each earphone including a microphone, speaker, and camera; generating a story by inputting a prompt to a data generation model that instructs the model to generate a story based on the user data; and playing the generated story from the speakers; wherein acquiring the user data includes sequentially acquiring the user data from the two earphones; and generating the story includes inputting a prompt to the data generation model that instructs the model to generate a continuation of the story based on the user data acquired immediately before, thereby repeatedly generating continuations of the story, thereby generating a series of stories.

28. A data processing device comprising: an input unit that acquires user data; a processing unit that performs specific processing using a data generation model that generates a predetermined inference result according to the user data; and an output unit that uses the result of the specific processing to play audio from speakers of two earphones, at least one of which is provided with a vibration generating unit and which are worn on the user's left and right ears, one each; wherein the input unit acquires, as the user data, learning content including learning content information that can play audio of the learning content performed by the user; the processing unit performs the processing of generating the vibration information as the specific processing by inputting, to the data generation model, a prompt based on the user data, instructing the data generation model to generate vibration information that vibrates the vibration generating unit in synchronization with the timing of the audio playback of important learning points in the learning content indicated by the learning content information; and the output unit plays the learning content indicated by the learning content information from the speakers of the earphones and vibrates the vibration generating unit using the vibration information.

29. The data processing device of claim 28, wherein the learning content further includes sound effect information capable of playing sound effects related to the learning content, and the output unit plays the learning content indicated by the learning content information from the speaker of one of the earphones, and plays sound effects indicated by the sound effect information from the speaker of the other of the earphones.

30. A data processing device as described in claim 28 or claim 29, wherein at least one of the two earphones is equipped with a sensor that detects at least one of the user's biometric data, the user's movement data, and the user's surrounding environmental data, and the input unit further acquires at least one of the biometric data, the movement data, and the environmental data from the sensor as the user data.

31. A data processing device as described in claim 28 or claim 29, wherein at least one of the two earphones is equipped with a microphone, the input unit further acquires ambient sounds of the user from the microphone as the user data, and the processing unit further performs, as the specific processing, a process of adjusting the intensity of the output target by the output unit in accordance with the ambient sounds.

32. The data processing device according to claim 31, wherein at least one of the two earphones further has a noise canceling function, and the processing unit further performs, as the specific processing, a process of adjusting the intensity of the noise canceling function in accordance with the ambient sound.

33. A data processing method in which a computer acquires user data, performs a specific processing using a data generation model that generates a predetermined inference result according to the user data, and uses the result of the specific processing to play audio from speakers of two earphones, at least one of which is provided with a vibration generating unit and which are worn on the user's left and right ears, one each, wherein the computer acquires, as the user data, learning content including learning content information that can play audio of the learning content performed by the user, inputs, based on the user data, a prompt to the data generation model that instructs the computer to generate vibration information that vibrates the vibration generating unit in synchronization with the timing of the audio playback of important learning points in the learning content indicated by the learning content information, thereby performing the processing to generate the vibration information as the specific processing, and plays the learning content indicated by the learning content information from the speakers of the earphones and vibrates the vibration generating unit using the vibration information.

34. A data processing program that causes a computer to execute the following processes: acquire user data; perform a specific process using a data generation model that generates a predetermined inference result according to the user data; and use the result of the specific process to play audio from speakers of two earphones, at least one of which is provided with a vibration generating unit and which are worn on the user's left and right ears, one each, wherein learning content including learning content information that can play audio of the learning content performed by the user is acquired as the user data; based on the user data, input a prompt to the data generation model that instructs the data generation model to generate vibration information that vibrates the vibration generating unit in synchronization with the timing of the audio playback of important learning points in the learning content indicated by the learning content information, thereby generating the vibration information as the specific process; and play the learning content indicated by the learning content information from the speaker of the earphones and vibrate the vibration generating unit using the vibration information.

35. A data processing device comprising: an input unit that inputs user data including biometric information collected by two earphones that include a microphone, a speaker, a camera, a biometric information sensor, and a vibration applying unit and are worn on the user's ears; a processing unit that performs specific processing using a data generation model that generates a predetermined inference result according to the user data; and an output unit that plays the result of the specific processing from the speaker and causes the vibration applying unit to apply vibration, wherein the input unit inputs the biometric information detected by the biometric information sensor as the user data, and the processing unit evaluates the user's condition based on the biometric information and inputs a prompt to the data generation model instructing it to generate music data and vibration pattern data according to the evaluated user's condition, thereby performing the specific processing to generate the music data and vibration pattern data as a result of the specific processing.

36. A data processing device as described in claim 35, wherein the processing unit generates the music data and vibration pattern data as a result of the specified processing by inputting a prompt to the data generation model instructing the data generation model to generate music data and vibration pattern data according to the user's state based on the user data, the user's life log in which the past biometric information, the music data, and the vibration pattern data are recorded, and the user's music preference data.

37. The data processing device according to claim 36, wherein the processing unit performs an integrated analysis of the life log and the music preference data, and personalizes the music data and the vibration pattern data in accordance with the characteristics of each individual user.

38. A data processing device according to claim 35, wherein the processing unit continuously monitors the biological information while providing music and vibrations, and performs feedback control to adjust the music data and vibration pattern data in real time in response to changes in the user's condition.

39. The data processing device according to claim 35, wherein the data generation model includes a machine learning or deep learning algorithm and is a model trained to assess the mental and physical state of the user from the biometric information.

40. A data processing method in which a computer inputs user data including biometric information collected by two earphones attached to the user's ears, each earphone comprising a microphone, a speaker, a camera, a biometric information sensor, and a vibration applying unit, and executes specific processing using a data generation model that generates a predetermined inference result corresponding to the user data, wherein the specific processing is executed by inputting biometric information detected by the biometric information sensor as the user data, evaluating the user's condition based on the biometric information, and inputting a prompt to the data generation model instructing the model to generate music data and vibration pattern data corresponding to the evaluated user's condition, thereby generating the music data and vibration pattern data as a result of the specific processing, and playing the result of the specific processing from the speaker and causing the vibration applying unit to apply vibration.

41. A data processing program that causes a computer to execute specific processing using a data generation model that inputs user data including biometric information collected by two earphones that include a microphone, speaker, camera, biometric information sensor, and vibration applying unit and are worn on the user's ears, and generates a predetermined inference result corresponding to the user data, wherein the specific processing comprises inputting biometric information detected by the biometric information sensor as the user data, evaluating the user's condition based on the biometric information, and inputting a prompt to the data generation model instructing it to generate music data and vibration pattern data corresponding to the evaluated user's condition, thereby generating the music data and vibration pattern data as a result of the specific processing, and playing the result of the specific processing from the speaker and applying vibration to the vibration applying unit.

42. An earphone comprising: a housing to be worn on the ear of a second sending user who communicates with a first receiving user; a data collection unit that collects operation data generated by detecting at least one of a swipe operation, a tap operation, and a long press operation on the housing; a processing unit that performs a specification process using a data generation model that generates a predetermined inference result according to the operation data, and generates vibration data for causing an earphone worn on the ear of the first receiving user to generate vibrations according to the result of the specification process; and an output unit that transmits the vibration data to the earphone worn on the ear of the first receiving user, wherein the processing unit executes, as the specification process, a process to infer at least one of the intention and emotion of the second sending user based on the operation data, thereby generating vibrations in the earphone worn on the ear of the first receiving user that reproduce at least one of the intention and the emotion according to the result of the specification process.

43. The earphone according to claim 42, wherein the processing unit performs the identification process taking into consideration an operation history of the second transmitting user on a housing worn on the ear of the second transmitting user.

44. The earphone according to claim 42, wherein the output section transmits the operation data by wireless communication.

45. The earphone described in claim 42, wherein the processing unit generates audio data to be played from earphones worn on the ears of the receiving first user, with at least one of the intention and the emotion as audio guidance according to the result of the specific processing while generating the vibration, and the output unit transmits the audio data via wireless communication.

46. ​​Earphones comprising: a vibration unit that vibrates a housing worn on the ear of a first receiving user; an input unit that receives operation data generated by detecting at least one of a swipe operation, a tap operation, and a long press operation performed by a second sending user communicating with the first receiving user on earphones worn on the ear of the second sending user; a processing unit that performs identification processing using a data generation model that generates a predetermined inference result according to the input operation data; and a control unit that causes the vibration unit to generate vibrations according to the result of the identification processing, wherein the processing unit executes, as the identification processing, a process of inferring at least one of the intention and emotion of the second sending user based on the operation data, and the control unit generates vibrations in the vibration unit that reproduce at least one of the intention and emotion according to the result of the identification processing.

47. The earphone according to claim 46, wherein the processing unit performs the specifying process taking into consideration an operation history of the second transmitting user on an earphone worn on the ear of the second transmitting user.

48. The earphone according to claim 46, wherein the input unit receives the operation data transmitted from an earphone worn on the ear of the transmitting-side second user via wireless communication.

49. The earphone of claim 46, wherein the control unit generates audio data for reproducing at least one of the intention and the emotion as audio guidance in accordance with the result of the specific processing while causing the vibration unit to generate vibrations.

50. A data processing device comprising: an input unit that receives operation data generated by detecting at least one of a swipe operation, a tap operation, and a long press operation performed by a second sending-side user communicating with a first receiving-side user wearing earphones on the earphones worn on the ears of the second sending-side user; a processing unit that performs identification processing using a data generation model that generates a predetermined estimation result according to the input operation data; a control unit that generates vibration data that causes the earphones worn on the ears of the first receiving-side user to generate vibrations according to the result of the identification processing; and an output unit that transmits the vibration data to the earphones worn on the ears of the first receiving-side user, wherein the processing unit executes, as the identification processing, a process of estimating at least one of the intention and emotion of the second sending-side user based on the operation data, and the control unit generates the vibration data that reproduces at least one of the intention and emotion according to the result of the identification processing.

51. A data processing device according to claim 50, wherein the processing unit performs the specification process taking into consideration an operation history of an earphone worn by the second sending user on the ear of the second sending user.

52. A data processing device according to claim 50, wherein the input unit receives the operation data transmitted from an earphone worn on the ear of the transmitting-side second user via wireless communication.

53. A data processing device as described in claim 50, wherein the control unit generates audio data that reproduces at least one of the intention and the emotion as audio guidance according to the result of the identification process while causing vibrations in earphones worn in the ears of the first receiving user, and the output unit transmits the audio guidance to the earphones worn in the ears of the first receiving user.

54. A data processing method comprising: at least one processor performing a determination process using a data generation model that generates a predetermined estimation result according to operation data generated by detecting at least one of a swipe operation, a tap operation, and a long press operation on earphones worn in the ears of a second sending-side user who is communicating with a first receiving-side user wearing earphones; generating vibration data that causes the earphones worn in the ears of the first receiving-side user to generate vibrations according to the result of the determination process; transmitting the vibration data to the earphones worn in the ears of the first receiving-side user; performing, as the determination process, a process that determines at least one of the intention and emotion of the second sending-side user based on the operation data; and generating the vibration data that reproduces at least one of the intention and emotion according to the result of the determination process.

55. A data processing program that causes at least one processor to execute processes including: performing a specification process using a data generation model that generates a predetermined estimation result according to operation data generated by detecting at least one of a swipe operation, a tap operation, and a long press operation on earphones worn in the ears of a second sending-side user who communicates with a first receiving-side user wearing earphones; generating vibration data that causes the earphones worn in the ears of the first receiving-side user to generate vibrations according to the result of the specification process; transmitting the vibration data to the earphones worn in the ears of the first receiving-side user; executing, as the specification process, a process that infers at least one of the intention and emotion of the second sending-side user based on the operation data; and generating the vibration data that reproduces at least one of the intention and emotion according to the result of the specification process.

56. A data processing device comprising: an acquisition unit that acquires video data including video of a user's surroundings; an extraction unit that extracts visual information intended for an unspecified number of people from the video data; a database that stores user profile data indicating the characteristics or attributes of the user; a selection unit that selects the visual information based on the user profile data; and a message generation unit that generates a message including the content of the selected visual information using a data generation model.

57. A data processing device according to claim 56, wherein said extraction unit extracts said visual information from an image of an information transmission medium included in said video data.

58. The data processing device according to claim 56, wherein the user profile data includes at least one of the user's hobbies, interests, behavioral history, web browsing history, application usage history, schedule, age, gender, and affiliation.

59. A data processing system comprising: a data processing device according to any one of claims 56 to 58; and earphones worn by the user, wherein the earphones include a camera that captures images of the user's surroundings and generates the video data; and a speaker that outputs the message, and the data processing device includes a communication unit that transmits audio data including the content of the message to the earphones.

60. The data processing system of claim 59, wherein the earphone includes a microphone and a vibration motor, and the communication unit transmits a control signal to the earphone together with the audio data to vibrate the vibration motor when the volume of the environmental sound input to the microphone exceeds a threshold.

61. The data processing system of claim 59, wherein the earphone includes a sensor for acquiring biometric information, and the selection unit selects the visual information based on the user's emotions estimated based on the biometric information.

62. A data processing method in which a computer acquires video data including images of a user's surroundings, extracts visual information intended for an unspecified number of people from the video data, selects the visual information based on user profile data that indicates the characteristics or attributes of the user, and generates a message including the content of the selected visual information using a data generation model.

63. A program for causing a computer to execute the following processes: acquire video data including images of the user's surroundings; extract visual information intended for an unspecified number of people from the video data; select the visual information based on user profile data indicating the characteristics or attributes of the user; and generate a message including the content of the selected visual information using a data generation model.

64. An audio device comprising: an input unit that inputs user data including sound, images, and biometric data collected by two earphones that include a microphone, a speaker, a camera, and a biometric sensor and are worn on the user's ears; a processing unit that performs specific processing using a data generation model that generates a predetermined inference result according to the user data; and an output unit that reproduces the result of the specific processing from the speaker; wherein the earphones are canal-type earphones or inner-ear type earphones, and only the user wearing the earphones can hear the sound output from the earphones; the earphones have a vibration function and are capable of outputting tactile feedback; and the input unit inputs the sound detected by the microphone, the image captured by the camera, and the biometric data acquired by the biometric sensor as the user data; When the processing unit receives an utterance from the user wearing the earphones regarding the user's memory or behavior, it suggests information corresponding to the content of the utterance to the user based on the user's life log in which the sounds and images associated with the user are recorded, and performs the specific processing of obtaining the information corresponding to the content of the utterance and the tactile feedback by inputting a prompt to a data generation model instructing the earphones to generate the tactile feedback.

65. A data processing device as described in claim 64, wherein, when the user wearing the earphones requests a message that will trigger a specific memory as the content of the utterance, the processing unit instructs the user in the prompt to suggest one or more messages selected based on the life log as information corresponding to the content of the utterance to the user.

66. A data processing device as described in claim 64, wherein, when the user wearing the earphones tweets a specific matter as the content of the speech, the processing unit instructs the user in the prompt to suggest to the user, as information corresponding to the content of the speech, actions that are recommended for the user regarding the matter based on the life log.

67. A data processing method in which a computer inputs user data including sound, images, and biometric data collected by two earphones that include a microphone, a speaker, a camera, and a biometric sensor and are worn on the user's ears, and executes specific processing using a data generation model that generates a predetermined inference result according to the user data, wherein the earphones are canal-type earphones or inner-ear type earphones, and only the user wearing the earphones can hear the sound output from the earphones, the earphones have a vibration function and are capable of outputting haptic feedback, and the sound detected by the microphone, the images captured by the camera, and the biometric data acquired by the earphones are input as the user data, A data processing method executed by the computer, in which, when an utterance relating to the user's memory or behavior is received from the user wearing the earphones, the computer suggests information corresponding to the content of the utterance to the user based on the user's life log in which the sounds and images associated with the user are recorded, and a prompt instructing the earphones to generate the haptic feedback is input into a data generation model, thereby executing the specific processing to obtain the information corresponding to the content of the utterance and the haptic feedback; and the computer plays back the results of the specific processing from the speaker and outputs haptic feedback using the vibration function.

68. A data processing program that causes a computer to execute specific processing using a data generation model that inputs user data including sound, images, and biometric data collected by two earphones that include a microphone, speaker, camera, and biometric sensor and are worn on the user's ears, and generates a predetermined inference result according to the user data, wherein the earphones are canal-type earphones or inner-ear type earphones, and only the user wearing the earphones can hear the sound output from the earphones, the earphones have a vibration function and are capable of outputting haptic feedback, and the sound detected by the microphone, the images captured by the camera, and the biometric data acquired by the earphones are input as the user data, A data processing program that causes the computer to execute, when an utterance related to the user's memory or behavior is received from the user wearing the earphones, a process of suggesting information corresponding to the content of the utterance to the user based on the user's life log in which the sounds and images associated with the user are recorded, and acquiring the information corresponding to the content of the utterance and the haptic feedback by inputting a prompt that instructs the earphones to generate the haptic feedback into a data generation model, as the specific process; and a process of playing the results of the specific process from the speaker and outputting haptic feedback using the vibration function.

69. A data processing device comprising: an input unit that inputs user data collected by earphones including a microphone, speaker, biometric sensor, and camera; a processing unit that performs specific processing using a data generation model that analyzes the user's emotional state based on the user data; an output unit that transmits the results of the specific processing to a partner user and outputs feedback according to the partner user's emotional state; and a communication unit that wirelessly communicates with the partner user's earphones, wherein the processing unit further generates at least one of a vibration pattern, a voice message, and visual feedback according to the partner user's emotional state, and the output unit outputs the generated at least one of the vibration pattern, voice message, and visual feedback.

70. A data processing device as described in claim 69, wherein the processing unit generates at least one of the vibration pattern, the voice message, and the visual feedback by inputting a prompt to a data generation model instructing the data generation model to generate at least one of the vibration pattern, the voice message, and the visual feedback in accordance with the emotional state of the other user.

71. A data processing device as described in claim 69, wherein the user data includes biometric information, audio data, and video data, and the processing unit analyzes the user's emotional state by inputting a prompt to the data generation model instructing the data generation model to integrate the biometric information, audio data, and video data to perform a multidimensional analysis of the user's emotional state.

72. A data processing program that causes a computer to operate as the data processing device according to any one of claims 69 to 71.

Citation Information

Patent Citations

  • ACTION SUPPORT DEVICE, ACTION SUPPORT METHOD, PROGRAM, AND STORAGE MEDIUM

    JP6111932B2

  • An AI assistant in glasses, a memory aid with image recognition capabilities.

    JP6808808B1

  • Attention tracking to augment focus transitions

    WO2023049746A1