Data processing device, data processing method, and data processing program

The data processing device facilitates dual audio playback by collecting and processing user data from both earphones to deliver different audio streams, allowing users to multitask effectively.

JP2026091166APending Publication Date: 2026-06-03SOFTBANK GROUP CORP

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-11-22
Publication Date
2026-06-03

AI Technical Summary

Technical Problem

Conventional technologies allow users to perform only one task at a time due to the same sound being provided from both earphones, limiting multitasking capabilities.

Method used

A data processing device that collects user data from microphones, cameras, and sensors in both earphones, generates and transmits different audio data to each earphone, and adjusts volume based on the content and importance, enabling dual audio playback.

Benefits of technology

Enables users to perform two tasks simultaneously by providing distinct audio streams to each earphone, optimizing time utilization and task management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026091166000001_ABST
    Figure 2026091166000001_ABST
Patent Text Reader

Abstract

The objective is to provide a device, method, and program that can provide audio from the left and right earphones, enabling the user to perform two tasks simultaneously. [Solution] The data processing device includes a processor, which collects user data output from at least one of a microphone, camera, and sensor provided in at least one of two earphones worn on the left and right ears of the user, and based on the user data, acquires two audio data with different content, transmits one of the two audio data to one of the two earphones, and transmits the other of the two audio data to the other of the two earphones.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] , , , ,

[0005] , , , , ,

[0001] The technology of the present disclosure relates to a data processing device, a data processing method, and a data processing program.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the conventional technology, since the same sound is provided only from the left and right earphones, there is a problem that the user can perform only one task at a time. Therefore, an object of the present disclosure is to provide a device, a method, and a program capable of providing sounds that allow the user to perform two tasks at a time from the left and right earphones.

Means for Solving the Problems

[0005] A data processing device according to a first aspect of the present disclosure includes a processor which collects user data output from at least one of a microphone, camera, and sensor provided in at least one of two earphones worn on the left and right ears of a user, and based on the user data, acquires two audio data with different content, transmits one of the two audio data to one of the two earphones, and transmits the other of the two audio data to the other of the two earphones.

[0006] A data processing device according to a second aspect of this disclosure, in a data processing device according to a first aspect, the processor generates two prompts based on the user data and acquires two audio data using the two prompts.

[0007] A data processing device according to a third aspect of this disclosure, in a data processing device according to the first or second aspect, the processor issues a command to adjust the volume of at least one of the two earphones in response to the two audio data.

[0008] A data processing device according to a fourth aspect of this disclosure is a data processing device according to any one of the first to third aspects, wherein the processor collects at least the user's biometric data as user data.

[0009] A data processing method relating to a fifth aspect of this disclosure comprises a computer collecting user data output from at least one of a microphone, camera, and sensor provided in at least one of two earphones worn on the left and right ears of a user; acquiring two audio data with different content based on the user data; transmitting one of the two audio data to one of the two earphones and transmitting the other of the two audio data to the other of the two earphones.

[0010] A data processing program according to a sixth aspect of this disclosure causes a computer to collect user data output from at least one of a microphone, camera, and sensor provided in at least one of two earphones worn on the left and right ears of a user, to acquire two audio data with different content based on the user data, to transmit one of the two audio data to one of the two earphones, and to transmit the other of the two audio data to the other of the two earphones. [Brief explanation of the drawing]

[0011] [Figure 1] Figure 1 is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] Figure 2 is a conceptual diagram showing an example of the main functions of a data processing device and earphones. [Figure 3A] Figure 3A shows an example of an earphone configuration. [Figure 3B] Figure 3B shows the user wearing earphones. [Figure 3C] Figure 3C is a diagram illustrating the field of view of camera 42. [Figure 3D] Figure 3D shows the user wearing the earphones. [Figure 3E] Figure 3E shows the user wearing the earphones. [Figure 3F] Figure 3F shows the user wearing earphones. [Figure 4] The functional configuration of a specific processing unit of a data processing device is shown in general terms. [Figure 5] This diagram outlines an example of the operation flow of a specific process performed by a data processing device. [Figure 6] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 7] This is a conceptual diagram showing an example of the operation flow of a specific process by a data processing device according to the second embodiment.

Embodiments for Carrying out the Invention

[0012] An example of an embodiment of a data processing apparatus, a data processing method, and a program according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0013] First, the terms used in the following description will be explained.

[0014] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), or an APU (Accelerated Processing Unit), etc.

[0015] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0016] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0017] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F controls communication between a plurality of computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0018] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.

[0019] <First Embodiment> FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0020] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and earphones 14. An example of the data processing device 12 is a server. In this embodiment, the data processing device 12 is an example of the "data processing device" according to the technology of the present disclosure.

[0021] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0022] The earphone 14 includes a computer 36, a microphone 38, a speaker 40, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 38, speaker 40, and camera 42 are also connected to the bus 52.

[0023] The microphone 38 receives voice signals from the user 20 and accepts instructions from the user 20. The microphone 38 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 40 outputs audio according to the instructions from the processor 46. Hereafter, the microphone 38 may be simply referred to as the microphone 38.

[0024] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0025] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0026] Figure 2 shows an example of the main functions of the data processing device 12 and the earphone 14.

[0027] As shown in Figure 2, in the data processing device 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "data processing program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0028] The storage 32 stores the data generation model 58. The data generation model 58 is used by the specific processing unit 290.

[0029] (Earphones 14) In the earphone 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0030] The earphone 14 may be interpreted as a canal-type earphone that is fitted into the ear canal of the user 20, as shown in Figure 3A. However, the earphone 14 is not limited to a canal type; it may also be an inner-ear type earphone that is inserted into the inner ear of the user 20, or a headphone type earphone that covers the entire ear of the user 20. Each of the two earphones 14 is equipped with a microphone 38, a speaker 40, and a camera 42. The sound and images collected by the two earphones 14 fitted into the ears of the user 20 may be recorded as a life log in the database 24.

[0031] The life log can be interpreted as a history of the user 20's actions in daily life, and may include sounds and images associated with the user 20, specifically sounds collected by the microphone 38 and images taken by the camera 42 during daily life. The life log may record sounds and images associated with the user 20, along with the date, time, and location in which they were acquired.

[0032] The sounds collected by the microphone 38 may include the voice of the person the user 20 is talking to, and sounds that occur around the user 20 while walking or cycling (such as the sound of cars driving, birds chirping, the babbling of a stream, and the sound of trees swaying in the wind).

[0033] As shown in Figure 3C, the camera 42 may capture images of the scenery within its field of view that is in front of the user 20, or it may capture images of scenery within its field of view that is not in front of the user 20, for example, to the side, behind, below, or above the user 20. The images captured by the camera 42 may include images of the person the user 20 is talking to, the scenery around the user 20 when they are walking or cycling, and images of the pet the user 20 is walking with.

[0034] Since each of the two earphones 14 is equipped with a camera 42, the two earphones 14 worn on the user's ears 20 are positioned at a specific distance apart, one on the left ear and the other on the right ear, as shown in Figure 3B. Therefore, compared to cases where two cameras are arranged side by side in a single housing, such as in a video camera, the spacing between the two cameras 42 can be increased, making 3D sensing easier. 3D sensing can be interpreted as measuring three-dimensional shapes.

[0035] Furthermore, when the two earphones 14 are placed in the user 20's ears, the two cameras 42 are positioned close to the user 20's left and right eyes, allowing images (captured images) that are nearly identical to those seen with the naked eye to be recorded as a life log in the database 24. Consequently, in specific processing, it becomes easier to reproduce information corresponding to inquiries from the user 20, that is, information corresponding to the content of the user 20's speech.

[0036] While the two earphones 14 are attached to the user 20, all or part of the images captured by the camera 42 may be recorded in the database 24 as a life log. Specifically, when the two earphones 14 are attached to the user 20, the recording of images captured by the camera 42 to the database 24 may begin, and when the two earphones 14 are removed from the user 20, the recording of those images to the database 24 may end.

[0037] While the two earphones 14 are worn by the user 20, all or part of the sound collected by the microphone 38 may be recorded as a lifelog in the database 24. Specifically, when the two earphones 14 are worn by the user 20, the recording of the sound collected by the microphone 38 to the database 24 may begin, and when the two earphones 14 are removed from the user 20, the recording of the sound to the database 24 may end.

[0038] Next, we will describe the processing of the specific processing unit 290 when the data processing device 12 receives an utterance from the user 20 wearing the earphones 14 regarding the user 20's memories or actions, and performs specific processing to propose information corresponding to the content of the user 20's utterance to the user 20.

[0039] (Specific processing) In this embodiment, the specific processing involves inputting user data and performing specific processing using a data generation model that generates predetermined inference results corresponding to the input user data. Specifically, in the specific processing, when utterances related to the user's memories or actions are received as user data from a user 20 wearing earphones 14, the system refers to the database 24 and performs processing to propose information corresponding to the content of the utterances to the user 20. Specifically, after a life log is recorded in the database 24, if the user 20 wearing earphones 14 makes an utterance related to the user's memories or actions, the specific processing may involve referring to the database 24 and proposing information corresponding to the content of the utterances to the user 20.

[0040] (Example of specific processing) If the user wearing the earphones requests a message that will trigger the recall of a specific memory, the specific processing unit 290 may propose one or more messages selected based on the life log to the user who made the request, as information corresponding to the content of the utterance (request).

[0041] For example, if user 20, wearing earphones 14, tries to recall their memory and asks, "What did I say to person A around [date] at [time]?", the identification processing unit 290, as part of its identification process, inputs this message as a prompt to the data generation model 58. The identification processing unit 290 may refer to the life log in database 24 and, based on the output obtained from the data generation model 58, generate a message such as, "I think you said, 'I found a nice restaurant, let's make a reservation.'" This message may be interpreted as an example of information corresponding to the content of user 20's utterance.

[0042] For example, if user 20 wearing earphones 14 tries to recall their memory and asks, "Who was I talking to around [date] at [time]?", the identification processing unit 290 will input this message as a prompt to the data generation model 58 as part of its identification process. The identification processing unit 290 may refer to the life log in database 24 and, based on the output obtained from the data generation model 58, generate a message such as, "It seems you were talking with two friends at that time, probably B and C." This message may be interpreted as an example of information corresponding to the content of user 20's utterance.

[0043] For example, if user 20, wearing earphones 14, tries to recall their emotions and says, "How did I feel when I was talking to person A around [date] at [time]?", the identification processing unit 290, as part of its identification process, inputs this message as a prompt to the data generation model 58. The identification processing unit 290 may refer to the life log in database 24 and, based on the output obtained from the data generation model 58, generate a message such as, "At that time, you were laughing a lot, so it seems you had a good impression of your friend and were very happy." This message may be interpreted as an example of information corresponding to the content of user 20's utterance.

[0044] (Example of specific processing, part 2) If a user 20 wearing earphones 14 mutters a specific matter as part of their utterance, the specific processing unit 290 may suggest to the user 20 who requested the message, based on their life log, recommended actions for the user 20 regarding that matter, as information corresponding to the content of their utterance (muttering).

[0045] For example, when user 20 wearing earphones 14 is shopping at a specific retail store and says, "What should I buy?", the specific processing unit 290 inputs this message as a prompt to the data generation model 58 as a specific processing step. The specific processing unit 290 may refer to the life log in the database 24 and, based on the output obtained by the data generation model 58, generate a message such as, "A few months ago, you purchased product A at this store and commented that it wasn't very tasty, so how about purchasing recently released products B and C this time?" This message may be interpreted as an example of information corresponding to the content of user 20's utterance.

[0046] (Third example of specific processing) As shown in Figure 3D, when user 20, wearing earphones 14, is operating a PC and says, "What was the name of product A that I searched for the day before yesterday?", the identification processing unit 290 inputs this message as a prompt to the data generation model 58 as part of its identification processing. The data generation model 58 refers to the life log in the database 24 and analyzes the video of the PC screen when user 20 was operating it in the past to generate a specific output. Based on the output obtained by the data generation model 58, the identification processing unit 290 may generate a message such as "Product A is ○○○". This message may be interpreted as an example of information corresponding to the content of user 20's utterance.

[0047] (Fourth example of specific processing) As shown in Figure 3E, if user 20, wearing earphones 14, says "There was a place nearby with a great view, but I wonder where it is?" while cycling, the identification processing unit 290 inputs this message as a prompt to the data generation model 58 as part of its identification process. The data generation model 58 refers to the life log in database 24 and analyzes places previously visited by user 20 and the route to those places to generate a specific output. Based on the output obtained by the data generation model 58, the identification processing unit 290 may generate a message such as "I think it's Cape XX, about 500m from here." This message can be interpreted as an example of information corresponding to the content of user 20's utterance.

[0048] (Example 5 of specific processing) As shown in Figure 3F, when user 20, wearing earphones 14, meets Mr. X at company A, the company he is visiting, and says, "Can you tell me this person's name?", the identification processing unit 290 inputs this message as a prompt to the data generation model 58 as part of the identification process. The data generation model 58 refers to the life log in database 24 and generates specific output from the history of people that user 20 met when he visited company A. Based on the output obtained from the data generation model 58, the identification processing unit 290 may generate a message such as, "I think his name is ○○." This message may be interpreted as an example of information corresponding to the content of user 20's utterance.

[0049] As shown in Figure 4, the specific processing unit 290 includes an input unit 291, a processing unit 292, and an output unit 293.

[0050] The input unit 291 acquires user input received through the earphone 14. Specifically, it acquires the user's voice received through the earphone 14.

[0051] The processing unit 292 performs specific processing using the data generation model 58. Specifically, it inputs voice from the user into the data generation model 58 and obtains a generation result. More specifically, when it receives an utterance from the user 20 wearing the earphones 14 regarding the user 20's memories or actions, it performs a specific processing step of proposing information corresponding to the content of the utterance to the user 20.

[0052] The output unit 293 transmits the result of the specific processing to the earphone 14. In the earphone 14, the control unit 46A causes the speaker 40 to output the result of the specific processing. The microphone 38 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0053] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0054] Next, the operation of the data processing system 10 will be explained.

[0055] An example of the flow of a specific processing method will be explained with reference to Figure 5. Note that the flow of a specific processing method shown in Figure 5 is an example of a "data processing method" related to the technology disclosed herein.

[0056] In step S300, the data processing device 12 receives user data, including sound and images collected by the two earphones 14.

[0057] In step S302, if the data processing device 12 receives an utterance from the user wearing the earphones 14 regarding the user's memories or actions, it executes a specific process to propose information corresponding to the content of the utterance to the user 20 based on the user's life log.

[0058] In step S303, the data processing device 12 executes a process to play back the result of a specific process from the speaker 40.

[0059] Figure 6 shows an example of the configuration of the data processing system 10 according to the second embodiment. In the first embodiment, the data processing system 10 was shown as an example applied to single audio playback, which plays one audio track. However, in single audio playback, only the same audio track is provided from the left and right earphones, which means the user can only perform one task at a time. To solve this problem, the second embodiment describes a case in which the data processing system 10 is applied to dual audio playback, which plays two different audio tracks from two earphones.

[0060] The data processing system 10 according to the second embodiment includes a data processing device 12, an earphone 14L for the left ear, and an earphone 14R for the right ear. Earphones 14L and 14R are connected to the data processing device 12 via a network 54 so as to be communicative.

[0061] In the second embodiment, two different sounds are played from the left earphone 14L and the right earphone 14R, respectively. Therefore, it is important to avoid mixing the left and right sounds. Accordingly, it is preferable that the earphones 14L and 14R are inner-ear type earphones rather than bone conduction type or open-ear type, and it is particularly preferable that they are canal-type earphones with a high degree of airtightness.

[0062] Each of the left earphone 14L and the right earphone 14R may have the same configuration as the earphone 14 according to the first embodiment. That is, the left earphone 14L may include a computer 36L, a microphone 38L, a speaker 40L, a camera 42L, and a communication I / F 44L. The computer 36L may also include a processor 46L, RAM 48L, and storage 50L. In addition, the left earphone 14L may include a sensor 45L. These components may be connected to a bus 52L and be able to communicate with each other.

[0063] Similarly, the right earphone 14R may include a computer 36R, a microphone 38R, a speaker 40R, a camera 42R, and a communication interface 44R. The computer 36R may also include a processor 46R, RAM 48R, and storage 50R. In addition, the right earphone 14R may include a sensor 45R. These components may be connected to a bus 52R and be able to communicate with each other.

[0064] Sensors 45L and 45R (collectively referred to as "Sensor 45") measure biometric data of users wearing earphones 14L and 14R. Examples of biometric data include body temperature and electroencephalogram (EEG).

[0065] In such a data processing system 10, the data processing device 12 may be configured to send and receive different data streams to and from earphones 14L and 14R, respectively. In this case, the data processing device 12 may employ a left / right independent transmission method with respect to earphones 14L and 14R, and transmit different audio data in the left and right data streams.

[0066] This allows the Earphone 14L and Earphone 14R to reproduce different audio. It should be noted that "reproducing different audio" here means reproducing separate audio from different source content, and should be interpreted as being entirely different from stereo playback, which reproduces the same content with different localizations.

[0067] Figure 7 shows an example of the operation flow of a specific process performed by the data processing device 12 according to the second embodiment. This flow may be started when the user presses a button to start dual audio playback or speaks to that effect.

[0068] In step S401, the processor 28 collects user data. For example, the processor 28 may collect user data output from at least one of the microphone 38L, camera 42L, sensor 45L, microphone 38R, camera 42R, and sensor 45R. In this case, the processor 28 may collect at least the user's biometric data as user data. For example, the processor 28 may collect user data output from at least one of the microphone, camera, and sensor provided in at least one of the two earphones worn on the left and right ears of the user.

[0069] In step S402L, the processor 28 generates a prompt for the left ear. Also, in step S402R, the processor 28 generates a prompt for the right ear.

[0070] For example, in step S401, the processor 28 detects a speech instructing the playback of two pieces of content from the output of at least one of the microphones 38L and 38R. For example, it may detect that the user has said, "Play music and schedule simultaneously?" The processor 28 also detects the user's actions from the output of at least one of the cameras 42L and 42R collected in step S401. For example, it may detect that the user is walking along the beach. The processor 28 also detects the user's state from the output of at least one of the sensors 45L and 45R collected in step S401. For example, it may detect that the user's stress level is high.

[0071] In this case, the processor 28 may generate a prompt for the left ear, "Access the calendar app and read out today's schedule," from the utterances instructing the playback of the two contents. The processor 28 may also generate a prompt for the right ear, "Play relaxing ocean-themed music," from the utterances instructing the playback of the two contents, the user's actions, and the user's state. The processor 28 may generate the two prompts based on user data in this way, for example.

[0072] In step S403L, the processor 28 acquires audio data for the left ear. Also, in step S403R, the processor 28 acquires audio data for the right ear.

[0073] For example, the processor 28 may input the prompt for the left ear generated in step S402L to the data generation model 58. In response, the processor 28 may obtain the following audio data from the data generation model 58 as audio data for the left ear: "Today's schedule consists of three points: a meeting in the first conference room from 10 to 11, lunch with Muhammad from 12 to 13, and submitting a report to Yang by 17."

[0074] Furthermore, the processor 28 may input the prompt for the right ear generated in step S402R to the data generation model 58. In response, the processor 28 may obtain music data of soothing ocean-themed music from the data generation model 58 as audio data for the right ear. Thus, the processor 28 may obtain two audio data sets using the two prompts. In this way, for example, the processor 28 can obtain two audio data sets with different content based on user data.

[0075] In step S404L, the processor 28 transmits audio data for the left ear to the left earphone 14L. Also, in step S404R, the processor 28 transmits audio data for the right ear to the right earphone 14R.

[0076] For example, the processor 28 may transmit the audio data for the left ear acquired in step S403L to the communication interface 44L via the communication interface 26 and the network 54. The processor 28 may also transmit the audio data for the right ear acquired in step S403R to the communication interface 44R via the communication interface 26 and the network 54. In this way, the processor 28 may transmit one of the two audio data to one of the two earphones and the other of the two audio data to the other of the two earphones.

[0077] As a result, the left earphone (14L) plays an audio message from speaker 40L stating, "Today's schedule consists of three points: a meeting in Conference Room 1 from 10:00 to 11:00, lunch with Muhammad from 12:00 to 13:00, and submitting a report to Yang by 17:00." The right earphone (14R) plays soothing ocean-themed music from speaker 40R.

[0078] Here, the audio data played in the right earphone 14R is for an entertainment-related task. On the other hand, the audio data played in the left earphone 14L is for a business-related task. Therefore, the audio data played in the left earphone 14L is more important than the audio data played in the right earphone 14R. So, the processor 28 then adjusts the balance of the left and right volume.

[0079] In step S405L, the processor 28 issues a command to adjust the volume for the left ear. Also, in step S405R, the processor 28 issues a command to adjust the volume for the right ear.

[0080] For example, the processor 28 issues a command to increase the volume of the left earphone 14L so that the audio data played in the left earphone 14L is louder than the audio data played in the right earphone 14R. The processor 28 also issues a command to decrease the volume of the right earphone 14R so that the audio data played in the right earphone 14R is quieter than the audio data played in the left earphone 14L. This optimizes the volume of both earphones. The processor 28 may issue commands to adjust the volume of at least one of the two earphones in accordance with the two audio data, for example in this manner.

[0081] Then, processor 28 terminates this flow. Note that the processing in steps S402L to S405L and the processing in steps S402R to S405R are executed independently of each other. Therefore, a dual AI assistant system can be provided that simultaneously activates different AI assistants in the left and right earphones.

[0082] While the above description uses music and schedule reading as an example, the technology disclosed can be applied to various combinations. Examples of such combinations include music and email reading, news reading and schedule reading, and music and navigation guidance.

[0083] As described above, the data processing device 12 according to the second embodiment collects user data output from at least one of the microphone, camera, and sensor provided in at least one of the two earphones worn on the user's left and right ears, acquires two audio data with different content based on the user data, transmits one of the two audio data to one of the two earphones and transmits the other of the two audio data to the other of the two earphones. As a result, the data processing device 12 according to the second embodiment can provide audio from both the left and right earphones that allows the user to perform two tasks at once, thus contributing to the effective use of time.

[0084] In this case, the data processing device 12 may generate two prompts based on the user data and acquire two audio data sets using the two prompts. This allows the data processing device 12 to operate the AI ​​assistant in parallel.

[0085] Furthermore, the data processing device 12 may issue a command to adjust the volume of at least one of the two earphones in response to the two audio data. This allows the data processing device 12 to optimize the left and right volume according to the importance of the content.

[0086] Furthermore, the data processing device 12 may collect at least the user's biometric data as user data. This allows the data processing device 12 to acquire and provide voice data that takes biometric data into account using only the output from the earphones, without requiring a separate means for collecting biometric data.

[0087] Furthermore, the technology according to the second embodiment can be modified or applied in various ways. Generally, it is known that the human brain has a left hemisphere that controls language and a right hemisphere that controls images. The structure of the nervous system from the human ear to the brain is in a crossed state, where information entering from the left ear is transmitted to the right hemisphere, and information entering from the right ear is transmitted to the left hemisphere.

[0088] Therefore, when generating prompts for the left and right ears, the processor 28 may also take these characteristics of the human body into consideration. That is, in order to transmit music from the left ear to the right brain, a prompt to play music may be generated as the prompt for the left ear. Similarly, in order to transmit text from the right ear to the left brain, a prompt to read aloud a schedule or email may be generated as the prompt for the right ear.

[0089] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0090] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.

[0091] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0092] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0093] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0094] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0095] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0096] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0097] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0098] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0099] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0100] Furthermore, the following additional information is disclosed regarding the above explanation.

[0101] (Note 1) An input unit that receives user data including sound and images collected by two earphones worn on the user's ears, including a microphone, speaker, and camera, A processing unit that performs specific processing using a data generation model that generates predetermined inference results according to the user data, An output unit that reproduces the result of the specified processing from the speaker, Equipped with, The input unit inputs the sound detected by the microphone and the image captured by the camera as user data. The processing unit, when it receives a speech from the user wearing the earphones, regarding the user's memories or actions, performs a process as the specific processing to propose information to the user corresponding to the content of the speech, based on the user's life log in which sounds and images associated with the user are recorded.

[0102] (Note 2) The processing unit, when the user wearing the earphones requests a message that will trigger the user to recall a specific memory as part of the utterance, proposes to the user one or more messages selected based on the life log as information corresponding to the content of the utterance, as described in Appendix 1.

[0103] (Note 3) The processing unit, when the user wearing the earphones mutters a specific matter as the content of the utterance, proposes to the user, based on the life log, recommended actions for the user in relation to the matter, as information corresponding to the content of the utterance, as described in Appendix 1 or 2.

[0104] (Note 4) A data processing method in which a computer performs specific processing using a data generation model that inputs user data including sound and images collected by two earphones worn on the user's ears, which include a microphone, speaker, and camera, and generates predetermined inference results according to the user data, The sound detected by the microphone and the image captured by the camera are input as user data. When the system receives a utterance from the user wearing the earphones, regarding the user's memories or actions, the system performs a process as the specified process to propose information corresponding to the content of the utterance to the user, based on the user's life log in which the sounds and images associated with the user are recorded. The process of playing back the result of the aforementioned specific process from the speaker is as follows: A data processing method performed by the aforementioned computer.

[0105] (Note 5) A data processing program that causes a computer to perform specific processing using a data generation model that inputs user data including sound and images collected by two earphones worn on the user's ears, including a microphone, speaker, and camera, and generates predetermined inference results according to the user data, The sound detected by the microphone and the image captured by the camera are input as user data. When the system receives a utterance from the user wearing the earphones, regarding the user's memories or actions, the system performs a process as the specified process to propose information corresponding to the content of the utterance to the user, based on the user's life log in which the sounds and images associated with the user are recorded. The process of playing back the result of the aforementioned specific process from the speaker is as follows: A data processing program to be executed by the aforementioned computer. [Explanation of symbols]

[0106] 10 Data Processing Systems 12 Data Processing Devices 14 Earphones 290 Specific Processing Unit 291 Input section 292 Processing Unit 293 Output section< / url:>

Claims

1. The processor comprises, The system collects user data output from at least one of the microphones, cameras, and sensors located in at least one of the two earphones worn in the user's left and right ears. Based on the user data mentioned above, two audio data sets with different content are obtained. One of the two audio data is transmitted to one of the two earphones, and the other of the two audio data is transmitted to the other of the two earphones. Data processing device.

2. The aforementioned processor, Based on the user data mentioned above, two prompts are generated, The two audio data sets are acquired using the two prompts mentioned above. The data processing device according to claim 1.

3. The aforementioned processor, In response to the two audio data, a command is issued to adjust the volume of at least one of the two earphones. The data processing device according to claim 1.

4. The processor collects at least the user's biometric data as user data. The data processing device according to claim 1.

5. Computers The system collects user data output from at least one of the microphones, cameras, and sensors located in at least one of the two earphones worn in the user's left and right ears, Based on the aforementioned user data, two audio data sets with different content are obtained, The device includes transmitting one of the two audio data to one of the two earphones, and transmitting the other of the two audio data to the other of the two earphones. Data processing method.

6. On the computer, The system collects user data output from at least one of the microphones, cameras, and sensors located in at least one of the two earphones worn in the user's left and right ears. Based on the user data, two audio data sets with different content are obtained. One of the two audio data is transmitted to one of the two earphones, and the other of the two audio data is transmitted to the other of the two earphones. Data processing program.