Data processing apparatus, data processing method, and data processing program
The data processing system using earphones to collect sounds and images for a data generation model improves chatbot responses by integrating user experiences, enhancing relevance and accuracy.
Patent Information
- Application Number
- JP2024107472
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-03
- Publication Date
- 2026-01-16
AI Technical Summary
Conventional chatbot systems fail to incorporate user actions and experiences from daily life into their responses, limiting the relevance and accuracy of suggested information.
A data processing system using earphones equipped with microphones and cameras to collect sounds and images, which are processed through a data generation model to suggest information relevant to the user's memories and behaviors based on their life logs.
Enhances the relevance and accuracy of chatbot responses by incorporating user-specific experiences and actions, providing personalized and contextually appropriate suggestions.
Smart Images

Figure 2026007529000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a data processing device, a data processing method, and a data processing program. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] However, in conventional technology, when generating utterances in response to words spoken by a user, the actions taken by the user in their daily lives are not reflected, so there is room for improvement in suggesting information that corresponds to the content of the user's utterance. [Means for solving the problem]
[0005] A first aspect of the technology disclosed herein is a data processing device comprising: an input unit that inputs user data including sounds and images collected by two earphones that include a microphone, a speaker, and a camera and are worn on the user's ears; a processing unit that performs specific processing using a data generation model that generates a predetermined inference result according to the user data; and an output unit that plays the results of the specific processing from the speaker, wherein the input unit inputs sounds detected by the microphone and images captured by the camera as the user data, and when the processing unit receives an utterance from the user wearing the earphones regarding the user's memory or behavior, the processing unit performs the specific processing by suggesting information corresponding to the content of the utterance to the user based on the user's life log in which the sounds and images linked to the user are recorded.
[0006] A second aspect of the technology disclosed herein is a data processing method in which a computer inputs user data including sounds and images collected by two earphones that include a microphone, a speaker, and a camera and are worn on the user's ears, and executes a specific process using a data generation model that generates a predetermined inference result according to the user data.The data processing method includes inputting sounds detected by the microphone and images captured by the camera as the user data, and when an utterance regarding the user's memory or behavior is received from the user wearing the earphones, executing as the specific process a process of suggesting information corresponding to the content of the utterance to the user based on the user's life log in which the sounds and images linked to the user are recorded, and executing a process in which the computer plays back the results of the specific process from the speaker.
[0007] A third aspect of the technology of the present disclosure is a data processing program that causes a computer to execute a specific process using a data generation model that inputs user data including sound and images collected by two earphones that include a microphone, a speaker, and a camera and are worn on the user's ears, and generates a predetermined inference result according to the user data.The data processing program inputs sound detected by the microphone and images captured by the camera as the user data, and when an utterance regarding the user's memory or behavior is received from the user wearing the earphones, the specific process is executed to suggest information corresponding to the content of the utterance to the user based on the user's life log in which the sound and the images linked to the user are recorded, and causes the computer to execute a process to play back the results of the specific process from the speaker. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a conceptual diagram showing an example of the configuration of a data processing system. [Figure 2] FIG. 2 is a conceptual diagram showing an example of the main functions of the data processing device and the earphone. [Figure 3A] FIG. 3A is a diagram showing an example of the configuration of an earphone. [Figure 3B] FIG. 3B is a diagram showing a state in which the user is wearing the earphones. [Figure 3C] FIG. 3C is a diagram for explaining the angle of view of the camera 42. As shown in FIG. [Figure 3D] FIG. 3D shows a state in which the user is wearing the earphones. [Figure 3E] FIG. 3E shows a state in which the user is wearing the earphones. [Figure 3F] FIG. 3F shows a state in which the user is wearing the earphones. [Figure 4] 2 shows a schematic functional configuration of a specific processing unit of the data processing device. [Figure 5] 10 is a diagram illustrating an example of an operational flow of specific processing by a data processing device. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, an example of an embodiment of a data processing device, a data processing method, and a program according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0010] First, the terms used in the following description will be explained.
[0011] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), or an APU (Accelerated Processing Unit).
[0012] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0013] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0014] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0016] FIG. 1 shows an example of the configuration of a data processing system 10 according to the embodiment.
[0017] 1, a data processing system 10 includes a data processing device 12 and earphones 14. An example of the data processing device 12 is a server. In this embodiment, the data processing device 12 is an example of a "data processing device" according to the technology of the present disclosure.
[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0019] The earphones 14 include a computer 36, a microphone 38, a speaker 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 38, the speaker 40, and the camera 42 are also connected to the bus 52.
[0020] The microphone 38 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 38 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 40 outputs audio in accordance with instructions from the processor 46. Hereinafter, the microphone 38 may be simply referred to as the mic 38.
[0021] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0022] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0023] FIG. 2 shows an example of the main functions of the data processing device 12 and the earphone 14.
[0024] As shown in FIG. 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "data processing program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0025] The storage 32 stores a data generation model 58. The data generation model 58 is used by the specific processing unit 290.
[0026] (Earphone 14) In the earphones 14, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0027] As shown in Fig. 3A, the earphones 14 may be interpreted as canal-type earphones that are fitted into the ear canals of the user 20. However, the earphones 14 are not limited to canal-type earphones, and may be inner-ear-type earphones that are fitted into the inner ears of the user 20, or headphone-type earphones that cover the entire ears of the user 20. Each of the two earphones 14 is provided with a microphone 38, a speaker 40, and a camera 42. Sounds and images collected by the two earphones 14 fitted into the ears of the user 20 may be recorded in a database 24 as a life log.
[0028] The life log may be interpreted as a history of actions taken by the user 20 in daily life, and may include sounds and images associated with the user 20, specifically, sounds collected by the microphone 38 in daily life and images taken by the camera 42. The life log may record sounds and images associated with the user 20 in association with the date, time, and location at which they were acquired.
[0029] The sounds collected by the microphone 38 may include the voice of the person with whom the user 20 is talking, sounds occurring around the user 20 while walking or cycling (such as the sound of cars passing by, birds chirping, the sound of a river flowing, and the sound of trees rustling in the wind).
[0030] 3C, camera 42 may capture an image of scenery within an angle of view that captures what is in front of user 20, or may capture an image of scenery within an angle of view that captures what is not in front of user 20, for example, what is to the side, behind, below, or above user 20. The image captured by camera 42 may include the image of someone with whom user 20 is talking, the image of the scenery around user 20 when taking a walk or cycling, the image of a pet walking with user 20, etc.
[0031] Because each of the two earphones 14 is provided with a camera 42, the two earphones 14 worn on the ears of the user 20 are positioned a specific distance apart on the left and right ears, as shown in FIG. 3B. Therefore, compared to when two cameras are arranged side by side in a single housing, such as a video camera, the distance between the two cameras 42 can be made wider, making 3D sensing easier. 3D sensing can be understood as measuring a three-dimensional shape.
[0032] Furthermore, when the two earphones 14 are attached to the ears of the user 20, the two cameras 42 are positioned close to the left and right eyes of the user 20, so that images (photographed images) that are substantially the same as those seen with the naked eye can be recorded as a life log in the database 24. Therefore, in the identification process, it becomes easier to reproduce information corresponding to an inquiry from the user 20, that is, information corresponding to the content of the utterances of the user 20.
[0033] While the two earphones 14 are worn by the user 20, all or part of the images captured by the camera 42 may be recorded as a life log in the database 24. Specifically, when the two earphones 14 are worn by the user 20, recording of the images captured by the camera 42 in the database 24 may start, and when the two earphones 14 are removed from the user 20, recording of the images in the database 24 may end.
[0034] While the two earphones 14 are worn by the user 20, all or part of the sounds collected by the microphone 38 may be recorded as a life log in the database 24. Specifically, when the two earphones 14 are worn by the user 20, recording of the sounds collected by the microphone 38 in the database 24 may start, and when the two earphones 14 are removed from the user 20, recording of the sounds in the database 24 may end.
[0035] Next, we will explain the processing of the specific processing unit 290 when the data processing device 12 performs specific processing to suggest information corresponding to the content of the user 20's utterance when it receives an utterance from the user 20 wearing the earphones 14 regarding the user's memory or behavior.
[0036] (Specific processing) In the identification process of this embodiment, user data is input and an identification process is performed using a data generation model that generates a predetermined inference result according to the input user data. Specifically, in the identification process, when an utterance related to the memory or behavior of user 20 is received as user data from user 20 wearing earphones 14, a process of suggesting information corresponding to the content of the utterance to user 20 by referring to database 24 is executed. Specifically, after a life log is recorded in database 24, when user 20 wearing earphones 14 makes an utterance related to the memory or behavior of user 20, a process of suggesting information corresponding to the content of the utterance to user 20 by referring to database 24 may be executed as the identification process.
[0037] (First example of specific processing) When a user wearing earphones requests a message that will trigger a specific memory as the content of the utterance, the specific processing unit 290 may suggest one or more messages selected based on the life log to the user who requested the message as information corresponding to the content (request) of the utterance.
[0038] For example, if the user 20 wearing the earphones 14 tries to recall his / her memory and utters, "What did I say to A on a certain date at around XX time?", the identification processing unit 290 inputs the message as a prompt into the data generation model 58 as an identification process. The identification processing unit 290 may refer to the life log in the database 24 and generate a message such as, "I think he / she said, 'I found a nice restaurant, so let's make a reservation.'" based on the output obtained by the data generation model 58. The message may be interpreted as an example of information corresponding to the content of the user 20's utterance.
[0039] For example, if the user 20 wearing the earphones 14 tries to recall his / her memory and utters, "Who were you talking to at around XX time on XX date?", the identification processing unit 290 inputs the message as a prompt into the data generation model 58 as an identification process. The identification processing unit 290 may refer to the life log in the database 24 and generate a message such as, "It seems that two friends, probably Mr. B and Mr. C, were talking at that time," based on the output obtained by the data generation model 58. The message may be interpreted as an example of information corresponding to the content of the user 20's utterance.
[0040] For example, if the user 20 wearing the earphones 14 tries to recall his or her own feelings and utters, "How did I feel when I was talking to A-san around XX on XX date?", the identification processing unit 290 inputs the message as a prompt into the data generation model 58 as an identification process. The identification processing unit 290 may refer to the life log in the database 24 and generate a message such as, "You were laughing a lot at that time, so you seemed to like your friend and be very happy." based on the output obtained by the data generation model 58. The message may be interpreted as an example of information corresponding to the content of the user 20's utterance.
[0041] (Second example of specific processing) When a user 20 wearing earphones 14 tweets a specific matter as the content of the speech, the specific processing unit 290 may suggest to the user 20 who requested the message, as information corresponding to the content of the speech (tweet), the behavior of the user 20 that is recommended for the matter based on the life log.
[0042] For example, if the user 20 wearing the earphones 14 utters, "What should I buy?" while shopping at a particular retail store, the identification processing unit 290, as an identification process, inputs the message as a prompt into the data generation model 58. The identification processing unit 290 may generate a message such as, "A few months ago, after purchasing product A at this store, you commented that it wasn't very tasty, so how about purchasing product B or product C, which were recently released, this time," based on the output obtained by the data generation model 58 by referring to the life log in the database 24. The message may be interpreted as an example of information corresponding to the content of the user 20's utterance.
[0043] (Third example of specific processing) As shown in FIG. 3D , when a user 20 wearing earphones 14 utters, "What was the name of product A I searched for the day before yesterday?" while operating a personal computer, the identification processing unit 290 inputs the message as a prompt into the data generation model 58 as an identification process. The data generation model 58 generates a specific output by referencing the life log in the database 24 and analyzing images of the personal computer screen when the user 20 was operating the computer in the past. The identification processing unit 290 may generate a message such as "Product A is XXX" based on the output obtained by the data generation model 58. The message may be interpreted as an example of information corresponding to the content of the user 20's utterance.
[0044] (Fourth example of specific processing) As shown in FIG. 3E, if a user 20 wearing earphones 14 utters, while cycling, "There's a place nearby with a spectacular view. Where is it?", the identification processing unit 290 inputs the message as a prompt to the data generation model 58 as an identification process. The data generation model 58 generates a specific output by referencing the life log in the database 24 and analyzing places the user 20 has previously visited and the route to those places. Based on the output obtained by the data generation model 58, the identification processing unit 290 may generate a message such as, "I think Cape XX is 500 meters from here." This message may be interpreted as an example of information corresponding to the content of the user 20's utterance.
[0045] (Fifth example of specific processing) As shown in FIG. 3F , when the user 20 wearing the earphones 14 meets Mr. X from Company A while visiting and utters, "What's his name?", the identification processing unit 290 inputs the message as a prompt into the data generation model 58 as an identification process. The data generation model 58 references the life log in the database 24 and generates a specific output based on the history of people the user 20 met while visiting Company A. The identification processing unit 290 may generate a message such as, "I think his name is XX" based on the output obtained by the data generation model 58. The message may be interpreted as an example of information corresponding to the content of the user 20's utterance.
[0046] As shown in FIG. 4, the specific processing unit 290 includes an input unit 291, a processing unit 292, and an output unit 293.
[0047] The input unit 291 acquires a user input received by the earphone 14. Specifically, the input unit 291 acquires the user's voice received by the earphone 14.
[0048] The processing unit 292 performs identification processing using the data generation model 58. Specifically, the processing unit 292 inputs a voice input by the user into the data generation model 58 and obtains a generation result. More specifically, when an utterance related to the memory or behavior of the user 20 is received from the user 20 wearing the earphones 14, the processing unit 292 performs the identification processing by suggesting information corresponding to the content of the utterance to the user 20.
[0049] The output unit 293 transmits the result of the specific processing to the earphone 14. In the earphone 14, the control unit 46A causes the speaker 40 to output the result of the specific processing. The microphone 38 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0050] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generative AI models. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0051] Next, the operation of the data processing system 10 will be described.
[0052] An example of the flow of the identification process will be described with reference to Fig. 5. The flow of the identification process shown in Fig. 5 is an example of a "data processing method" according to the technology of the present disclosure.
[0053] In step S300, the data processing device 12 receives user data including sounds and images collected by the two earphones 14.
[0054] In step S302, when the data processing device 12 receives a speech from a user wearing the earphones 14 regarding the memory or behavior of the user 20, the data processing device 12 executes a specific process to suggest information corresponding to the content of the speech to the user 20 based on the life log of the user 20.
[0055] In step S303, the data processing device 12 executes a process of reproducing the result of the specific process from the speaker 40.
[0056] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[0057] In the above embodiment, an example was given in which a specific process is performed by one computer 22, but the technology disclosed herein is not limited to this, and distributed processing of the specific process may be performed by multiple computers including computer 22.
[0058] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[0059] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0060] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[0061] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[0062] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[0063] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[0064] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[0065] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[0066] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[0067] In addition, the following supplementary notes are provided in relation to the above description.
[0068] (Appendix 1) an input unit for inputting user data including sounds and images collected by two earphones including a microphone, a speaker, and a camera and worn on the user's ears; a processing unit that performs a specific process using a data generation model that generates a predetermined inference result according to the user data; an output unit that reproduces the result of the specific processing from the speaker; Equipped with the input unit inputs, as the user data, the sound detected by the microphone and the image captured by the camera; When the processing unit receives an utterance from the user wearing the earphones regarding the user's memory or behavior, the processing unit performs the specific processing of suggesting information to the user corresponding to the content of the utterance based on the user's life log in which the sounds and images associated with the user are recorded.
[0069] (Appendix 2) The data processing device described in Appendix 1, wherein when the user wearing the earphones requests a message that will trigger a specific memory as the content of the utterance, the processing unit suggests to the user one or more messages selected based on the life log as information corresponding to the content of the utterance.
[0070] (Appendix 3) The data processing device described in Appendix 1 or 2, wherein when the user wearing the earphones tweets a specific matter as the content of the speech, the processing unit suggests to the user recommended actions for the user regarding the matter based on the life log as information corresponding to the content of the speech.
[0071] (Appendix 4) A data processing method in which a computer executes a specific process using a data generation model that inputs user data including sounds and images collected by two earphones that include a microphone, a speaker, and a camera and are attached to the user's ears, and generates a predetermined inference result according to the user data, inputting the sound detected by the microphone and the image captured by the camera as the user data; When an utterance relating to the memory or behavior of the user is received from the user wearing the earphones, the identification process is performed by suggesting information corresponding to the content of the utterance to the user based on a life log of the user in which the sounds and images associated with the user are recorded; a process of reproducing a result of the specific process from the speaker; A data processing method executed by the computer.
[0072] (Appendix 5) A data processing program that causes a computer to execute a specific process using a data generation model that inputs user data including sounds and images collected by two earphones that include a microphone, a speaker, and a camera and are attached to the user's ears, and generates a predetermined inference result according to the user data, inputting the sound detected by the microphone and the image captured by the camera as the user data; When an utterance relating to the memory or behavior of the user is received from the user wearing the earphones, the identification process is performed by suggesting information corresponding to the content of the utterance to the user based on a life log of the user in which the sounds and images associated with the user are recorded; a process of reproducing a result of the specific process from the speaker; A data processing program to be executed by the computer. [Explanation of symbols]
[0073] 10 Data Processing System 12 Data Processing Device 14 Earphones 290 Special Processing Department 291 Input section 292 Processing section 293 Output Section< / url:>
Claims
1. an input unit for inputting user data including sounds and images collected by two earphones including a microphone, a speaker, and a camera and worn on the user's ears; a processing unit that performs a specific process using a data generation model that generates a predetermined inference result according to the user data; an output unit that reproduces the result of the specific processing from the speaker; Equipped with the input unit inputs, as the user data, the sound detected by the microphone and the image captured by the camera; When the processing unit receives an utterance from the user wearing the earphones regarding the user's memory or behavior, the processing unit performs the specific processing of suggesting information to the user corresponding to the content of the utterance based on the user's life log in which the sounds and images associated with the user are recorded.
2. The data processing device according to claim 1, wherein when the user wearing the earphones requests a message that will trigger a specific memory as the content of the utterance, the processing unit suggests to the user one or more messages selected based on the life log as information corresponding to the content of the utterance.
3. The data processing device according to claim 1, wherein when the user wearing the earphones tweets a specific matter as the content of the speech, the processing unit suggests to the user recommended actions for the matter based on the life log as information corresponding to the content of the speech.
4. A data processing method in which a computer executes a specific process using a data generation model that inputs user data including sounds and images collected by two earphones that include a microphone, a speaker, and a camera and are attached to the user's ears, and generates a predetermined inference result according to the user data, inputting the sound detected by the microphone and the image captured by the camera as the user data; When an utterance relating to the memory or behavior of the user is received from the user wearing the earphones, the identification process is performed by suggesting information corresponding to the content of the utterance to the user based on a life log of the user in which the sounds and images associated with the user are recorded; a process of reproducing a result of the specific process from the speaker; A data processing method executed by the computer.
5. A data processing program that causes a computer to execute a specific process using a data generation model that receives user data including sounds and images collected by two earphones that include a microphone, a speaker, and a camera and are attached to the user's ears, and generates a predetermined inference result according to the user data, inputting the sound detected by the microphone and the image captured by the camera as the user data; When an utterance relating to the memory or behavior of the user is received from the user wearing the earphones, the identification process is performed by suggesting information corresponding to the content of the utterance to the user based on a life log of the user in which the sounds and images associated with the user are recorded; a process of reproducing a result of the specific process from the speaker; A data processing program to be executed by the computer.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A