Data processing device, data processing method, and data processing program
The data processing device enhances specific sounds by identifying the user's focus through image analysis, addressing the issue of uniform noise reduction in noise cancellation technologies, allowing users to concentrate on desired voices.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-11-21
- Publication Date
- 2026-06-02
Smart Images

Figure 2026089991000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a data processing device, a data processing method, and a data processing program.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
[0003] On the other hand, in earphones, attempts have been made to reduce ambient noise using noise cancellation technology.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] However, noise cancellation technology uniformly reduces ambient noise, and even when a user wants to concentrate on a specific voice, the voice they want to hear is also reduced.
[0006] The present disclosure has been made in view of the above circumstances, and aims to enable a user to concentrate on a specific voice.
Means for Solving the Problems
[0007] A first aspect of the technology of this disclosure includes an input unit that inputs audio data and image data collected by two earphones worn on the user's ears, and a microphone, speaker, and camera. A processing unit that performs specific processing using a data generation model that generates predetermined inference results corresponding to the audio data and the image data, An output unit that reproduces the result of the specified processing from the speaker, Equipped with, The processing unit is a data processing device that identifies an object of interest to the user wearing the earphones by analyzing the image data, and performs a process to emphasize the sound from the object in the audio data as the identification process.
[0008] A second aspect of the technology of this disclosure is a data processing method in which a computer performs specific processing using a data generation model that includes a microphone, a speaker, and a camera, inputs audio data and image data collected by two earphones worn on a user's ears, and generates predetermined inference results corresponding to the audio data and image data, The audio data detected by the microphone and the image data captured by the camera are input. The object that the user wearing the earphones is focusing on is identified by analyzing the image data, and the identification process is performed to emphasize the sound from the object in the audio data. This is a data processing method that reproduces the results of the aforementioned specific processing through the speaker.
[0009] A third aspect of the technology of this disclosure is a data processing program that causes a computer to perform specific processing using a data generation model that includes a microphone, a speaker, and a camera, inputs audio data and image data collected by two earphones worn on a user's ears, and generates predetermined inference results corresponding to the audio data and image data, A procedure for inputting audio data detected by the microphone and image data captured by the camera, A procedure for identifying an object of interest to a user wearing the aforementioned earphones by analyzing the aforementioned image data, and for performing a process to emphasize the sound from the aforementioned object in the audio data as the identification process, This is a data processing program that causes a computer to perform the procedure of playing back the results of the aforementioned specific processing through the speaker.
[0010] Audio enhancement processing refers to making a target audio sound louder compared to other audio sounds. For example, this includes making the gain of the target audio louder than the gain of other audio sounds, making the gain of other audio sounds quieter, and making the gain of the target audio louder than the gain of other audio sounds, and also making the gain of the target audio louder than the gain of other audio sounds, and making the gain of the other audio sounds quieter. [Brief explanation of the drawing]
[0011] [Figure 1] Figure 1 is a conceptual diagram showing an example of the configuration of a data processing system. [Figure 2] Figure 2 is a conceptual diagram showing an example of the main functions of a data processing device and earphones. [Figure 3A] Figure 3A shows an example of an earphone configuration. [Figure 3B] Figure 3B shows the user wearing earphones. [Figure 3C] Figure 3C is a diagram illustrating the camera's field of view. [Figure 3D] Figure 3D shows the user wearing the earphones. [Figure 3E] Figure 3E shows the user wearing the earphones. [Figure 3F] Figure 3F shows the user wearing earphones. [Figure 4] The functional configuration of a specific processing unit of a data processing device is shown in general terms. [Figure 5] An example of the operation flow of a specific process performed by the data processing device according to the first embodiment is schematically shown. [Figure 6]An example of the operation flow of the specific processing by the data processing device according to the second embodiment is schematically shown.
Mode for Carrying Out the Invention
[0012] Hereinafter, an example of an embodiment of a data processing device, a data processing method, and a program according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0013] First, the terms used in the following description will be explained.
[0014] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of the arithmetic unit include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), or an APU (Accelerated Processing Unit), etc.
[0015] In the following embodiments, the signed RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0016] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of the non-volatile storage device include a flash memory (SSD (Solid State Drive)), a magnetic disk (e.g., a hard disk), or a magnetic tape, etc.
[0017] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0018] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0019] <First Embodiment> Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0020] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and earphones 14. An example of the data processing device 12 is a server. In this embodiment, the data processing device 12 is an example of a "data processing device" according to the technology of this disclosure.
[0021] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0022] The earphone 14 includes a computer 36, a microphone 38, a speaker 40, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 38, speaker 40, and camera 42 are also connected to the bus 52.
[0023] The microphone 38 receives voice signals from the user 20 and accepts instructions from the user 20. The microphone 38 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 40 outputs audio according to the instructions from the processor 46. Hereafter, the microphone 38 may be simply referred to as the microphone 38.
[0024] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision). The lens closest to the subject among the lenses constituting camera 42 may be an ultra-wide-angle lens, such as a fisheye lens.
[0025] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0026] Figure 2 shows an example of the main functions of the data processing device 12 and the earphone 14.
[0027] As shown in Figure 2, in the data processing device 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "data processing program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0028] The storage 32 stores the data generation model 58. The data generation model 58 is used by the specific processing unit 290.
[0029] (Earphones 14) In the earphone 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0030] The earphone 14 may be interpreted as a canal-type earphone that is fitted into the ear canal of the user 20, as shown in Figure 3A. However, the earphone 14 is not limited to a canal type; it may also be an inner-ear type earphone that is inserted into the inner ear of the user 20, or a headphone type earphone that covers the entire ear of the user 20. Each of the two earphones 14 is equipped with a microphone 38, a speaker 40, and a camera 42. The sound and images collected by the two earphones 14 fitted into the ears of the user 20 may be recorded as a life log in the database 24.
[0031] The life log can be interpreted as a history of the user 20's actions in daily life, and may include sounds and images associated with the user 20, specifically sounds collected by the microphone 38 and images taken by the camera 42 during daily life. The life log may record sounds and images associated with the user 20, along with the date, time, and location in which they were acquired.
[0032] The sounds collected by the microphone 38 may include the voice of the person the user 20 is talking to, and sounds that occur around the user 20 while walking or cycling (such as the sound of cars driving, birds chirping, the babbling of a stream, and the sound of trees swaying in the wind).
[0033] As shown in Figure 3C, the camera 42 may capture images of the scenery within a field of view that captures the area in front of the user 20, or it may capture images of scenery within a field of view that captures areas other than in front of the user 20, such as to the side, behind, below, or above the user 20. The images captured by the camera 42 may include the appearance of the person the user 20 is talking to, the scenery around the user 20 when they are walking or cycling, and the appearance of the pet the user 20 is walking with. If the camera 42 in this embodiment has an ultra-wide-angle lens, the camera 42 can also capture images of the user 20's eyes.
[0034] Since each of the two earphones 14 is equipped with a camera 42, the two earphones 14 worn on the user's ears 20 are positioned at a specific distance apart, one on the left ear and the other on the right ear, as shown in Figure 3B. Therefore, compared to cases where two cameras are arranged side by side in a single housing, such as in a video camera, the spacing between the two cameras 42 can be increased, making 3D sensing easier. 3D sensing can be interpreted as measuring three-dimensional shapes.
[0035] Furthermore, when the two earphones 14 are placed in the user 20's ears, the two cameras 42 are positioned close to the user 20's left and right eyes, allowing images (captured images) that are nearly identical to those seen with the naked eye to be recorded as a life log in the database 24. Consequently, in specific processing, it becomes easier to reproduce information corresponding to inquiries from the user 20, that is, information corresponding to the content of the user 20's speech.
[0036] While the two earphones 14 are attached to the user 20, all or part of the images captured by the camera 42 may be recorded in the database 24 as a life log. Specifically, when the two earphones 14 are attached to the user 20, the recording of images captured by the camera 42 to the database 24 may begin, and when the two earphones 14 are removed from the user 20, the recording of those images to the database 24 may end.
[0037] While the two earphones 14 are worn by the user 20, all or part of the sound collected by the microphone 38 may be recorded as a lifelog in the database 24. Specifically, when the two earphones 14 are worn by the user 20, the recording of the sound collected by the microphone 38 to the database 24 may begin, and when the two earphones 14 are removed from the user 20, the recording of the sound to the database 24 may end.
[0038] Next, in the first embodiment, when the data processing device 12 receives an utterance from the user 20 wearing the earphones 14 regarding the user 20's memories or actions, the processing of the specific processing unit 290 when it performs specific processing to propose information corresponding to the content of the user 20's utterance to the user 20 will be described.
[0039] (Specific processing) In the first embodiment, the identification process uses a data generation model that takes user data as input and generates predetermined inference results corresponding to the input user data. Specifically, in the identification process, when utterances related to the user 20's memories or actions are received as user data from a user 20 wearing earphones 14, the process is executed to propose information corresponding to the content of the utterances to the user 20 by referring to the database 24. Specifically, after a life log is recorded in the database 24, if the user 20 wearing earphones 14 makes an utterance related to the user 20's memories or actions, the identification process may be executed to propose information corresponding to the content of the utterances to the user 20 by referring to the database 24.
[0040] (Example of specific processing) If the user wearing the earphones requests a message that will trigger the recall of a specific memory, the specific processing unit 290 may propose one or more messages selected based on the life log to the user who made the request, as information corresponding to the content of the utterance (request).
[0041] For example, if user 20, wearing earphones 14, tries to recall their memory and asks, "What did I say to person A around [date] at [time]?", the identification processing unit 290, as part of its identification process, inputs this message as a prompt to the data generation model 58. The identification processing unit 290 may refer to the life log in database 24 and, based on the output obtained from the data generation model 58, generate a message such as, "I think you said, 'I found a nice restaurant, let's make a reservation.'" This message may be interpreted as an example of information corresponding to the content of user 20's utterance.
[0042] For example, if user 20 wearing earphones 14 tries to recall their memory and asks, "Who was I talking to around [date] at [time]?", the identification processing unit 290 will input this message as a prompt to the data generation model 58 as part of its identification process. The identification processing unit 290 may refer to the life log in database 24 and, based on the output obtained from the data generation model 58, generate a message such as, "It seems you were talking with two friends at that time, probably B and C." This message may be interpreted as an example of information corresponding to the content of user 20's utterance.
[0043] For example, if user 20, wearing earphones 14, tries to recall their emotions and says, "How did I feel when I was talking to person A around [date] at [time]?", the identification processing unit 290, as part of its identification process, inputs this message as a prompt to the data generation model 58. The identification processing unit 290 may refer to the life log in database 24 and, based on the output obtained from the data generation model 58, generate a message such as, "At that time, you were laughing a lot, so it seems you had a good impression of your friend and were very happy." This message may be interpreted as an example of information corresponding to the content of user 20's utterance.
[0044] (Example of specific processing, part 2) If a user 20 wearing earphones 14 mutters a specific matter as part of their utterance, the specific processing unit 290 may suggest to the user 20 who requested the message, based on their life log, recommended actions for the user 20 regarding that matter, as information corresponding to the content of their utterance (muttering).
[0045] For example, when user 20 wearing earphones 14 is shopping at a specific retail store and says, "What should I buy?", the specific processing unit 290 inputs this message as a prompt to the data generation model 58 as a specific processing step. The specific processing unit 290 may refer to the life log in the database 24 and, based on the output obtained by the data generation model 58, generate a message such as, "A few months ago, you purchased product A at this store and commented that it wasn't very tasty, so how about purchasing recently released products B and C this time?" This message may be interpreted as an example of information corresponding to the content of user 20's utterance.
[0046] (Third example of specific processing) As shown in Figure 3D, when user 20, wearing earphones 14, is operating a PC and says, "What was the name of product A that I searched for the day before yesterday?", the identification processing unit 290 inputs this message as a prompt to the data generation model 58 as part of its identification processing. The data generation model 58 refers to the life log in the database 24 and analyzes the video of the PC screen when user 20 was operating it in the past to generate a specific output. Based on the output obtained by the data generation model 58, the identification processing unit 290 may generate a message such as "Product A is ○○○". This message may be interpreted as an example of information corresponding to the content of user 20's utterance.
[0047] (Fourth example of specific processing) As shown in Figure 3E, if user 20, wearing earphones 14, says "There was a place nearby with a great view, but I wonder where it is?" while cycling, the identification processing unit 290 inputs this message as a prompt to the data generation model 58 as part of its identification process. The data generation model 58 refers to the life log in database 24 and analyzes places previously visited by user 20 and the route to those places to generate a specific output. Based on the output obtained by the data generation model 58, the identification processing unit 290 may generate a message such as "I think it's Cape XX, about 500m from here." This message can be interpreted as an example of information corresponding to the content of user 20's utterance.
[0048] (Example 5 of specific processing) As shown in Figure 3F, when user 20, wearing earphones 14, meets Mr. X at company A, the company he is visiting, and says, "Can you tell me this person's name?", the identification processing unit 290 inputs this message as a prompt to the data generation model 58 as part of the identification process. The data generation model 58 refers to the life log in database 24 and generates specific output from the history of people that user 20 met when he visited company A. Based on the output obtained from the data generation model 58, the identification processing unit 290 may generate a message such as, "I think his name is ○○." This message may be interpreted as an example of information corresponding to the content of user 20's utterance.
[0049] As shown in Figure 4, the specific processing unit 290 includes an input unit 291, a processing unit 292, and an output unit 293.
[0050] The input unit 291 acquires user input received through the earphone 14. Specifically, it acquires the user's voice received through the earphone 14.
[0051] The processing unit 292 performs specific processing using the data generation model 58. Specifically, it inputs voice from the user into the data generation model 58 and obtains a generation result. More specifically, when it receives an utterance from the user 20 wearing the earphones 14 regarding the user 20's memories or actions, it performs a specific processing step of proposing information corresponding to the content of the utterance to the user 20.
[0052] The output unit 293 transmits the result of the specific processing to the earphone 14. In the earphone 14, the control unit 46A causes the speaker 40 to output the result of the specific processing. The microphone 38 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0053] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0054] Next, the operation of the data processing system 10 in the first embodiment will be described.
[0055] An example of the flow of a specific processing method will be explained with reference to Figure 5. Note that the flow of a specific processing method shown in Figure 5 is an example of a "data processing method" related to the technology disclosed herein.
[0056] In step S300, the data processing device 12 receives user data, including sound and images collected by the two earphones 14.
[0057] In step S302, if the data processing device 12 receives an utterance from the user wearing the earphones 14 regarding the user's memories or actions, it executes a specific process to propose information to the user 20 that corresponds to the content of the utterance, based on the user's life log.
[0058] In step S303, the data processing device 12 executes a process to play back the result of a specific process from the speaker 40.
[0059] Next, a second embodiment of the present disclosure will be described. The configuration of the data processing system according to the second embodiment is the same as that of the data processing system 10 according to the first embodiment, so a detailed explanation will be omitted here.
[0060] In the second embodiment, the camera 42 of the earphone 14 has an ultra-wide-angle lens and is capable of capturing the user's eyes.
[0061] Next, in the second embodiment, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described.
[0062] In the identification process in the second embodiment, audio data and image data collected by the earphone 14 are input, and an identification process is performed using a data generation model that generates data that produces a predetermined inference result corresponding to the input audio data and image data. Specifically, in the identification process, the image data is analyzed to identify the object that the user is paying attention to, and the process of emphasizing the sound from the object in the audio data is performed as the identification process.
[0063] The specific processing unit 290 extracts the user's pupil area from two image data acquired by the cameras 42 of the left and right earphones 14 and derives the user's gaze direction. The gaze direction can be derived using 3D gaze measurement technology. The specific processing unit 290 may also derive the user's gaze movement and blinking frequency. The derived user's gaze direction is represented by the pixel position in the image data.
[0064] Once the user's gaze direction is determined, the identification processing unit 290 inputs the result of the user's gaze direction determination, along with audio data, image data, and the prompt "Identify the object at the end of the gaze direction from the image data, and generate audio data emphasizing the sound emitted by the object from the audio data" to the data generation model 58. The data generation model 58 identifies the object in the direction of the user's gaze in the image represented by the image data. In this case, the data generation model 58 extracts all subjects included in the image and identifies the object in the direction of the user's gaze from among all subjects as the target.
[0065] The specific processing unit 290 uses the data generation model 58 to analyze the audio data and separate the voice emitted by the target from the audio data. As a method for separating the voice emitted by the target, for example, see "https: / / medium.com / axinc / voicefilter-%E4%BB%BB%E6%84%8F%E3%81%AE%E4%BA%BA%E7%89%A9%E3%81%AE%E5%A3%B0%E3%82%92%E6%8A%BD%E5%87%BA%E3%81%A7%E3%81%8D%E3%82%8B% You can use AI model-based methods proposed at "E9%9F%B3%E5%A3%B0%E5%88%86%E9%9B%A2%E3%83%A2%E3%83%87%E3%83%AB-d5b88a8549d9" or "https: / / crystal-method.com / solutions / sound-source-separation-system / ". The former is a method that separates a specific person's voice from audio where multiple people are speaking by treating all sounds other than the specific person's voice as noise and feeding the specific person's voice to the AI model. The latter is a method that separates a specific sound from audio where various sounds are mixed by training the AI model with specific sounds such as human voices, machine sounds, and musical instrument sounds.
[0066] The specific processing unit 290 uses the data generation model 58 to enhance the target audio by emphasizing the separated audio. For example, it generates audio data that emphasizes the separated audio by increasing the gain of the separated audio, decreasing the gain of audio other than the separated audio, or increasing the gain of the separated audio and decreasing the gain of audio other than the separated audio. Note that decreasing the gain of audio other than the separated audio effectively reduces noise. The specific processing unit 290 may also optimize the sound quality of the emphasized audio to make it easier to listen to.
[0067] The specific processing unit 290 may adjust parameters according to the user's situation when emphasizing the sound. For example, it may add a prompt to the data generation model 58 that says, "Determine whether the user is focused or distracted based on the user's eye movements or blinking frequency, and if the user is distracted, increase the gain to emphasize the sound."
[0068] The specific processing unit 290 outputs audio data with the separated audio enhanced to the earphone 14. The user can then hear the enhanced audio through the earphone 14.
[0069] Next, the operation of the data processing system 10 in the second embodiment will be described.
[0070] An example of the flow of a specific processing method will be explained with reference to Figure 6. Note that the flow of a specific processing method shown in Figure 6 is an example of a "data processing method" related to the technology disclosed herein.
[0071] In step S400, the data processing device 12 receives user data, including sound and images collected by the two earphones 14.
[0072] In step S402, the data processing device 12 analyzes the input image to identify the object that the user is interested in.
[0073] In step S404, the data processing device 12 highlights the sound generated by the identified target.
[0074] In step S406, the data processing device 12 performs a process to play the enhanced audio from the speaker 40.
[0075] In the second embodiment, a database including the user's gaze patterns, voice history, and user preferences may be generated and stored in the data processing device 12. By additionally inputting prompts to the data generation model 58 that instruct it to optimize the voice according to the user's preferences by referring to such a database, the user can be provided with voice that meets their individual needs.
[0076] The specific processing unit 290 may, if necessary, output the status of eye tracking and voice enhancement to the earphone 14 via voice guidance. Alternatively, it may notify the user via an app to a mobile device other than the earphone 14. This allows the user to understand the current settings and status.
[0077] Furthermore, the specific processing unit 290 may acquire user environment information and operation information, and additionally input prompts to the data generation model 58 instructing it to adjust the parameters of the specific processing (for example, parameters related to the degree to which the gain of the separated audio is increased, and parameters related to the degree to which the gain of audio other than the separated audio is decreased) in real time based on the environment information and operation information. This makes it possible to continuously improve the accuracy and performance of the audio enhancement processing.
[0078] Furthermore, the data generation model 58 may learn the user's past gaze patterns and voice history and optimize voice emphasis according to the user's preferences.
[0079] In the second embodiment, the camera 42 has an ultra-wide-angle lens, but is not limited to this. The camera 42 may extract a subject located in the center of the image of the user's front acquired by the camera.
[0080] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0081] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.
[0082] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0083] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0084] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0085] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0086] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0087] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0088] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0089] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0090] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[0091] Furthermore, the following additional information is disclosed regarding the above explanation. (Additional note 1) The system includes a microphone, speaker, and camera, and an input section that receives audio and image data collected by two earphones worn by the user. A processing unit that performs specific processing using a data generation model that generates predetermined inference results corresponding to the audio data and the image data, An output unit that reproduces the result of the specified processing from the speaker, Equipped with, The processing unit is a data processing device that identifies an object of interest to a user wearing the earphones by analyzing the image data, and performs a process to emphasize the sound from the object in the audio data as the identification process. (Additional note 2) The processing unit is a data processing device according to Appendix 1, which acquires user environment information and operation information, and uses the data generation model to adjust the parameters of the specific processing in real time based on the environment information and operation information. (Additional note 3) The processing unit is a data processing device according to Appendix 1 or 2, which uses the data generation model to optimize the voice according to the user's preferences based on the user's past gaze patterns and voice history. (Additional note 4) A data processing method in which a computer performs specific processing using a data generation model that includes a microphone, speaker, and camera, inputs audio data and image data collected by two earphones worn on the user's ears, and generates predetermined inference results corresponding to the audio data and image data, The audio data detected by the microphone and the image data captured by the camera are input. The object that the user wearing the earphones is focusing on is identified by analyzing the image data, and the identification process is performed to emphasize the sound from the object in the audio data. The result of the aforementioned specific processing is played back from the speaker. Data processing method. (Additional note 5) A data processing program that causes a computer to perform specific processing using a data generation model that includes a microphone, speaker, and camera, inputs audio data and image data collected by two earphones worn on the user's ears, and generates predetermined inference results corresponding to the audio data and image data, A procedure for inputting audio data detected by the microphone and image data captured by the camera, A procedure for identifying an object of interest to a user wearing the aforementioned earphones by analyzing the aforementioned image data, and for performing a process to emphasize the sound from the aforementioned object in the audio data as the identification process, A data processing program that causes a computer to perform a procedure for playing back the results of the aforementioned specific processing through the speaker. [Explanation of symbols]
[0092] 10 Data Processing Systems 12 Data Processing Devices 14 Earphones 42 cameras 290 Specific Processing Unit 291 Input section 292 Processing Unit 293 Output section< / url:>
Claims
1. It includes a microphone, speaker, and camera, and an input section that receives audio and image data collected by two earphones worn on the user's ears, A processing unit that performs specific processing using a data generation model that generates predetermined inference results corresponding to the audio data and the image data, An output unit that reproduces the result of the specified processing from the speaker, Equipped with, The processing unit is a data processing device that identifies an object of interest to a user wearing the earphones by analyzing the image data, and performs a process to emphasize the sound from the object in the audio data as the identification process.
2. The data processing device according to claim 1, wherein the processing unit acquires user environment information and operation information, and uses the data generation model to adjust the parameters of the specific processing in real time based on the environment information and operation information.
3. The processing unit optimizes the voice according to the user's preferences based on the user's past gaze patterns and voice history using the data generation model, as described in claim 1 or 2.
4. A data processing method in which a computer performs specific processing using a data generation model that includes a microphone, speaker, and camera, inputs audio data and image data collected by two earphones worn on the user's ears, and generates predetermined inference results corresponding to the audio data and image data, The audio data detected by the microphone and the image data captured by the camera are input. The object that the user wearing the earphones is focusing on is identified by analyzing the image data, and the identification process is performed to emphasize the sound from the object in the audio data. The result of the aforementioned specific processing is played back from the speaker. Data processing method.
5. A data processing program that causes a computer to perform specific processing using a data generation model that includes a microphone, speaker, and camera, inputs audio data and image data collected by two earphones worn on the user's ears, and generates predetermined inference results corresponding to the audio data and image data, A procedure for inputting audio data detected by the microphone and image data captured by the camera, A procedure for identifying an object of interest to a user wearing the aforementioned earphones by analyzing the aforementioned image data, and for performing a process to emphasize the sound from the aforementioned object in the audio data as the identification process, A data processing program that causes a computer to perform a procedure for playing back the results of the aforementioned specific processing through the speaker.