Information processing device, information processing method, and program
The information processing apparatus addresses the challenge of balancing computing resources and response latency in voice processing systems by utilizing a second neural network that integrates outputs from a first neural network and diverse sensing data, achieving efficient and accurate speech processing.
Patent Information
- Application Number
- PCT/JP2024/038484
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-22
- Filing Date
- 2024-10-29
- Publication Date
- 2025-06-26
AI Technical Summary
Existing voice processing systems face challenges in balancing computing resources and response latency, particularly when performing speech recognition on cloud servers, which often results in increased processing costs and reduced efficiency.
An information processing apparatus that includes a second task processing unit with a second neural network, which takes the output from a first neural network and second sensing data as inputs to perform a second task related to speech processing. The first and second sensing data are acquired by different types of sensors, allowing for efficient preprocessing and reduced overall processing costs.
This configuration enables speech processing with reduced computing resources, improving efficiency and reducing response latency while maintaining accurate speech recognition and estimation tasks.
Smart Images

Figure JP2024038484_26062025_PF_FP_ABST
Abstract
Description
Information processing device, information processing method, and program
[0001] The present disclosure relates to an information processing device, an information processing method, and a program.
[0002] In recent years, systems have been developed that collect voice data related to user utterances and perform processing based on the voice data. For example, Patent Literature 1 discloses a system that performs voice recognition based on collected voice data.
[0003] JP 2023-120068 A
[0004] Processing based on speech, such as speech recognition, is significantly affected by computing resources and response delays.
[0005] According to one aspect of the present disclosure, there is provided an information processing device comprising: a second task processing unit including a first neural network that receives as input first sensing data related to an utterance and executes a first task corresponding to the utterance, the first task result being output from the first neural network; and a second neural network that receives as input second sensing data related to the utterance and executes a second task corresponding to the utterance, wherein the first sensing data and the second sensing data are acquired by sensors of different types.
[0006] According to another aspect of the present disclosure, there is provided an information processing method including a processor inputting first sensing data related to an utterance to a first neural network, the first neural network receiving the first sensing data related to the utterance and executing a first task corresponding to the utterance, and inputting the result of the first task output from the first neural network and second sensing data related to the utterance to a second neural network, and executing a second task corresponding to the utterance, wherein the first sensing data and the second sensing data are acquired by sensors of different types.
[0007] According to another aspect of the present disclosure, there is provided a program that causes a computer to function as an information processing device, comprising: a second task processing unit that receives first sensing data related to an utterance as input and executes a first task corresponding to the utterance, the first task result being output from a first neural network that receives second sensing data related to the utterance as input and executes a second task corresponding to the utterance, and the first sensing data and the second sensing data being acquired by sensors of different types.
[0008] FIG. 1 is a block diagram showing an example of a schematic functional configuration of an information processing device 10 according to an embodiment of the present disclosure. A more specific example of a functional configuration of the information processing device 10 according to the embodiment will be described. A diagram for explaining an example of training of a speech neural network 113 and a fusion neural network 122 according to the embodiment. A diagram showing an example of a training dataset according to the embodiment. A diagram showing an example of a training dataset according to the embodiment. A diagram for explaining an example of training when a second task according to the embodiment is emotion estimation. A flowchart showing an example of the operation flow of the information processing device 10 according to the embodiment. A block diagram showing an example of the hardware configuration of an information processing device 90 according to the embodiment. A diagram for explaining computational resources and response delay for each processing entity.
[0009] Preferred embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functional configurations are designated by the same reference numerals, and redundant description will be omitted.
[0010] The description will be given in the following order: 1. Embodiment 1.1. Background 1.2. Functional configuration example 1.3. Learning example 1.4. Operation example 2. Hardware configuration example 3. Summary
[0011] 1. Embodiments 1.1. Background As described above, in recent years, systems have been developed that collect voice data related to user utterances and perform processing based on the voice data.
[0012] For example, Patent Document 1 discloses a system in which voice data collected by a meeting device equipped with a microphone and a speaker is uploaded to the cloud via a PC (Personal Computer), and voice recognition based on the voice data is performed on the cloud.
[0013] The system disclosed in Patent Document 1 can reduce the processing load on meeting devices and PCs by performing speech recognition in the cloud, which has abundant computing resources.
[0014] However, when processing based on voice data is performed on the cloud, as in the system disclosed in Patent Document 1, response delays generally increase.
[0015] FIG. 9 is a diagram for explaining the computational resources and response delays for each processing entity.
[0016] Figure 9 shows an edge terminal 50 that collects speech data related to the speech of a user (hereinafter also referred to as the speaker) and provides the user with the results of some processing based on the speech data, a server 70 located in the cloud, and a smartphone 60.
[0017] The edge terminal 50 may be, for example, an earphone or other earphone with separate left and right earphones.
[0018] The smartphone 60 is connected to the edge terminal 50 and the server 70 so as to be able to communicate information with them.
[0019] In this case, the smartphone 60 can receive the voice data collected by the edge terminal 50 , perform processing based on the voice data, and transmit the results of the processing to the edge terminal 50 .
[0020] In addition, the server 70 can receive the voice data collected by the edge terminal 50 via the smartphone 60, perform processing based on the voice data, and transmit the results of the processing to the edge terminal 50 via the smartphone 60.
[0021] The edge terminal 50 can also perform processing based on the voice data it collects and provide the results of the processing to the user.
[0022] When the edge terminal 50 performs processing based on voice data, the response delay is smallest compared to when the smartphone 60 performs the processing and when the server 70 performs the processing. However, the edge terminal 50 generally has smaller computing resources than the smartphone 60 and the server 70, and therefore there is a possibility that the computing resources will be insufficient.
[0023] When the smartphone 60 performs processing based on voice data, the response delay is medium compared to when the edge terminal 50 performs the processing and when the server 70 performs the processing. In addition, the smartphone 60 generally has medium-level computing resources compared to the edge terminal 50 and the server 70.
[0024] When the server 70 performs processing based on voice data, the response delay is greatest compared to when the edge terminal 50 performs the processing and when the smartphone 60 performs the processing. However, the server 70 generally has larger computing resources than the edge terminal 50 and the smartphone 60.
[0025] As illustrated above, it can be said that, in general, there is a trade-off between computational resources and response delays with respect to a processing entity that performs processing based on audio data.
[0026] The technical concept of one embodiment of the present disclosure was conceived with the above points in mind, and makes it possible to perform speech-based processing using smaller computational resources.
[0027] An example of a functional configuration that achieves the above will be described in detail below.
[0028] <<1.2. Functional Configuration Example>> FIG. 1 is a block diagram showing a schematic functional configuration example of an information processing device 10 according to an embodiment of the present disclosure.
[0029] The information processing device 10 according to this embodiment may be, for example, an earphone or other earphone-like separate earphones, or other head-mounted wearable devices.
[0030] On the other hand, the information processing device 10 according to the present embodiment may be, for example, a stationary device or a smartphone that provides a voice agent function.
[0031] As shown in FIG. 1, the information processing device 10 according to this embodiment includes a first task processing unit 110, a second task processing unit 120, a task result processing unit 130, an audio input / output processing unit 140, an audio output unit 150, and a sensor unit 160.
[0032] (First Task Processing Unit 110) The first task processing unit 110 according to this embodiment includes a first neural network that receives first sensing data related to an utterance as input and executes a first task corresponding to the utterance.
[0033] (Second Task Processing Unit 120) The second task processing unit 120 according to this embodiment includes a second neural network that receives as input the result of the first task output from the first neural network and second sensing data related to the utterance, and executes a second task corresponding to the utterance. The first sensing data and the second sensing data are acquired by sensors of different types.
[0034] Here, the first task is a partial task that is performed prior to the second task (the final task) in order to perform the second task.
[0035] The information processing device 10 according to an embodiment of the present disclosure makes it possible to reduce the overall processing cost by processing a part of a certain second task in advance as a first task.
[0036] Specific examples of the first task and the second task will be described later.
[0037] (Task Result Processing Unit 130) The task result processing unit 130 according to this embodiment performs some kind of processing based on the result of the second task.
[0038] The functional configuration of the task result processing unit 130 may be designed appropriately depending on the content of the processing to be performed.
[0039] (Audio Input / Output Processing Unit 140) The audio input / output processing unit 140 according to this embodiment performs processing related to audio input and output.
[0040] Examples of the processing performed by the audio input / output processor 140 include various types of filtering and DNC (Digital Noise Cancelling).
[0041] (Audio Output Unit 150 ) The audio output unit 150 according to this embodiment includes a speaker that outputs the audio processed by the audio input / output processing unit 140 .
[0042] (Sensor Unit 160) The sensor unit 160 according to this embodiment includes various sensors that collect first sensing data or second sensing data.
[0043] Examples of the first sensing data and the second sensing data include voice data, acceleration data, angular velocity data, image data, and vital data (for example, data relating to pulse, blood pressure, body temperature, and electrocardiogram).
[0044] The above has described an example of the general functional configuration of the information processing device 10 according to this embodiment. Note that the functional configuration described above with reference to Fig. 1 is merely an example, and the functional configuration example of the information processing device 10 according to this embodiment is not limited to this example.
[0045] For example, the information processing device 10 may further include a communication unit that communicates information with other devices, an operation input unit that accepts operations by the user, and the like.
[0046] The functional configuration of the information processing device 10 according to this embodiment can be flexibly modified according to specifications, operation, and the like.
[0047] Next, a more specific example of the functional configuration of the information processing device 10 according to this embodiment will be described with reference to FIG.
[0048] Note that Figure 2 illustrates an example of a functional configuration in which the second task is speech recognition (estimation of text data including consonants and vowels) based on speech data (an example of first sensing data), and the first task is extraction of consonants based on the speech data.
[0049] As shown in FIG. 2, the information processing device 10 according to this embodiment may include a DSP (Digital Signal Processor) 111, an NPU (Neural Processing Unit) 121, audio 141, a speaker 151, a microphone 161, an IMU (Inertial Measurement Unit) 162, and a camera 163.
[0050] (DSP 111 ) The DSP 111 according to this embodiment is an example of the first task processing unit 110 .
[0051] The DSP 111 according to this embodiment includes a preprocessing unit 112 and a speech neural network (also referred to as speech NN) 113 .
[0052] (Pre-processing unit 112) The pre-processing unit 112 according to this embodiment performs fast Fourier transform, extraction of Mel frequency cepstrum coefficients, etc. on the voice data input from the audio 141, and inputs the processed voice data to the voice NN 113.
[0053] (Audio Neural Network 113) The audio neural network 113 according to this embodiment is an example of a first neural network.
[0054] The speech neural network 113 according to this embodiment extracts consonants corresponding to speech from the speech data input from the preprocessing unit 112 as a first task.
[0055] The details of the learning of the speech neural network 113 according to this embodiment will be described later.
[0056] (NPU 121 ) The NPU 121 according to this embodiment is an example of the second task processing unit 120 .
[0057] The NPU 121 according to this embodiment includes a fusion neural network (also referred to as a fusion NN) 122 .
[0058] (Fusion Neural Network 122) The fusion neural network 122 according to this embodiment receives as input the result of the first task output by the speech neural network 113 and the second sensing data, performs estimation of text data including consonants and vowels as the second task, and outputs the result of the second task 2R.
[0059] Examples of the second sensing data include acceleration data and angular velocity data collected by the IMU 162 provided in the sensor unit 160, and image data of the speaker's mouth obtained by the camera 163 provided in the sensor unit 160.
[0060] The learning process of the fusion neural network 122 according to this embodiment will be described in detail later.
[0061] (Audio 141) The audio 141 according to this embodiment is an example of the audio input / output processing unit 140.
[0062] The audio 141 according to this embodiment may perform filtering using a filter 143 and DNC processing by a DNC engine 142 on audio data collected by a microphone 161 included in a sensor unit 160 .
[0063] Furthermore, the audio 141 according to this embodiment may output audio data that has been subjected to filtering using the filter 143 and DNC processing by the DNC engine 142 to a speaker 151 provided in the audio output unit 150 .
[0064] <<1.3. Learning Example>> Next, with reference to FIG. 3, a learning example of the speech neural network 113 and the fusion neural network 122 shown in FIG. 2 will be described.
[0065] 3 is a computer used for training the speech neural network 113 and the fusion neural network 122. The training device 20 may be, for example, a PC.
[0066] The training device 20 includes a consonant classifier 210. The consonant classifier 210 is a trained neural network that classifies consonants from input data.
[0067] 4 is a diagram showing an example of a training data set according to this embodiment. As shown in FIG. 4, the training data set according to this embodiment includes, for example, speech data 171, IMU data 172, and correct text data 2T.
[0068] The IMU data 172 may be sensing data collected by the IMU 162 attached to the speaker's head during the same period as the voice data 171 .
[0069] The IMU data 172 includes acceleration data. The acceleration data collected by the IMU 162 attached to the speaker's head may reveal characteristics of the speaker's mouth movements corresponding to vowels. The IMU data 172 may also include angular velocity data in addition to acceleration data.
[0070] The correct text data 2T is data obtained by accurately converting the speech content corresponding to the voice data 171 into text, and includes information on consonants and vowels.
[0071] The description will continue with reference to FIG.
[0072] In the learning process, the speech neural network 113 executes a first task based on the input speech data 171 and outputs the results of the first task. The speech data 171 input to the speech neural network 113 may be subjected to various processes by the preprocessing unit 112.
[0073] The result of the first task output by the speech neural network 113 is input to the consonant classifier 210 provided in the training device 20 .
[0074] The consonant classifier 210 according to this embodiment receives the output from the speech neural network 113, that is, the result of the first task, as input, and outputs a consonant classification result 1R, which is text data consisting of only consonants.
[0075] The speech neural network 113 of this embodiment performs supervised learning based on the consonant identification result 1R estimated based on the result of the first task to be output and the correct consonant text data 1T generated from the correct text data 2T.
[0076] That is, the speech neural network 113 according to this embodiment performs training so as to reduce the difference between the consonant identification result 1R and the correct consonant text data 1T (for example, loss function=L(1R-1T)).
[0077] Through such learning, the speech neural network 113 becomes able to output data that preferentially retains information about consonants contained in the input speech data 171, i.e., data that strongly extracts the characteristics of consonants, as the result of the first task.
[0078] Next, the learning of the fusion neural network 122 according to this embodiment will be described.
[0079] In the training, the output from the speech neural network 113, i.e., the result of the first task, and the IMU data 172 are input to the fusion neural network 122.
[0080] As described above, the results of the first task are data in which the characteristics of consonants are strongly extracted, and the IMU data 172 are data in which the characteristics of the speaker's mouth movements corresponding to vowels appear.
[0081] The fusion neural network 122 executes the second task based on the above two pieces of data and outputs the result of the second task, 2R.
[0082] The result 2R of the second task is text data including consonants and vowels.
[0083] The fusion neural network 122 according to this embodiment performs supervised learning based on the output result 2R of the second task and the correct answer text data 2T.
[0084] That is, the fusion neural network 122 according to this embodiment performs training so as to reduce the difference between the result 2R of the second task and the correct text data 2T (for example, loss function=L(2R-2T)).
[0085] Through such learning, the fusion neural network 122 is able to output text data containing consonants and vowels as the result 2R of the second task, which is obtained by integrating the consonant information contained in the result of the input first task and the vowel information contained in the IMU data 172.
[0086] Furthermore, the results of the training of the fusion neural network 122 may be reflected in the training of the speech neural network 113 .
[0087] In this case, the loss function used in training the speech neural network 113 may be, for example, L(1R-1T)×N+L(2R-2T)×(1-N) (N<1).
[0088] It is expected that the above loss function will enable the speech neural network 113 to extract the features of consonants more strongly.
[0089] The learning of the speech neural network 113 and the fusion neural network 122 has been explained above using specific examples.
[0090] However, the learning described above is merely an example, and the learning according to this embodiment is not limited to this example.
[0091] For example, in the above description, the IMU data 172 is used for training the fusion neural network 122, but the image data 173 may also be used for training the fusion neural network 122.
[0092] 5 is a diagram showing an example of a training data set including image data 173 according to this embodiment. The training data set shown in FIG. 5 includes speech data 171, image data 173, and correct text data 2T.
[0093] The image data 173 may be a group of images captured by the camera 163 during the same period as the audio data 171 .
[0094] The image data 173 may strongly show the characteristics of the speaker's mouth movements according to the vowels.
[0095] Therefore, by learning using image data 173, the fusion neural network 122 can output text data including consonants and vowels as the result of the second task 2R based on the input result of the first task and image data 173.
[0096] When the second task is performed using the image data 173 as input, the information processing device 10 according to this embodiment may be a stationary device, a smartphone, a PC, AR / VR goggles, or the like.
[0097] Furthermore, in the above description, as an example of the first neural network, the speech neural network 113 is described, which receives speech data 171 as input and outputs the result of the first task in which the feature amount of consonants is strongly extracted.
[0098] On the other hand, the first neural network may be a neural network that outputs the result of a first task in which a feature amount of a vowel is strongly extracted using at least one of the IMU data 172 and the image data 173 as input. That is, the first task may be extraction of a vowel corresponding to an utterance.
[0099] In this case, the fusion neural network 122 receives the result of the first task output by the first neural network and the audio data 171 as input, and outputs text data including consonants and vowels as the result of the second task.
[0100] Furthermore, the first task is not limited to extracting features of consonants or vowels. For example, the first task may be extracting features such as voice quality and intonation. In this case, the second task may be estimating the content of the utterance and the speaker.
[0101] Furthermore, the second task is not limited to estimating the content of an utterance, but may be, for example, emotion estimation.
[0102] FIG. 6 is a diagram illustrating an example of learning when the second task according to this embodiment is emotion estimation.
[0103] The training device 20 shown in Fig. 6 includes an intonation classifier 215. The intonation classifier 215 is a trained neural network that identifies intonation from input data.
[0104] The intonation classifier 215 receives the output from the speech neural network 113 and outputs a child intonation classification result 1Rb.
[0105] The speech neural network 113 shown in FIG. 6 performs supervised learning based on the intonation classification result 1Rb estimated based on the output result of the first task and the correct intonation data 1Tb.
[0106] Through such learning, the speech neural network 113 becomes able to output data that preferentially retains the intonation information contained in the input speech data 171, i.e., data that strongly extracts intonation characteristics, as the result of the first task.
[0107] In the example shown in FIG. 6, the output from the speech neural network 113, i.e., the result of the first task, and pulse wave data (an example of vital data) collected by the pulse wave sensor 164 are input to the fusion neural network 122.
[0108] The fusion neural network 122 performs emotion estimation based on the above two pieces of data as a second task and outputs an emotion estimation result 2Rb.
[0109] The fusion neural network 122 shown in FIG. 6 performs learning so as to reduce the difference between the emotion estimation result 2Rb and the correct emotion data 2Tb.
[0110] Through such learning, the fusion neural network 122 shown in FIG. 6 becomes able to estimate the emotion of the speaker based on the intonation information and vital data contained in the input result of the first task.
[0111] As described above, the first task may be a partial task that is performed prior to the second task (the final task) in order to perform the second task, and the second task is not limited to estimating the content of an utterance.
[0112] The information processing device 10 according to an embodiment of the present disclosure makes it possible to reduce the overall processing cost by processing a part of a certain second task in advance as a first task.
[0113] Furthermore, according to the above-described processing, the size of the neural network that performs the first task and the neural network that performs the second task can be reduced, which is expected to reduce power consumption and improve estimation accuracy.
[0114] In the above description, the information processing device 10 is described as including both the first task processing unit 110 and the second task processing unit 120. However, the first task processing unit 110 and the second task processing unit 120 may be provided in different devices. Even in this case, the effect of reducing the size of the neural network can be similarly obtained.
[0115] <<1.4. Operation Example>> Next, the flow of operations of the information processing device 10 according to this embodiment will be described.
[0116] FIG. 7 is a flowchart showing an example of the flow of operations of the information processing device 10 according to this embodiment.
[0117] In the example shown in FIG. 7, the information processing device 10 first collects first sensing data and second sensing data (S101).
[0118] Next, the information processing apparatus 10 performs a first task based on the first sensing data collected in step S101 (S102).
[0119] Next, the information processing apparatus 10 performs a second task based on the result of the first task performed in step S102 and the second sensing data collected in step S101 (S103).
[0120] Next, the information processing device 10 performs a predetermined process based on the result of the second task performed in step S103 (S104).
[0121] 2. Hardware Configuration Example Next, a hardware configuration example of the information processing device 90 according to an embodiment of the present disclosure will be described. Fig. 8 is a block diagram showing a hardware configuration example of the information processing device 90 according to an embodiment of the present disclosure. The information processing device 90 may be a device having a hardware configuration equivalent to that of the information processing device 10 according to an embodiment of the present disclosure.
[0122] 8 , the information processing device 90 includes, for example, a processor 871, a ROM 872, a RAM 873, a host bus 874, a bridge 875, an external bus 876, an interface 877, an input device 878, an output device 879, a storage 880, a drive 881, a connection port 882, and a communication device 883. Note that the hardware configuration shown here is an example, and some of the components may be omitted. Furthermore, the information processing device 90 may include further components in addition to the components shown here.
[0123] (Processor 871) The processor 871 functions, for example, as an arithmetic processing device or control device, and controls the overall operation of each component or part of it based on various programs recorded in the ROM 872, RAM 873, storage 880, or removable storage medium 901.
[0124] (ROM 872, RAM 873) The ROM 872 is a means for storing programs to be read into the processor 871, data to be used for calculations, etc. The RAM 873 temporarily or permanently stores, for example, the programs to be read into the processor 871 and various parameters that change as appropriate when the programs are executed.
[0125] (Host bus 874, bridge 875, external bus 876, interface 877) The processor 871, ROM 872, and RAM 873 are connected to one another via, for example, a host bus 874 that is capable of high-speed data transmission. On the other hand, the host bus 874 is connected to, for example, an external bus 876 that has a relatively low data transmission speed via a bridge 875. Furthermore, the external bus 876 is connected to various components via an interface 877.
[0126] (Input Device 878) For example, a mouse, keyboard, touch panel, button, switch, lever, etc. are used as the input device 878. Furthermore, a remote controller (hereinafter referred to as a remote control) capable of transmitting control signals using infrared rays or other radio waves may also be used as the input device 878. The input device 878 also includes an audio input device such as a microphone.
[0127] (Output Device 879) The output device 879 is a device capable of visually or audibly notifying the user of acquired information, such as a display device such as a CRT (Cathode Ray Tube), LCD, or organic EL, an audio output device such as a speaker or headphones, a printer, a mobile phone, or a facsimile. The output device 879 according to the present disclosure also includes various vibration devices capable of outputting tactile stimulation.
[0128] (Storage 880) The storage 880 is a device for storing various types of data. For example, a magnetic storage device such as a hard disk drive (HDD), a semiconductor storage device, an optical storage device, or a magneto-optical storage device may be used as the storage 880.
[0129] (Drive 881) The drive 881 is a device that reads information recorded on a removable storage medium 901 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, or writes information to the removable storage medium 901.
[0130] (Removable storage medium 901) The removable storage medium 901 is, for example, a DVD medium, a Blu-ray (registered trademark) medium, an HD DVD medium, various semiconductor storage media, etc. Of course, the removable storage medium 901 may also be, for example, an IC card equipped with a contactless IC chip, an electronic device, etc.
[0131] (Connection Port 882) The connection port 882 is a port for connecting an external device 902, such as a USB (Universal Serial Bus) port, an IEEE 1394 port, a SCSI (Small Computer System Interface), an RS-232C port, or an optical audio terminal.
[0132] (Externally Connected Device 902) The externally connected device 902 is, for example, a printer, a portable music player, a digital camera, a digital video camera, or an IC recorder.
[0133] (Communication device 883) The communication device 883 is a communication device for connecting to a network, such as a communication card for wired or wireless LAN, Bluetooth (registered trademark), or WUSB (Wireless USB), a router for optical communication, a router for ADSL (Asymmetric Digital Subscriber Line), or a modem for various types of communication.
[0134] 3. Summary As described above, the information processing device 10 according to an embodiment of the present disclosure includes the second task processing unit 120 including a first neural network that receives as input first sensing data related to an utterance and executes a first task corresponding to the utterance, the first task result being output from the first neural network, and the second sensing data related to the utterance and executes a second task corresponding to the utterance. Furthermore, one of the features of the information processing device 10 is that the first sensing data and the second sensing data are acquired by sensors of different types.
[0135] According to the above configuration, it is possible to perform processing based on utterances using smaller computational resources.
[0136] Although the preferred embodiments of the present disclosure have been described in detail above with reference to the accompanying drawings, the technical scope of the present disclosure is not limited to such examples. It is clear that a person skilled in the art of the present disclosure can conceive of various modified or altered examples within the scope of the technical idea described in the claims, and it is understood that these also naturally fall within the technical scope of the present disclosure.
[0137] Furthermore, the steps of the processes described in this disclosure do not necessarily have to be processed in chronological order according to the order shown in the flowcharts or sequence diagrams. For example, the steps of the processes of each device may be processed in an order different from the order shown, or may be processed in parallel.
[0138] Furthermore, the series of processes performed by each device described in this disclosure may be realized by a program stored in a non-transitory computer-readable storage medium. Each program is, for example, loaded into RAM when executed by a computer and executed by a processor such as a CPU. The storage medium may be, for example, a magnetic disk, an optical disk, a magneto-optical disk, or a flash memory. The program may also be distributed, for example, via a network, without using a storage medium.
[0139] Furthermore, the effects described herein are merely descriptive or exemplary and are not limiting. That is, the technology according to the present disclosure may achieve other effects that will be apparent to those skilled in the art from the description of this specification, in addition to or in place of the above-described effects.
[0140] Note that the following configurations also fall within the technical scope of the present disclosure. (1) An information processing device comprising: a second task processing unit including a second neural network that receives as input first sensing data related to an utterance and executes a first task corresponding to the utterance, the second task processing unit receiving as input second sensing data related to the utterance and a result of the first task output from the first neural network, the second task processing unit executing a second task corresponding to the utterance, the first sensing data and the second sensing data being acquired by sensors of different types. (2) The information processing device described in (1), wherein the first task is either extracting consonants corresponding to the utterance or extracting vowels corresponding to the utterance, and the second task processing unit receives as input the result of the first task and the second sensing data and estimates, as the second task, text data including consonants and vowels corresponding to the utterance. (3) The information processing device according to (1), wherein the first task is extraction of consonants corresponding to the utterance, and the second task processing unit receives a result of the first task and the second sensing data as input, and estimates, as the second task, text data including consonants and vowels corresponding to the utterance. (4) The information processing device according to any one of (1) to (3), wherein the first sensing data is audio data related to the utterance. (5) The information processing device according to (4), wherein the second sensing data includes acceleration data collected during the same period as the audio data related to the utterance. (6) The information processing device according to (5), wherein the acceleration data is collected by a device worn on the head of a speaker who makes the utterance. (7) The information processing device according to (6), wherein the device is worn on the head of the speaker. (8) The information processing device according to (7), wherein the device is an earpiece. (9) The information processing device according to (4), wherein the second sensing data includes image data of a mouth of a speaker making the utterance, the image data being collected during the same period as the voice data relating to the utterance. (10) The information processing device according to (3), wherein the first neural network performs supervised learning based on consonant text data estimated based on the output result of the first task and correct consonant text data.(11) The information processing device according to (3), wherein the second neural network performs supervised learning based on the output result of the second task and correct answer text data. (12) The information processing device according to any one of (1) to (11), comprising: a first task processing unit including the first neural network. (13) The information processing device according to (4), comprising: a microphone that collects audio data related to the utterance. (14) An information processing method including: a processor inputting first sensing data related to an utterance and executing a first task corresponding to the utterance, the first task result output from a first neural network and second sensing data related to the utterance into a second neural network, and executing a second task corresponding to the utterance, wherein the first sensing data and the second sensing data are acquired by sensors of different types. (15) A program that causes a computer to function as an information processing device, comprising: a second task processing unit including a second neural network that receives as input first sensing data related to an utterance, the first neural network receiving as input first sensing data related to the utterance and executing a first task corresponding to the utterance, the second neural network receiving as input second sensing data related to the utterance and executing a second task corresponding to the utterance; and
[0141] REFERENCE SIGNS LIST 10 Information processing device 110 First task processing unit 111 DSP 113 Audio neural network 120 Second task processing unit 121 NPU 122 Fusion neural network 150 Audio output unit 151 Speaker 160 Sensor unit 161 Microphone 162 IMU 163 Camera 164 Pulse wave sensor
Claims
1. An information processing device comprising: a second task processing unit including a second neural network that receives as input first sensing data related to an utterance and executes a first task corresponding to the utterance, the second task processing unit outputting a result of the first task from a first neural network that receives as input second sensing data related to the utterance and executes a second task corresponding to the utterance, wherein the first sensing data and the second sensing data are acquired by sensors of different types.
2. The information processing device of claim 1, wherein the first task is either extracting consonants corresponding to the utterance or extracting vowels corresponding to the utterance, and the second task processing unit receives the result of the first task and the second sensing data as input, and estimates, as the second task, text data including consonants and vowels corresponding to the utterance.
3. The information processing device of claim 1, wherein the first task is to extract consonants corresponding to the utterance, and the second task processing unit receives the result of the first task and the second sensing data as input, and estimates, as the second task, text data including consonants and vowels corresponding to the utterance.
4. The information processing device according to claim 1, wherein the first sensing data is voice data relating to the utterance.
5. The information processing device according to claim 4, wherein the second sensing data includes acceleration data collected during the same period as the voice data relating to the utterance.
6. The information processing device according to claim 5, wherein the acceleration data is collected by a device worn on the head of a speaker making the utterance.
7. The information processing device according to claim 6, which is a device worn on the head of the speaker.
8. The information processing device according to claim 7, which is an earrable device.
9. The information processing device according to claim 4, wherein the second sensing data includes image data of the mouth of the speaker making the utterance, collected during the same period as the voice data relating to the utterance.
10. The information processing device according to claim 3, wherein the first neural network performs supervised learning based on consonant text data estimated based on the result of the first task to be output and correct consonant text data.
11. The information processing device according to claim 3, wherein the second neural network performs supervised learning based on the output result of the second task and correct answer text data.
12. The information processing device according to claim 1, further comprising: a first task processing unit including the first neural network.
13. The information processing device according to claim 4, further comprising: a microphone for collecting voice data relating to the speech.
14. An information processing method comprising: a processor inputting first sensing data related to an utterance and executing a first task corresponding to the utterance from a first neural network, the first task result being output from the first neural network, and inputting second sensing data related to the utterance into a second neural network, and executing a second task corresponding to the utterance; wherein the first sensing data and the second sensing data are obtained by sensors of different types.
15. A program that causes a computer to function as an information processing device, comprising: a second task processing unit that receives as input first sensing data related to an utterance, the first task processing unit outputting a result of the first task from a first neural network that executes a first task corresponding to the utterance, and second sensing data related to the utterance, the second task processing unit executing a second task corresponding to the utterance, the first sensing data and the second sensing data being acquired by sensors of different types.
Citation Information
Patent Citations
Learning apparatus, learning method, program, learnt model and lip reading apparatus
JP2019204147A
Lip reading device and lip reading method
JP2021086274A
Nasal and oral respiration sensor
JP2023119038A