Information processing apparatus, information processing method, and program
By calibrating the utterance detection function and adjusting preprocessing parameters based on sensor signals, the system addresses the challenge of accurately detecting utterances across individual differences and wearing styles, thereby improving detection accuracy.
Patent Information
- Application Number
- PCT/JP2024/040765
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-05
- Filing Date
- 2024-11-18
- Publication Date
- 2025-06-12
AI Technical Summary
Existing technologies face challenges in accurately detecting utterances by wearers of devices such as earphones, due to individual differences and variations in wearing styles, which affect the signal-to-noise ratio of acceleration signals.
The proposed solution involves calibrating the utterance detection function based on sensor signals from sensors that detect physical amounts related to utterances, and adjusting parameters for preprocessing to optimize detection accuracy.
This approach improves the detection rate of utterances by adjusting for individual differences and wearing styles, enhancing the signal-to-noise ratio and overall accuracy of utterance detection.
Smart Images

Figure JP2024040765_12062025_PF_FP_ABST
Abstract
Description
INFORMATION PROCESSING APPARATUS, INFORMATION PROCESSING METHOD, AND PROGRAM
[0001] The present technology relates to an information processing apparatus, an information processing method, and a program, and more particularly, to an information processing apparatus, an information processing method, and a program capable of improving a detection rate of utterance by a wearer of a device.
[0002] <CROSS REFERENCE TO RELATED APPLICATIONS> This application claims the benefit of Japanese Priority Patent Application JP 2023-205061 filed on December 5, 2023, the entire contents of which are incorporated herein by reference.
[0003] Many technologies have been developed to improve user experience (UX) of a device to be worn on the ears by a user, such as earphones (inner ear headphones), a true wireless stereo (TWS), and a hearing aid. PTL 1 describes a technique for notifying a wearer of a wearing state favorable for improving sound quality to be reproduced from earphones. In addition, there is an increasing demand for detecting an environment with a device to improve UX of the device.
[0004] For example, if the wearer of the earphones utters, the earphones detect the utterance of the wearer, mute music that is being reproduced, or transition to a mode for capturing external sound. Even if the wearer does not control a smartphone, or the like, the earphones execute various functions seamlessly according to presence or absence of utterance of the wearer, so that the wearer can have a conversation with a person in front of the wearer, for example, while wearing the earphones.
[0005] [PTL 1]JP 2020-150320A
[0006] Utterance of a wearer is detected on the basis of, for example, a sound signal acquired by a microphone (microphone) mounted on the earphones or a sensor signal acquired by a sensor. In order to accurately detect utterance of the wearer, it is important to increase a signal-to-noise ratio (S / N) of a signal to be acquired by the microphone or the sensor.
[0007] For example, in a case where vibration generated by the utterance of the wearer is detected by an acceleration sensor, the vibration propagates in the head and reaches the earphones, and thus, an S / N of an acceleration signal greatly changes depending on individual differences and a way of wearing the earphones. By determining a parameter to be used for utterance detection on the basis of an average value of acceleration signals for each individual, a detection rate of utterance by an average wearer can be improved, but a detection rate of utterance by a wearer other than the average wearer decreases.
[0008] With the technique described in PTL 1, it is not possible to improve an S / N of an acceleration signal indicating a detection result of vibration generated by utterance of a wearer, propagated in the head, and reached the earphones.
[0009] The present technology has been made in view of such circumstances, and it is desirable to improve a detection rate of utterance by a wearer of a device.
[0010] A system comprising an information processing apparatus according to a first aspect of the present technology includes circuitry configured to perform calibration of an utterance detection function based on a sensor signal acquired by a sensor that detects a physical amount related to an utterance by a wearer of a device, and a detection result of presence or absence of the utterance by the utterance detection function that detects presence or absence of the utterance based on the sensor signal.
[0011] An information processing method according to a first aspect of the present technology is an information processing method to be performed by a system comprising an information processing apparatus, and includes performing calibration of an utterance detection function based on a sensor signal acquired by a sensor that detects a physical amount related to an utterance by a wearer of a device, and a detection result of presence or absence of the utterance by the utterance detection function that detects presence or absence of the utterance based on the sensor signal.
[0012] At least one non-transitory computer-readable storage medium according to a first aspect of the present technology having instructions encoded thereon, that when executed by circuitry of a system comprising an information processing apparatus, cause the circuitry to perform an information processing method comprising performing calibration of an utterance detection function based on a sensor signal acquired by a sensor that detects a physical amount related to an utterance by a wearer of a device, and a detection result of presence or absence of the utterance by the utterance detection function that detects presence or absence of the utterance based on the sensor signal.
[0013] A system comprising an information processing apparatus according to a second aspect of the present technology includes circuitry configured to determine a manner of wearing a device or a volume of an utterance by a wearer of the device where presence or absence of the utterance is detected based on a sensor signal acquired by a sensor that detects a physical amount related to the utterance of the wearer of the device.
[0014] In the first aspect of the present technology, the calibration of the utterance detection function is performed on the basis of the sensor signal acquired by the sensor that detects the physical amount related to the utterance of the wearer of the device and the detection result of presence or absence of the utterance by the utterance detection function that detects presence or absence of the utterance on the basis of the sensor signal.
[0015] In the second aspect of the present technology, the way of wearing the device or the volume of the utterance when presence or absence of the utterance is detected is determined on the basis of the sensor signal acquired by the sensor that detects the physical amount related to the utterance of the wearer of the device.
[0016] Fig. 1 is a view illustrating a configuration example of an information processing system to which the present technology is applied.Fig. 2 is a view for explaining flow of utterance detection.Fig. 3 is a view illustrating a configuration example of an ear device.Fig. 4 is a block diagram illustrating a functional configuration example of the ear device in a case where utterance detection is performed on the basis of an acceleration signal.Fig. 5 is a view illustrating an example of vibration in a specific direction.Fig. 6 is a block diagram illustrating a functional configuration example of an information processing system according to a first embodiment.Figs. 7A to 7C are views illustrating examples of frequency characteristics of sound signals.Fig. 8 is a flowchart for explaining processing to be performed by the information processing system having the configuration of Fig. 6.Figs. 9A and 9B are views illustrating examples of a method for presenting a guide for recommending an appropriate way of wearing the ear device and an appropriate volume of utterance.Figs. 10A to 10C are views illustrating examples of a graph for recommending an appropriate way of wearing the ear device and an appropriate volume of utterance.Fig. 11 is a view illustrating a configuration example of the ear device on which a microphone is not mounted.Fig. 12 is a view for explaining flow of optimizing a parameter of preprocessing.Fig. 13 is a block diagram illustrating a functional configuration example of an information processing system according to a second embodiment.Fig. 14 is a flowchart for explaining processing to be performed by the information processing system having the configuration of Fig. 13.Fig. 15 is a block diagram illustrating a functional configuration example of the information processing system in a case where the ear device optimizes the parameter of the preprocessing.Fig. 16 is a block diagram illustrating a functional configuration example of the information processing system in a case where an entire utterance detection function is optimized by a cloud.Fig. 17 is a flowchart for explaining processing to be performed by the information processing system having the configuration of Fig. 16.Fig. 18 is a flowchart for explaining processing of performing calibration of the utterance detection function during a call.Fig. 19 is a block diagram illustrating a configuration example of hardware of a computer.
[0017] Hereinafter, a mode for carrying out the present technology will be described. The description will be provided in the following order. 1. Utterance detection by ear device 2. First embodiment (example of recommending an appropriate way of wearing ear device) 3. Second embodiment (example of optimizing a parameter of utterance detection)
[0018] <1. Utterance detection by ear device> Fig. 1 is a view illustrating a configuration example of an information processing system to which the present technology is applied.
[0019] The information processing system in Fig. 1 includes TWS units 1L and 1R to be respectively worn on the left and right ears, and an external terminal 2.
[0020] The TWS units 1L and 1R are an example of a wearable device to be worn on the ears by the user, and hereinafter, such a wearable device will be also referred to as an ear device 1. The ear device 1 includes earphones, a hearing aid, and the like, in addition to the TWS. The TWS units 1L and 1R perform wireless communication using, for example, Bluetooth low energy (BLE) with the external terminal 2 in a state of being paired with each other and output sound of music supplied from the external terminal 2.
[0021] The external terminal 2 includes a smartphone, a tablet terminal, a PC, and the like. The external terminal 2 reproduces music data and supplies the resultant sound of the music to the TWS units 1L and 1R.
[0022] The information processing system of Fig. 1 has an utterance detection function of detecting presence or absence of utterance by a wearer of the ear device 1.
[0023] Fig. 2 is a view for explaining flow of utterance detection.
[0024] If the wearer utters, a sensor mounted on the ear device 1 detects a physical amount related to the utterance and acquires a sensor signal. Specifically, as illustrated in #1 of Fig. 2, an acceleration sensor mounted on the ear device 1 detects vibration (utterance vibration) generated by the utterance and acquires an acceleration signal, and a microphone collects sound of the utterance and acquires a sound signal.
[0025] Next, for example, an arithmetic block in a digital signal processor (DSP) or a central processing unit (CPU) mounted on the ear device 1 performs utterance detection processing on the basis of the acceleration signal or the sound signal acquired by the acceleration sensor or the microphone, as illustrated in #2 of Fig. 2. For example, by using a learning model, presence or absence of utterance by the wearer of the ear device 1 is detected.
[0026] Finally, the ear device 1 executes various functions on the basis of the detection result of the presence or absence of utterance. For example, the ear device 1 mutes the music that is being reproduced or transitions to a mode of capturing external sound.
[0027] Fig. 3 is a view illustrating a configuration example of the ear device 1.
[0028] As illustrated in Fig. 3, the ear device 1 includes a driver 11, an outer microphone 12A, an inner microphone 12B, an acceleration sensor 14 provided on a substrate 13, and a CPU / DSP 15.
[0029] The driver 11 outputs sound of music supplied from the external terminal 2.
[0030] The outer microphone 12A is mounted on an outer side of a housing of the ear device 1 (side opposite to the ear side in the ear device 1) and collects sound of utterance by the wearer and external sound to acquire a sound signal. The inner microphone 12B is mounted on an inner side of the housing of the ear device 1 (ear side of the ear device 1) and collects the sound of the utterance by the wearer propagated in the head to acquire a sound signal. Note that, hereinafter, in a case where it is not necessary to distinguish between the outer microphone 12A and the inner microphone 12B, they are simply referred to as microphones 12.
[0031] The acceleration sensor 14 detects utterance vibration that has propagated in the head of the wearer and reached the ear device 1, and acquires an acceleration signal.
[0032] The CPU / DSP performs utterance detection processing on the basis of the sound signal acquired by the microphone 12, the acceleration signal acquired by the acceleration sensor 14, and the like, and executes various functions on the basis of a detection result of the presence or absence of utterance.
[0033] In general, in earphones, a TWS, a hearing aid, or the like, utterance detection is often performed on the basis of the acceleration signal acquired by the acceleration sensor, and in a headphone, or the like, utterance detection is often performed on the basis of the sound signal acquired by the microphone.
[0034] Fig. 4 is a block diagram illustrating a functional configuration example of the ear device 1 in a case where the utterance detection is performed on the basis of the acceleration signal.
[0035] As illustrated in Fig. 4, the ear device 1 includes the acceleration sensor 14, a preprocessing unit 31, a detector 32, and a function execution unit 33. The preprocessing unit 31, the detector 32, and the function execution unit 33 illustrated in Fig. 4 are implemented, for example, by execution of a predetermined program by the DSP / CPU 15 in Fig. 2. Note that Fig. 4 illustrates components related to utterance detection among the components of the ear device 1.
[0036] If a wearer U1 of the ear device 1 utters, utterance vibration is input to the acceleration sensor 14 through the inside of the head of the wearer. The acceleration sensor 14 detects, for example, acceleration in three axial directions of an x axis, a y axis, and a z axis, and acquires a three-axis signal indicating acceleration in the three axial directions. The acceleration sensor 14 supplies the three-axis signal to the preprocessing unit 31.
[0037] The preprocessing unit 31 performs preprocessing on the three-axis signal. Specifically, the preprocessing unit 31 performs preprocessing on the three-axis signal by extracting vibration information indicating vibration in a specific direction from the three-axis signal. The vibration information is information to be input to a learning model.
[0038] Fig. 5 is a view illustrating an example of the vibration in the specific direction.
[0039] As illustrated in Fig. 5, it is assumed that the acceleration sensor 14 acquires an acceleration signal x1 in the x-axis direction, an acceleration signal y1 in the y-axis direction, and an acceleration signal z1 in the z-axis direction. As the preprocessing, the preprocessing unit 31 extracts (generates), as vibration information, an acceleration signal p1 corresponding to a principal component direction of the utterance vibration propagated through the head and input to the acceleration sensor 14 by weighting and synthesizing the acceleration signals.
[0040] Returning to Fig. 4, the preprocessing unit 31 supplies the vibration information in the specific direction to the detector 32.
[0041] The detector 32 detects presence or absence of the utterance of the wearer by inputting the vibration information supplied from the preprocessing unit 31 to the learning model generated by machine learning. Specifically, the detector 32 acquires a probability (utterance probability) that the vibration input to the acceleration sensor 14 is utterance vibration and determines that the wearer has uttered in a case where the utterance probability is greater than a predetermined threshold. The detector 32 supplies a detection result of the presence or absence of the utterance to the function execution unit 33.
[0042] As described above, the utterance detection function of the ear device 1 is implemented by the preprocessing unit 31 and the detector 32.
[0043] The function execution unit 33 executes various functions according to the detection result of the presence or absence of utterance by the detector 32.
[0044] Here, in order to improve accuracy of utterance detection by the detector 32, quality of the acceleration signal (three-axis signal) to be acquired by the acceleration sensor 14 and quality of the vibration information to be extracted by the preprocessing unit 31 are important.
[0045] For example, in a case where the ear device 1 is not appropriately worn and the ears and the ear device 1 are not in close contact with each other, utterance vibration is not sufficiently input to the acceleration sensor 14, and a ratio of a noise component included in the acceleration signal increases.
[0046] It is known that, when vibration having acceleration in a certain direction is sensed, vibration information with a high S / N can be obtained by appropriately weighting the three-axis signal. The three-axis signal also includes a noise component, and thus, if the three-axis signal acquired by the acceleration sensor 14 is directly synthesized to generate vibration information, the ratio of the noise component included in the vibration information increases.
[0047] For example, in a case where vibration having acceleration only in the z-axis direction is input to the acceleration sensor 14, the acceleration signal in the x-axis direction and the acceleration signal in the y-axis direction do not include a vibration component but include only a noise component. If the three-axis signals are synthesized as they are, the noise component included in the vibration information is increased by the acceleration signals in the x-axis direction and the y-axis direction, and the S / N of the vibration information decreases. In this case, by generating the vibration information using only the acceleration signal in the z-axis direction without using the acceleration signals in the x-axis direction and the y-axis direction, it is possible to acquire the vibration information having a higher S / N than in a case where the vibration information is generated using the three-axis signal.
[0048] If utterance vibration actually detected by an acceleration sensor mounted on the earphones, or the like, is analyzed, it is known that there are individual differences and wearing errors in the principal component direction of the utterance vibration. By determining a parameter (weight) of the preprocessing on the basis of an average value of the principal component direction of the utterance vibration for each individual, a detection rate of utterance by an average wearer can be improved, but a detection rate of utterance by a wearer other than the average wearer decreases.
[0049] A first embodiment of the present technology has been conceived focusing on the above points, and proposes a technology capable of improving a detection rate of utterance by a wearer by recommending an appropriate way of wearing the ear device to the wearer. Furthermore, a second embodiment proposes a technology capable of improving a detection rate of utterance by a wearer by calibrating the parameter of preprocessing for each individual wearer. Hereinafter, the first embodiment and the second embodiment will be described in detail.
[0050] <2. First embodiment (example of recommending an appropriate way of wearing ear device)> Fig. 6 is a block diagram illustrating a functional configuration example of an information processing system according to the first embodiment.
[0051] The information processing system in Fig. 6 causes the wearer to manually calibrate the utterance detection function. Specifically, the information processing system causes the wearer to perform calibration by recommending an appropriate way of wearing the ear device or an appropriate volume of utterance to the wearer. For example, the information processing system recommends, as an appropriate way of wearing the ear device, a way such that utterance vibration in a direction matching a recommended direction (recommended direction) is input to the acceleration sensor 14. Note that Fig. 4 illustrates components related to recommendation of the way of wearing the ear device among the components of the ear device 1 and the external terminal 2.
[0052] In Fig. 6, the same components as those in Fig. 4 are denoted by the same reference numerals. Redundant description will be omitted as appropriate. The ear device 1 of Fig. 6 is different from the ear device 1 of Fig. 4 in that the outer microphone 12A, the inner microphone 12B, a sound information analysis unit 51, and a recommendation determination unit 52 are provided. The sound information analysis unit 51 and the recommendation determination unit 52 illustrated in Fig. 6 are implemented, for example, by execution of a predetermined program by the DSP / CPU 15 in Fig. 2.
[0053] If the wearer U1 of the ear device 1 utters, a sound wave of the utterance propagates, for example, in the air and is input to the outer microphone 12A and propagates, for example, in the head and is input to the inner microphone 12B. The outer microphone 12A and the inner microphone 12B detect the input sound wave to acquire a sound signal and supply the sound signal to the sound information analysis unit 51.
[0054] The sound information analysis unit 51 analyzes the sound signal supplied from the outer microphone 12A and the sound signal supplied from the inner microphone 12B. Specifically, the sound information analysis unit 51 analyzes a volume of the utterance of the wearer on the basis of the sound signal acquired by outer microphone 12A. In addition, the sound information analysis unit 51 analyzes frequency characteristics of the sound signals acquired by the outer microphone 12A and the inner microphone 12B.
[0055] Figs. 7A to 7C are views illustrating examples of the frequency characteristics of the sound signals. Figs. 7A to 7C indicate a frequency on a horizontal axis and indicate a sound pressure on a vertical axis.
[0056] Fig. 7A illustrates an example of the frequency characteristics of the sound signal acquired by the outer microphone 12A. The sound wave emitted from the mouth of the wearer propagates and is input to the outer microphone 12A, and thus, the sound signal acquired by the outer microphone 12A becomes a broadband signal as illustrated in Fig. 7A.
[0057] Fig. 7B illustrates an example of the frequency characteristics of the sound signal acquired by the inner microphone 12B in a case where the ear device 1 is appropriately worn. In a case where the ear device 1 is appropriately worn, the inner microphone 12B is arranged in an acoustic condition sealed by the housing of the ear device 1 and the ear canal, and thus, in the inner microphone 12B, energy of the sound wave propagated in the air and input decreases, and the sound wave propagated through the head is also input. Thus, as illustrated in Fig. 7B, the sound signal acquired by the inner microphone 12B becomes a narrower band signal than the sound signal acquired by the outer microphone 12A.
[0058] Fig. 7C illustrates an example of the frequency characteristics of the sound signal acquired by the inner microphone 12B in a case where the ear device 1 is not appropriately worn. In a case where the ear device 1 is not appropriately worn, the ear canal is not sealed by the housing of the ear device 1, so that the sound wave is easily input to the inner microphone 12B by propagating in the air. Thus, as illustrated in Fig. 7C, the sound signal acquired by the inner microphone 12B becomes a signal in a band similar to the band of the sound signal acquired by the outer microphone 12A.
[0059] As described above, the frequency characteristics of the sound signal acquired by the inner microphone 12B differ depending on whether or not the ear device 1 is appropriately worn. The ear device 1 can determine whether or not the ear device 1 is appropriately worn by comparing the frequency characteristics of the sound signals acquired by the outer microphone 12A and the inner microphone 12B.
[0060] Returning to Fig. 6, the sound information analysis unit 51 supplies an analysis result of the sound signal to the recommendation determination unit 52 as sound information.
[0061] To the recommendation determination unit 52, the vibration information in the specific direction is supplied from the preprocessing unit 31, and the detection result of the presence or absence of utterance is supplied from the detector 32. The recommendation determination unit 52 recommends an appropriate way of wearing the ear device and an appropriate volume of utterance according to the detection result of the presence or absence of utterance.
[0062] In a case where utterance is detected, the recommendation determination unit 52 presents, for example, that the utterance is normally detected to the wearer via the external terminal 2.
[0063] In a case where utterance is not detected, the recommendation determination unit 52 functions as a determination unit that determines an appropriate way of wearing the ear device 1 and an appropriate volume of the utterance on the basis of the vibration information in the specific direction and the analysis result of the sound signal in order to improve the detection rate of the utterance. The recommendation determination unit 52 presents a guide indicating the determined way of wearing the ear device 1 and the determined volume of the utterance to the wearer via the external terminal 2. For example, the recommendation determination unit 52 urges the wearer to change the way of wearing the ear device 1 or increase the volume of the utterance by presenting a guide by a message, a graph, a voice, or the like.
[0064] Note that a guide for recommending an appropriate way of wearing the ear device and an appropriate volume of utterance is created by the recommendation determination unit 52 or the external terminal 2 on the basis of the vibration information in the specific direction, the detection result of the presence or absence of utterance, and the analysis result of the sound signal, and is presented via the external terminal 2.
[0065] The external terminal 2 includes a UI control unit 61. The UI control unit 61 controls a user interface (UI) related to calibration of the utterance detection function. For example, in a case where the wearer of the ear device 1 gives an instruction to start calibration through predetermined operation on the UI, the UI control unit 61 controls the ear device 1 to start calibration of the utterance detection function.
[0066] Furthermore, for example, in the middle of the calibration, the UI control unit 61 presents a guide indicating an appropriate way of wearing the ear device and an appropriate volume of the utterance to the wearer of the ear device 1 according to the detection result of the presence or absence of utterance by the detector 32.
[0067] Next, processing to be performed by the information processing system having the configuration of Fig. 6 will be described with reference to the flowchart of Fig. 8. The processing of Fig. 8 is started, for example, in a case where the wearer of the ear device 1 gives an instruction to start calibration of the utterance detection function.
[0068] In step S1, the UI control unit 61 of the external terminal 2 presents a guide to urge the wearer of the ear device 1 to utter. The wearer utters with a volume desired to be detected as utterance according to the guide.
[0069] In step S2, the acceleration sensor 14 of the ear device 1 detects utterance vibration and acquires a three-axis signal. The preprocessing unit 31 of the ear device 1 extracts vibration information in a specific direction from the three-axis signal.
[0070] In step S3, the detector 32 of the ear device 1 performs utterance detection processing on the basis of the vibration information in the specific direction extracted from the three-axis signal.
[0071] In step S4, the recommendation determination unit 52 of the ear device 1 determines whether or not utterance has been detected by the detector 32.
[0072] In a case where it is determined in step S4 that utterance has been detected, it is considered that there is no problem regarding the way of wearing the ear device 1 and the volume of the utterance, and thus, the recommendation determination unit 52 completes the calibration of the utterance detection function in step S5. After the calibration of the utterance detection function is completed, the processing ends.
[0073] On the other hand, in a case where it is determined in step S4 that no utterance is detected, in step S6, the sound information analysis unit 51 of the ear device 1 analyzes the sound signal acquired by each of the outer microphone 12A and the inner microphone 12B.
[0074] In step S7, the recommendation determination unit 52 determines whether or not the volume of the utterance is larger than a threshold.
[0075] In a case where it is determined in step S7 that the volume of the utterance is larger than the threshold, in step S8, the recommendation determination unit 52 compares frequency characteristics of the sound signal acquired by the outer microphone 12A with frequency characteristics of the sound signal acquired by the inner microphone 12B.
[0076] In step S9, the recommendation determination unit 52 specifies how to change the way of wearing the ear device 1 so that the utterance can be detected, on the basis of the comparison result of the frequency characteristics and recommends an appropriate way of wearing the ear device to the wearer. In a case where the volume of the utterance is larger than the threshold, it is considered that there is no problem with the volume of the utterance, and thus, the recommendation determination unit 52 recommends the way of wearing the ear device.
[0077] On the other hand, in a case where it is determined in step S7 that the volume of the utterance is smaller than the threshold, the recommendation determination unit 52 recommends an appropriate volume of the utterance in step S10. In a case where the volume of the utterance is smaller than the threshold, it is considered that there is a problem regarding the volume of the utterance, and thus, the recommendation determination unit 52 recommends the volume of the utterance.
[0078] After the appropriate way of wearing the ear device and the appropriate volume of the utterance are recommended in step S9 or step S10, the UI control unit 61 determines whether or not to perform calibration again in step S11. For example, the UI control unit 61 displays a button asking whether or not to perform calibration again. In a case where the wearer of the ear device 1 selects to perform the calibration again, the UI control unit 61 determines to perform the calibration again.
[0079] In a case where it is determined in step S11 to perform the calibration again, the processing returns to step S1, and the subsequent processing is performed. By repeatedly performing the processing of steps S1 to S10, the wearer can improve the volume of the utterance and the way of wearing the ear device and increase the S / N of the three-axis signal acquired by the acceleration sensor.
[0080] On the other hand, in a case where it is determined in step S11 not to perform the calibration again, the processing ends.
[0081] As described above, in the information processing system according to the first embodiment of the present technology, the way of wearing the ear device 1 or the volume of the utterance when detecting presence or absence of utterance is determined on the basis of the acceleration signal acquired by the acceleration sensor or the sound signal acquired by the microphone 12. With such processing, the information processing system can find a volume problem and a wearing problem that greatly affect utterance detection, feed back the volume problem and the wearing problem to the wearer of the ear device 1, and recommend an appropriate volume of utterance and an appropriate way of wearing the ear device. As a result of the volume of the utterance and the way of wearing the ear device being improved, the S / N of the three-axis signal acquired by the acceleration sensor 14 increases, and the detection rate of the utterance can be improved.
[0082] The information processing system performs processing while separating the problem into the volume problem and the wearing problem, such as determining whether the utterance is not detected due to the volume of the utterance using the outer microphone 12A and determining whether the utterance is not detected due to the way of wearing the ear device 1 using the inner microphone 12B. As a result, the information processing system can accurately grasp the volume problem and the wearing problem.
[0083] Figs. 9A and 9B are views illustrating examples of a method for presenting a guide for recommending an appropriate way of wearing the ear device and an appropriate volume of utterance.
[0084] As illustrated in Fig. 9A, for example, a message for recommending an appropriate way of wearing the ear device or an appropriate volume of utterance is displayed on an upper portion of a display of the external terminal 2, and a graph G1 is displayed below the message.
[0085] As illustrated in Fig. 9B, for example, a speech such as “Please make the volume of voice a little larger” and “Please rotate the housing by 30 degrees” is output from a speaker of the external terminal 2. Note that the speech for recommending an appropriate way of wearing the ear device or an appropriate volume of utterance may be output from the ear device 1 instead of the speaker of the external terminal 2.
[0086] Figs. 10A to 10C are views illustrating examples of a graph for recommending an appropriate way of wearing the ear device and an appropriate volume of utterance.
[0087] As illustrated in Fig. 10A, arrows indicating a direction of vibration actually input to the acceleration sensor 14 (a specific direction indicated by the vibration information generated by the preprocessing) and a recommended direction are superimposed on an illustration of the headphone and displayed in real time, for example. In the example of Fig. 10A, the direction of the vibration actually input to the acceleration sensor 14 is indicated by a solid arrow, and the recommended direction is indicated by a dotted arrow.
[0088] For example, the wearer of the ear device 1 rotates the ear device, for example, so that the solid arrow matches the dotted arrow while viewing the graph illustrated in Fig. 10A.
[0089] As illustrated in Fig. 10B, a bar graph indicating an utterance probability (probability that vibration input to the acceleration sensor 14 is vibration due to voice) is displayed in real time. Furthermore, a threshold serving as a reference for detecting presence or absence of utterance is displayed.
[0090] For example, the wearer of the ear device 1 improves the volume of the utterance and the way of wearing the ear device 1 so that the utterance probability indicated by the graph illustrated in Fig. 10B exceeds the threshold.
[0091] As illustrated in Fig. 10C, a line graph indicating a sound pressure level of utterance at each time (time) is displayed in real time. Furthermore, a band-shaped region indicating a sound pressure level recommended for detecting utterance is displayed.
[0092] For example, the wearer of the ear device 1 checks how much the sound pressure level (volume) of the utterance should be increased by looking at the line graph illustrated in Fig. 10C and improves the volume of the utterance.
[0093] Note that the first embodiment of the present technology can also be applied to the ear device 1 on which a microphone is not mounted.
[0094] Fig. 11 is a view illustrating a configuration example of the ear device 1 on which a microphone is not mounted.
[0095] In Fig. 11, the same components as those in Fig. 3 are denoted by the same reference numerals. Redundant description will be omitted as appropriate. The ear device 1 of Fig. 11 is different from the ear device 1 of Fig. 3 in that the outer microphone 12A and the inner microphone 12B are not provided.
[0096] In a case where a microphone is not provided in the ear device 1, the processing flow of the information processing system is simplified, and only the recommendation based on the three-axis signal acquired by the acceleration sensor 14 is performed according to the detection result of the presence or absence of utterance during the calibration. For example, the graph as illustrated in Fig. 10A and a message regarding the wearing angle of the ear device 1 are displayed.
[0097] <3. Second embodiment (example of optimizing parameter of utterance detection)> The information processing system according to the second embodiment automatically calibrates the utterance detection function by adjusting and optimizing the parameter of the preprocessing for each individual wearer.
[0098] Fig. 12 is a view for explaining flow of optimization of the parameter of the preprocessing.
[0099] First, the external terminal 2 urges the wearer of the ear device 1 to utter. As illustrated in a right portion of Fig. 12, if the wearer of the ear device 1 utters, vibration and a sound wave generated by the utterance are input to the ear device 1.
[0100] Next, the ear device 1 transmits the vibration, an acceleration signal obtained by detecting the sound wave, and a sound signal to the external terminal 2 as illustrated in a left portion of Fig. 12. Next, the external terminal 2 analyzes the acceleration signal and the sound signal transmitted from the ear device 1 and optimizes the parameter to be used in the preprocessing of utterance detection on the basis of the analysis result.
[0101] Finally, the external terminal 2 transmits the optimized parameter to the ear device 1 and updates the parameter to be used in the preprocessing by the preprocessing unit 31 of the ear device 1.
[0102] As described above, the ear device 1 and the external terminal 2 cooperate with each other, whereby the utterance detection function can be calibrated.
[0103] Fig. 13 is a block diagram illustrating a functional configuration example of the information processing system according to the second embodiment. Note that Fig. 13 illustrates components related to parameter optimization among the components of the ear device 1 and the external terminal 2.
[0104] In Fig. 13, the same components as those in Fig. 6 are denoted by the same reference numerals. Redundant description will be omitted as appropriate. The ear device 1 of Fig. 13 is different from the ear device 1 of Fig. 6 in that the microphone 12 is provided instead of the outer microphone 12A and the inner microphone 12B. Furthermore, the ear device of Fig. 13 is different from the ear device of Fig. 6 in that a transmission unit 101 is provided and the sound information analysis unit 51 and the recommendation determination unit 52 are not provided.
[0105] If the wearer U1 of the ear device 1 utters, a sound wave generated by the utterance propagates, for example, in the air and is input to the microphone 12. The microphone 12 detects the input sound wave to acquire a sound signal and supplies the sound signal to the transmission unit 101.
[0106] To the transmission unit 101, the three-axis signal acquired by the acceleration sensor 14 is supplied. The transmission unit 101 encodes the sound signal acquired by the microphone 12 and the three-axis signal acquired by the acceleration sensor 14 to generate sound data and acceleration data, respectively. The transmission unit 101 transmits the sound data and the acceleration data to the external terminal 2.
[0107] The external terminal 2 in Fig. 13 is different from the external terminal 2 in Fig. 4 in that a reception unit 111, a preprocessing unit 112, a detector 113, a sound information analysis unit 114, and a parameter optimization unit 115 are provided.
[0108] The reception unit 111 receives the acceleration data and the sound data transmitted from the ear device 1 and decodes the acceleration data and the sound signal to generate the three-axis signal and the sound signal. The reception unit 111 supplies the three-axis signal to the preprocessing unit 112 and the parameter optimization unit 115, and supplies the sound signal to the sound information analysis unit 114.
[0109] The preprocessing unit 112 corresponds to the preprocessing unit 31 of the ear device 1, performs preprocessing similar to that of the preprocessing unit 31, and supplies vibration information in a specific direction to the detector 113.
[0110] The detector 113 corresponds to the detector 32 of the ear device 1, performs utterance detection processing, and supplies a detection result of presence or absence of utterance to the parameter optimization unit 115, similarly to the detector 32.
[0111] The sound information analysis unit 114 analyzes the sound signal supplied from the reception unit 111. Specifically, the sound information analysis unit 114 analyzes a volume of the utterance of the wearer on the basis of the sound signal and specifies time (period) at which the wearer is uttering. The sound information analysis unit 114 supplies the analysis result of the sound signal to the parameter optimization unit 115 as sound information.
[0112] The parameter optimization unit 115 searches for a parameter of preprocessing that maximizes an utterance probability on the basis of the three-axis signal supplied from the reception unit 111 and the detection result of the presence or absence of utterance by the detector 113. The three-axis signal to be used for searching for the parameter is, for example, a signal at the time when the wearer is uttering, specified by the sound information analysis unit 114.
[0113] The parameter optimization unit 115 causes the preprocessing unit 112 to perform preprocessing using the searched parameter. The parameter optimization unit 115 searches again for a parameter that maximizes the utterance probability on the basis of the detection result of the presence or absence of utterance performed on the basis of the vibration information in the specific direction extracted by the preprocessing.
[0114] In this manner, the preprocessing, the utterance detection, and the parameter search are repeated to optimize the parameter. For example, in a case where the utterance probability is greater than a predetermined threshold, the parameter optimization unit 115 determines that the optimization of the parameter has been completed, and applies the optimized parameter to the preprocessing in the preprocessing unit 31 of the ear device 1.
[0115] The parameter optimization unit 115 functions as a calibration unit that calibrates the utterance detection function by optimizing the parameter of the preprocessing.
[0116] Next, processing to be performed by the information processing system having the configuration of Fig. 13 will be described with reference to the flowchart of Fig. 14. The processing of Fig. 14 is started, for example, in a case where the wearer of the ear device 1 gives an instruction to start calibration of the utterance detection function.
[0117] In step S51, the UI control unit 61 of the external terminal 2 presents a guide to urge the wearer of the ear device 1 to utter. The wearer utters for a predetermined period according to the guide of the external terminal 2.
[0118] In step S52, the acceleration sensor 14 of the ear device 1 detects utterance vibration and acquires a three-axis signal.
[0119] In step S53, the parameter optimization unit 115 of the external terminal 2 optimizes the parameter of the preprocessing and applies the optimized parameter to the preprocessing in the preprocessing unit 31 of the ear device 1.
[0120] As described above, in the external terminal 2 according to the second embodiment of the present technology, calibration of the utterance detection function (optimization of the parameter of the preprocessing) is performed on the basis of the acceleration signal acquired by the acceleration sensor 14 and the detection result of the presence or absence of utterance by the utterance detection function that detects the presence or absence of utterance on the basis of the acceleration signal. By optimizing the parameter of the preprocessing, the S / N of the vibration information in the specific direction generated by the preprocessing unit 31 is increased, and the detection rate of the utterance can be improved.
[0121] By using a CPU with high performance mounted on the external terminal 2, the parameter of the preprocessing can be optimized at high speed.
[0122] In a case where a processing memory of the ear device 1 is not abundant, if an attempt is made to optimize the parameter by the ear device 1 itself, the wearer has to continue to utter during optimization of the parameter. In the second embodiment of the present technology, the parameter is optimized in the external terminal 2, so that the information processing system can optimize the parameter by the wearer uttering only for a predetermined period without continuing to utter.
[0123] Note that the parameter of the preprocessing may be optimized by the ear device 1.
[0124] Fig. 15 is a block diagram illustrating a functional configuration example of the information processing system in a case where the ear device 1 optimizes the parameter of the preprocessing.
[0125] In Fig. 15, the same components as the components in Fig. 6 are denoted by the same reference signs. Redundant description will be omitted as appropriate. The ear device 1 of Fig. 15 is different from the ear device 1 of Fig. 6 in that the microphone 12 is provided instead of the outer microphone 12A and the inner microphone 12B. Furthermore, the ear device of Fig. 15 is different from the ear device of Fig. 6 in that the parameter optimization unit 151 is provided and the recommendation determination unit 52 is not provided.
[0126] To the parameter optimization unit 151, the three-axis signal acquired by the acceleration sensor 14, the detection result of the presence or absence of utterance by the detector 32, and the analysis result of the sound signal by the sound information analysis unit 51 are supplied.
[0127] The parameter optimization unit 115 searches for a parameter of preprocessing that maximizes an utterance probability on the basis of the three-axis signal, the detection result of the presence or absence of utterance, and the sound information. Specifically, while the wearer U1 of the ear device 1 is uttering, the parameter optimization unit 115 searches for an optimal parameter by controlling the parameter to be used for the preprocessing of the preprocessing unit 31.
[0128] For example, while the wearer U1 of the ear device 1 is uttering, the parameter optimization unit 151 causes the preprocessing unit 112 to perform preprocessing using each of a plurality of parameter candidates. The parameter optimization unit 151 compares utterance probabilities acquired on the basis of the vibration information in the specific direction extracted by the preprocessing and searches for (selects) a parameter with the maximum utterance probability. The parameter optimization unit 151 applies the optimized parameter to the preprocessing in the preprocessing unit 31.
[0129] During the search for the optimal parameter, the wearer U1 of the ear device 1 has to continue to utter, but difficulty of implementation is lower than in a case of performing in cooperation with the external terminal 2. In a case where a memory of the ear device 1 is abundant, the ear device 1 can temporarily store the three-axis signal and the sound signal in the memory and can optimize the parameter of the preprocessing using the stored three-axis signal and sound signal.
[0130] Note that, although the first embodiment and the second embodiment of the present technology can be implemented independently, it is possible to further improve the detection rate of utterance by sequentially integrating recommendation of an appropriate way of wearing the ear device and an appropriate volume of the utterance (first embodiment), and optimization of the parameter of the preprocessing (second embodiment).
[0131] For example, after the appropriate way of wearing the ear device and the appropriate volume of the utterance are recommended and the wearer of the ear device 1 improves the way of wearing the ear device and volume of the utterance, the parameter of the preprocessing is optimized.
[0132] Although the example in which the ear device 1 or the external terminal 2 optimizes the parameter has been described above, the cloud can also optimize the parameter.
[0133] Furthermore, in the above, an example has been described in which the parameter to be used for the preprocessing is optimized. However, the entire utterance detection function to be implemented by the preprocessing by the preprocessing unit 31 and the utterance detection processing by the detector 32 may be optimized by the ear device 1, the external terminal 2, or the cloud. The optimization of the entire utterance detection function includes, for example, adjustment of the parameter of the preprocessing, relearning of a learning model to be used for the utterance detection processing, and selection of a learning model to be used for the utterance detection processing from a plurality of candidates for the learning model prepared in advance.
[0134] Fig. 16 is a block diagram illustrating a functional configuration example of an information processing system in a case where the entire utterance detection function is optimized by the cloud.
[0135] The information processing system in Fig. 16 includes the ear device 1, the external terminal 2, and a cloud 201. The external terminal 2 and the cloud 201 are connected via a network 202.
[0136] In Fig. 16, the same components as those in Fig. 12 are denoted by the same reference numerals. Redundant description will be omitted as appropriate. The ear device 1 of Fig. 16 is different from the ear device 1 of Fig. 12 in that the preprocessing unit 31 is not provided.
[0137] The three-axis signal acquired by the acceleration sensor 14 is supplied to the detector 32. The detector 32 detects presence or absence of utterance of the wearer by inputting the three-axis signal to the learning model generated by machine learning. Here, the learning model using the three-axis signal as an input is a model different from the learning model using the vibration information in the specific direction as an input, and is a model that collectively executes preprocessing of extracting the vibration information in the specific direction from the three-axis signal and utterance detection processing of detecting presence or absence of utterance on the basis of the vibration information.
[0138] The external terminal 2 in Fig. 16 is different from the external terminal 2 in Fig. 12 in that a communication unit 211 is provided and that the reception unit 111, the preprocessing unit 112, the detector 113, the sound information analysis unit 114, and the parameter optimization unit 115 are not provided.
[0139] The communication unit 211 receives the acceleration data and the sound data transmitted from the ear device 1 and transmits the acceleration data and the sound data to the cloud 201 via the network 202. Furthermore, the communication unit 211 receives the relearned learning model transmitted from the cloud 201 via the network 202 and transmits the relearned learning model to the ear device 1.
[0140] The cloud 201 includes a reception unit 221, a sound information analysis unit 222, and a relearning unit 223.
[0141] The reception unit 221 receives the acceleration data and the sound data transmitted from the external terminal 2 and decodes the acceleration data and the sound signal to generate the three-axis signal and the sound signal. The reception unit 221 supplies the three-axis signal to the relearning unit 223 and supplies the sound signal to the sound information analysis unit 222.
[0142] The sound information analysis unit 222 analyzes the sound signal supplied from the reception unit 221. Specifically, the sound information analysis unit 114 analyzes a volume of the utterance of the wearer on the basis of the sound signal and specifies time (period) at which the wearer is uttering. The sound information analysis unit 222 supplies the analysis result of the sound signal to the relearning unit 223 as sound information.
[0143] The relearning unit 223 relearns the learning model to be used by the detector 32 on the basis of the three-axis signal supplied from the reception unit 221. The three-axis signal to be used for relearning is, for example, a signal at the time when the wearer is uttering, which is specified by the sound information analysis unit 222.
[0144] The relearning unit 223 transmits the relearned learning model to the detector 32 of the ear device 1 via the external terminal 2 and updates the learning model to be used for the utterance detection processing by the detector 32.
[0145] Next, processing to be performed by the information processing system having the configuration of Fig. 16 will be described with reference to the flowchart of Fig. 17. The processing of Fig. 17 is started, for example, in a case where the wearer of the ear device 1 gives an instruction to start calibration of the utterance detection function.
[0146] In step S101, the UI control unit 61 of the external terminal 2 presents a guide to urge the wearer of the ear device 1 to utter. The wearer utters for a predetermined period according to the guide of the external terminal 2.
[0147] In step S102, the acceleration sensor 14 of the ear device 1 detects utterance vibration and acquires a three-axis signal.
[0148] In step S103, the transmission unit 101 of the ear device 1 transmits the acceleration data obtained by encoding the three-axis signal to the cloud 201 via the external terminal 2.
[0149] In step S104, the relearning unit 223 relearns the learning model to be used for the utterance detection processing by the detector 32 of the ear device 1 on the basis of the three-axis signal obtained by decoding the acceleration data transmitted from the ear device 1.
[0150] In step S105, the relearning unit 223 transmits the relearned learning model to the detector 32 of the ear device 1 via the external terminal 2 and updates the learning model to be used for the utterance detection processing by the detector 32.
[0151] Note that the cloud 201 may perform relearning of the learning model using the vibration information in the specific direction as an input instead of the learning model using the three-axis signal as an input. In this case, for example, a component corresponding to the preprocessing unit 31 is also provided in the cloud 201, and the relearning unit 223 performs relearning on the basis of the vibration information in the specific direction extracted by the component.
[0152] Although the example in which the wearer gives an instruction to start calibration of the utterance detection function has been described above, the information processing system can automatically calibrate the utterance detection function while the wearer is on the phone. In this case, the wearer does not have to perform special operation to give an instruction to start the calibration, and the external terminal 2 does not have to present a guide that urges the wearer to utter.
[0153] Processing of calibrating the utterance detection function during a call will be described with reference to a flowchart in Fig. 18.
[0154] In step S151, the external terminal 2 determines whether or not the wearer of the ear device 1 has started a call and waits until the wearer starts a call.
[0155] In a case where it is determined in step S151 that the wearer has started a call, in step S152, the acceleration sensor 14 of the ear device 1 detects utterance vibration during the call and acquires a three-axis signal.
[0156] In step S153, the parameter optimization unit 115 (Fig. 12) of the external terminal 2 optimizes the parameter of the preprocessing and applies the optimized parameter to the preprocessing in the preprocessing unit 31 of the ear device 1.
[0157] As described above, the information processing system of the present technology can calibrate the utterance detection function in the background of the call even if an instruction to start calibration is not given by the wearer of the ear device 1. Note that the wearer of the ear device 1 can set in advance whether or not to calibrate the utterance detection function in the background of the call.
[0158] <Regarding computer> The above-described series of processing can be performed by hardware or can be performed by software. In a case where the series of processing are executed by software, a program included in the software is installed from a program recording medium on a computer incorporated in dedicated hardware, a general-purpose personal computer, or the like. In some embodiments, the program comprises at least one non-transitory computer-readable storage medium having instructions encoded thereon that, when executed by circuitry, cause the circuitry to perform a method (e.g., any of the methods described herein).
[0159] Fig. 19 is a block diagram illustrating a configuration example of the hardware of the computer that executes the above-described series of processing by the program. The external terminal 2 and the cloud 201 include, for example, an information processing apparatus having a configuration similar to the configuration of the computer illustrated in Figs. 10A to 10C.
[0160] A CPU 501, a read only memory (ROM) 502, and a RAM 503 are mutually connected by a bus 504.
[0161] An input / output interface 505 is further connected to the bus 504. An input unit 506 including a keyboard, a mouse, and the like, and an output unit 507 including a display, a speaker, and the like, are connected to the input / output interface 505. Furthermore, a storage unit 508 including a hard disk, a nonvolatile memory, or the like, a communication unit 509 including a network interface, or the like, and a drive 510 that drives a removable medium 511 are connected to the input / output interface 505.
[0162] In the computer configured as described above, for example, the CPU 501 loads a program stored in the storage unit 508 into the RAM 503 via the input / output interface 505 and the bus 504 and executes the program to execute the above-described series of processing.
[0163] For example, the program to be executed by the CPU 501 is recorded in the removable medium 511, or provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital broadcasting, and then installed in the storage unit 508.
[0164] Note that the program to be executed by the computer may be a program in which processing is performed in time series in the order described in the present specification or may be a program in which processing is performed in parallel or at necessary timing such as when a call is made.
[0165] Note that, in the present specification, a system means an assembly of a plurality of components (devices, modules (parts), and the like), and it does not matter whether or not all the components are located in the same housing. Therefore, a plurality of devices housed in separate housings and connected to each other via a network and one device in which a plurality of modules is housed in one housing are both systems. Systems and techniques for performing calibration of an utterance detection function are described herein. Such techniques may be performed by circuitry (e.g., including at least one processor) which forms at least a part of the system. The circuitry may, in some embodiments, be part of a server. That is, the techniques described herein for performing a calibration of an utterance detection function may be performed by the server on the cloud. In some embodiments, the circuitry may be located on a device worn by a user (e.g., on a user’s ears). In such embodiments, the calibration of the utterance detection function may be performed by the device. In some embodiments, the circuitry may be part of an external apparatus external to the device worn by the user and the calibration of the utterance detection function may be performed by the external apparatus that is external to the device worn by the user. For example, the external apparatus may be a client device (e.g., a mobile device such as a smart phone, laptop, personal computer, etc.). In such embodiments where the circuitry is not located on the device worn by the user, the device worn by the user may communicate with the component (e.g., the server, external apparatus) which performs the calibration of the utterance detection function (e.g., by providing the sensor signal and / or the detection result to the circuitry). In some embodiments, the circuitry further obtains the sensor signal and / or the detection result.
[0166] Note that the effects described herein are only examples, and the effects of the present technology are not limited to these effects. Additional effects may also be obtained.
[0167] An embodiment of the present technology is not limited to the embodiments described above, and various modifications can be made without departing from the scope of the present technology.
[0168] For example, the present technology can have a cloud computing configuration in which one function is shared and processed in cooperation by a plurality of devices via a network.
[0169] Furthermore, each step described in the flowchart described above may be executed by one device, or can be executed by a plurality of devices in a shared manner.
[0170] Moreover, in a case where a plurality of pieces of processing is included in one step, the plurality of pieces of processing included in the one step can be executed by one device or executed by a plurality of devices in a shared manner.
[0171] <Combination examples of configurations> The present technology can also have the following configurations.
[0172] (1) A system comprising an information processing apparatus including: circuitry configured to perform calibration of an utterance detection function based on a sensor signal acquired by a sensor that detects a physical amount related to an utterance by a wearer of a device, and a detection result of presence or absence of the utterance by the utterance detection function that detects presence or absence of the utterance based on of the sensor signal. (2) The system according to (1), in which the device is worn on ears of the wearer. (3) The system according to (1) or (2), in which the sensor is an acceleration sensor that detects vibration generated by the utterance. (4) The system according to any one of (1) to (3), in which in the utterance detection function, the presence or absence of the utterance is detected using a learning model, the calibration performed by the circuitry includes adjustment of a parameter of preprocessing on the sensor signal, and the preprocessing includes processing of generating information to be input to the learning model based on the sensor signal. (5) The system according to (4), in which the sensor signal comprises a respective sensor signal for each of three axes; and the preprocessing generates information indicating vibration in a specific direction to be input to the learning model by weighting each of the sensor signals of the three axes and synthesizing the sensor signals. (6) The system according to (4) or (5), in which the circuitry is provided in an apparatus external to the device. (7) The system according to (6), in which the circuitry is configured to perform the calibration based on the sensor signal indicating a detection result of the physical amount related to the utterance by the wearer according to a guide presented by the apparatus external to the device and the detection result of the presence or absence of the utterance by the wearer according to the guide by the utterance detection function. (8) The system according to (4) or (5), in which the circuitry is provided in the device. (9) The system according to (8), in which the circuitry is configured to select the parameter to be used in the preprocessing from a plurality of candidates. (10) The system according to any one of (1) to (3), in which the calibration includes relearning of a learning model to be used for detecting the presence or absence of the utterance. (11) The system according to any one of (1) to (3), in which the calibration includes selection of a learning model to be used for detecting the presence or absence of the utterance from a plurality of candidates. (12) The system according to any one of (1) to (11), in which the circuitry is configured to perform the calibration based on the sensor signal indicating the detection result of the physical amount related to the utterance by the wearer during a call and the detection result of the presence or absence of the utterance by the wearer during the call by the utterance detection function. (13) The system according to any one of (1) to (12), in which the calibration unit determines a manner of wearing the device or a volume of the utterance when the presence or absence of the utterance is detected based on the sensor signal. (14) An information processing method to be performed by a system comprising an information processing apparatus, the information processing method including: performing calibration of an utterance detection function based on a sensor signal acquired by a sensor that detects a physical amount related to an utterance by a wearer of a device, and a detection result of presence or absence of the utterance by the utterance detection function that detects presence or absence of the utterance based on the sensor signal. (15) At least one non-transitory computer-readable storage medium having instructions encoded thereon, that when executed by circuitry of a system comprising an information processing apparatus, cause the circuitry to perform an information processing method comprising: performing calibration of an utterance detection function based on a sensor signal acquired by a sensor that detects a physical amount related to an utterance by a wearer of a device, and a detection result of presence or absence of the utterance by the utterance detection function that detects presence or absence of the utterance based on the sensor signal. (16) An system comprising: an information processing apparatus including: circuitry configured to determine a manner of wearing a device or a volume of an utterance by the wearer of the device where presence or absence of the utterance is detected based on a sensor signal acquired by a sensor that detects a physical amount related to the utterance of the wearer of the device. (17) The system according to (16), in which the sensor is an acceleration sensor that detects vibration generated by the utterance. (18) The system according to (16) or (17), in which the sensor is at least one microphone that detects a sound wave generated by the utterance. (19) The system according to (18), in which the device is worn on ears of the wearer, and the at least one microphone is mounted inside a housing of the device and outside the housing. (20) The system according to (19), wherein the circuitry is further configured to: analyze frequency characteristics of the sensor signal acquired by a first microphone of the at least one microphone inside the housing and frequency characteristics of the sensor signal acquired by a second microphone of the at least one microphone outside the housing, at least in part by comparing the frequency characteristics of the sensor signal acquired by the first microphone inside the housing with the frequency characteristics of the sensor signal acquired by the second microphone outside the housing and determine the manner of wearing the device based on the analyzing. (21) The system according to (16), in which the circuitry is configured to present a guide indicating the determined manner of wearing the device or the determined volume of the utterance by the wearer.
[0173] It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and alterations may occur depending on design requirements and other factors insofar as they are within the scope of the appended claims or the equivalents thereof.
[0174] 1 Ear device 2 External terminal 11 Driver 12A Outer microphone 12B Inner microphone 13 Substrate 14 Acceleration sensor 15 CPU / DSP 31 Preprocessing unit 32 Detector 33 Function execution unit 51 Sound information analysis unit 52 Recommendation determination unit 61 UI control unit 101 Transmission unit 111 Reception unit 112 Preprocessing unit 113 Detector 114 Sound information analysis unit 115, 151 Parameter optimization unit 201 Cloud 211 Communication unit 221 Reception unit 222 Sound information analysis unit 223 Relearning unit
Claims
1. A system comprising: an information processing apparatus comprising: circuitry configured to perform calibration of an utterance detection function based on: a sensor signal acquired by a sensor that detects a physical amount related to an utterance by a wearer of a device; and a detection result of presence or absence of the utterance by the utterance detection function that detects presence or absence of the utterance based on the sensor signal.
2. The system according to claim 1, wherein the device is worn on ears of the wearer.
3. The system according to claim 1, wherein the sensor is an acceleration sensor that detects vibration generated by the utterance.
4. The system according to claim 1, wherein in the utterance detection function, the presence or absence of the utterance is detected using a learning model, the calibration performed by the circuitry includes adjustment of a parameter of preprocessing on the sensor signal, and the preprocessing includes processing of generating information to be input to the learning model based on the sensor signal.
5. The system according to claim 4, wherein the sensor signal comprises a respective sensor signal for each of three axes; and the preprocessing generates information indicating vibration in a specific direction to be input to the learning model by weighting each of the sensor signals of the three axes and synthesizing the sensor signals.
6. The system according to claim 4, wherein the circuitry is provided in an apparatus external to the device.
7. The system according to claim 6, wherein the circuitry is configured to perform the calibration based on the sensor signal indicating a detection result of the physical amount related to the utterance by the wearer according to a guide presented by the apparatus external to the device and the detection result of the presence or absence of the utterance by the wearer according to the guide by the utterance detection function.
8. The system according to claim 4, wherein the circuitry is provided in the device.
9. The system according to claim 8, wherein the circuitry is configured to select the parameter to be used in the preprocessing from a plurality of candidates.
10. The system according to claim 1, wherein the calibration includes relearning of a learning model to be used for detecting the presence or absence of the utterance.
11. The system according to claim 1, wherein the calibration includes selection of a learning model to be used for detecting the presence or absence of the utterance from a plurality of candidates.
12. The system according to claim 1, wherein the circuitry is configured to perform the calibration based on the sensor signal indicating the detection result of the physical amount related to the utterance by the wearer during a call and the detection result of the presence or absence of the utterance by the wearer during the call by the utterance detection function.
13. The system according to claim 1, wherein the calibration unit determines a manner of wearing the device or a volume of the utterance when the presence or absence of the utterance is detected based on the sensor signal.
14. An information processing method to be performed by a system comprising an information processing apparatus, the information processing method comprising: performing calibration of an utterance detection function based on: a sensor signal acquired by a sensor that detects a physical amount related to an utterance by a wearer of a device; and a detection result of presence or absence of the utterance by the utterance detection function that detects presence or absence of the utterance based on the sensor signal.
15. At least one non-transitory computer-readable storage medium having instructions encoded thereon, that when executed by circuitry of a system comprising an information processing apparatus, cause the circuitry to perform an information processing method comprising: performing calibration of an utterance detection function based on: a sensor signal acquired by a sensor that detects a physical amount related to an utterance by a wearer of a device; and a detection result of presence or absence of the utterance by the utterance detection function that detects presence or absence of the utterance based on the sensor signal.
16. A system comprising: an information processing apparatus comprising: circuitry configured to determine a manner of wearing a device or a volume of an utterance by a wearer of the device where presence or absence of the utterance is detected based on a sensor signal acquired by a sensor that detects a physical amount related to the utterance of the wearer of the device.
17. The system according to claim 16, wherein the sensor is an acceleration sensor that detects vibration generated by the utterance.
18. The system according to claim 16, wherein the sensor is at least one microphone that detects a sound wave generated by the utterance.
19. The system according to claim 18, wherein the device is worn on ears of the wearer, and the at least one microphone is mounted inside a housing of the device and outside the housing.
20. The system according to claim 19, wherein the circuitry is further configured to: analyze frequency characteristics of the sensor signal acquired by a first microphone of the at least one microphone inside the housing and frequency characteristics of the sensor signal acquired by a second microphone of the at least one microphone outside the housing, at least in part by comparing the frequency characteristics of the sensor signal acquired by the first microphone inside the housing with the frequency characteristics of the sensor signal acquired by the second microphone outside the housing; and determine the manner of wearing the device based on the analyzing.
Citation Information
Patent Citations
Information processing device, information processing method, and program
EP3457399A1
Speech recognizing device and its control method
JP2005140860A
Acoustic control system and acoustic control method
JP2023091448A