Information processing equipment, information processing method and program

The information processing apparatus addresses the challenge of speech detection by calibrating the speech detection function based on sensor signals and determining optimal wearing methods or speech volumes, resulting in improved detection rates and user experience.

JP2025090076APending Publication Date: 2025-06-17SONY GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023205061
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-05
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

Existing technologies face challenges in accurately detecting speech by wearers of devices such as earphones, due to variations in signal-to-noise ratio caused by individual differences and improper wearing methods.

Method used

The proposed solution involves an information processing apparatus with a calibration unit that adjusts the speech detection function based on sensor signals from sensors detecting physical quantities related to speech, and a determination unit that determines the wearing method or speech volume for improved detection.

Benefits of technology

This approach enhances the detection rate of speech by wearers by optimizing the signal quality and adapting to individual variations, thereby improving the overall user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025090076000001_ABST
    Figure 2025090076000001_ABST
Patent Text Reader

Abstract

To improve a detection rate of utterance by a person with a device put on.SOLUTION: Information processing equipment comprises a calibration unit which calibrates an utterance detection function based upon a sensor signal acquired by a sensor which detects a physical quantity related to utterance of a person with a device put on and a detection result of whether the utterance is detected by an utterance detection function of detecting whether the utterance is given based upon the sensor signal. The present technology is applicable to an earphone which detects the utterance by the person with the device put on.SELECTED DRAWING: Figure 12
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present technology relates to an information processing apparatus, an information processing method, and a program, and particularly relates to an information processing apparatus, an information processing method, and a program that can improve the detection rate of speech by the wearer of a device.

Background Art

[0002] Many technologies have been developed to improve the UX (User Experience) of devices that users wear on their ears, such as earphones (inner ear headphones), TWS (True Wireless Stereo), and hearing aids. Patent Document 1 describes a technology for notifying a wearer of a good wearing state in order to improve the sound quality reproduced from earphones. In addition, in order to improve the UX of devices, the demand for environmental detection by devices is increasing.

[0003] For example, when the wearer of earphones speaks, the earphones detect the speech of the wearer and mute the music being played or transition to a mode of capturing external sounds. Even if the wearer does not control a smartphone or the like, the earphones seamlessly execute various functions according to the presence or absence of the wearer's speech, so that the wearer can, for example, talk to a person in front of them while wearing the earphones.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] The speech of the wearer is detected based on, for example, an audio signal acquired by a microphone (microphone) mounted on the earphone or a sensor signal acquired by a sensor. In order to accurately detect the speech of the wearer, it is important to increase the S / N (Signal-to-Noise ratio) of the signals acquired by the microphone or the sensor.

[0006] For example, when detecting the vibration generated by the wearer's speech with an acceleration sensor, since the vibration propagates through the head and reaches the earphone, the S / N of the acceleration signal varies greatly depending on individual differences and the way the earphone is worn. By determining the parameters used for speech detection based on the average value of the acceleration signals for each individual, the detection rate of speech by an average wearer can be improved, but the detection rate of speech by a wearer who deviates from the average decreases.

[0007] In the technology described in Patent Document 1, it is not possible to improve the S / N of the acceleration signal indicating the detection result of the vibration generated by the wearer's speech and propagating through the head to reach the earphone.

[0008] The present technology has been made in view of such a situation, and aims to improve the detection rate of speech by the wearer of the device.

Means for Solving the Problem

[0009] The information processing apparatus according to the first aspect of the present technology includes a calibration unit that calibrates the speech detection function based on a sensor signal acquired by a sensor that detects a physical quantity related to the speech of the wearer of the device and a detection result of the presence or absence of the speech by the speech detection function that detects the presence or absence of the speech based on the sensor signal.

[0010] The information processing method according to the first aspect of the present technology is such that the information processing apparatus calibrates the speech detection function based on a sensor signal acquired by a sensor that detects a physical quantity related to the speech of the wearer of the device and a detection result of the presence or absence of the speech by the speech detection function that detects the presence or absence of the speech based on the sensor signal.

[0011] The program of the first aspect of the present technology causes a computer to execute a process of calibrating the speech detection function based on a sensor signal acquired by a sensor that detects a physical quantity related to the speech of the wearer of the device and a detection result of the presence or absence of the speech by the speech detection function that detects the presence or absence of the speech based on the sensor signal.

[0012] The information processing apparatus of the second aspect of the present technology includes a determination unit that determines the wearing method of the device or the volume of the speech when detecting the presence or absence of the speech based on a sensor signal acquired by a sensor that detects a physical quantity related to the speech of the wearer of the device.

[0013] In the first aspect of the present technology, calibration of the speech detection function is performed based on a sensor signal acquired by a sensor that detects a physical quantity related to the speech of the wearer of the device and a detection result of the presence or absence of the speech by the speech detection function that detects the presence or absence of the speech based on the sensor signal.

[0014] In the second aspect of the present technology, the wearing method of the device or the volume of the speech when detecting the presence or absence of the speech is determined based on a sensor signal acquired by a sensor that detects a physical quantity related to the speech of the wearer of the device.

Brief Description of the Drawings

[0015]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Embodiments for Carrying Out the Invention

[0016] Hereinafter, embodiments for carrying out the present technology will be described. The description will be made in the following order. 1. Speech detection by ear device 2. First embodiment (example of recommending an appropriate wearing method) 3. Second embodiment (example of optimizing speech detection parameters)

[0017] <1. Speech detection by ear device> FIG. 1 is a diagram showing a configuration example of an information processing system to which the present technology is applied.

[0018] The information processing system in FIG. 1 is composed of TWS units 1L and 1R worn on each of the left and right ears and an external terminal 2.

[0019] The TWS units 1L and 1R are an example of wearable devices that a user wears on the ears. Hereinafter, such wearable devices are also referred to as ear devices 1. The ear device 1 includes, in addition to TWS, earphones, hearing aids, and the like. The TWS units 1L and 1R perform wireless communication with the external terminal 2 using, for example, BLE (Bluetooth Low Energy) in a paired state, and output the sound of music supplied from the external terminal 2.

[0020] The external terminal 2 is composed of a smartphone, a tablet terminal, a PC, or the like. The external terminal 2 plays music data and supplies the resulting music sound to the TWS units 1L and 1R.

[0021] The information processing system in FIG. 1 has a speech detection function for detecting the presence or absence of speech by the wearer of the ear device 1.

[0022] FIG. 2 is a diagram for explaining the speech detection process.

[0023] When the wearer speaks, the sensors mounted on the ear device 1 detect physical quantities related to the speech and acquire sensor signals. Specifically, as shown in #1 of FIG. 2, the acceleration sensor mounted on the ear device 1 detects vibrations (speech vibrations) generated by the speech and acquires an acceleration signal, or the microphone collects the speech of the speech and acquires a voice signal.

[0024] Next, for example, the arithmetic blocks in the DSP (Digital Signal Processor) or CPU (Central Processing Unit) mounted on the ear device 1 perform speech detection processing based on the acceleration signal and the voice signal acquired by the acceleration sensor and the microphone, as shown in #2 of FIG. 2. For example, by using a learning model, the presence or absence of speech by the wearer of the ear device 1 is detected.

[0025] Finally, the ear device 1 executes various functions based on the detection result of the presence or absence of speech. For example, the ear device 1 mutes the music being played or transitions to a mode of capturing external sounds.

[0026] FIG. 3 is a diagram showing a configuration example of the ear device 1.

[0027] As shown in FIG. 3, the ear device 1 is composed of a driver 11, an outer microphone 12A, an inner microphone 12B, and an acceleration sensor 14 and a CPU / DSP 15 provided on a substrate 13.

[0028] The driver 11 outputs the voice of the music supplied from the external terminal 2.

[0029] The outer microphone 12A is mounted on the outside of the housing of the ear device 1 (the side opposite to the ear side in the ear device 1), and collects the voice of the wearer's speech and external sounds to acquire a voice signal. The inner microphone 12B is mounted inside the housing of the ear device 1 (the ear side in the ear device 1), and collects the voice of the wearer's speech propagated inside the head to acquire a voice signal. Hereinafter, when it is not necessary to distinguish between the outer microphone 12A and the inner microphone 12B, they are simply referred to as the microphone 12.

[0030] The acceleration sensor 14 detects the speech vibration that has propagated inside the wearer's head and reached the ear device 1, and acquires an acceleration signal.

[0031] The CPU / DSP performs speech detection processing based on the audio signal acquired by the microphone 12, the acceleration signal acquired by the acceleration sensor 14, etc., and executes various functions based on the detection result of the presence or absence of speech.

[0032] Generally, in earphones, TWS, hearing aids, etc., speech detection is often performed based on the acceleration signal acquired by the acceleration sensor, and in headphones, etc., speech detection is often performed based on the audio signal acquired by the microphone.

[0033] FIG. 4 is a block diagram showing a functional configuration example of the ear device 1 when performing speech detection based on an acceleration signal.

[0034] As shown in FIG. 4, the ear device 1 is composed of an acceleration sensor 14, a preprocessing unit 31, a detector 32, and a function execution unit 33. The preprocessing unit 31, the detector 32, and the function execution unit 33 shown in FIG. 4 are realized, for example, by executing a predetermined program by the DSP / CPU 15 in FIG. 2. Note that FIG. 4 shows the configuration of the part related to speech detection among the configurations of the ear device 1.

[0035] When the wearer U1 of the ear device 1 speaks, the speech vibration is input to the acceleration sensor 14 through the wearer's head. The acceleration sensor 14 detects the acceleration in the three-axis directions of, for example, the x-axis, y-axis, and z-axis, and acquires a three-axis signal indicating the acceleration in the three-axis directions. The acceleration sensor 14 supplies the three-axis signal to the preprocessing unit 31.

[0036] The preprocessing unit 31 performs preprocessing on the three-axis signal. Specifically, the preprocessing unit 31 performs preprocessing on the three-axis signal by extracting vibration information indicating vibration in a specific direction from the three-axis signal. The vibration information is information for input to the learning model.

[0037] FIG. 5 is a diagram showing an example of vibration in a specific direction.

[0038] As shown in FIG. 5, it is assumed that the acceleration sensor 14 has acquired an acceleration signal x1 in the x-axis direction, an acceleration signal y1 in the y-axis direction, and an acceleration signal z1 in the z-axis direction. The preprocessing unit 31, as preprocessing, weights and synthesizes each acceleration signal to extract (generate), as vibration information, an acceleration signal p1 corresponding to the main component direction of the speech vibration that has propagated through the head and been input to the acceleration sensor 14.

[0039] Returning to FIG. 4, the preprocessing unit 31 supplies vibration information in a specific direction to the detector 32.

[0040] The detector 32 inputs the vibration information supplied from the preprocessing unit 31 to a learning model generated by machine learning to detect the presence or absence of the wearer's speech. Specifically, the detector 32 acquires the probability (speech probability) that the vibration input to the acceleration sensor 14 is speech vibration, and determines that the wearer has spoken when the speech probability is greater than a predetermined threshold. The detector 32 supplies the detection result of the presence or absence of speech to the function execution unit 33.

[0041] Thus, the speech detection function of the ear device 1 is realized by the preprocessing unit 31 and the detector 32.

[0042] The function execution unit 33 executes various functions according to the detection result of the presence or absence of speech by the detector 32.

[0043] Here, in order to improve the accuracy of speech detection by the detector 32, the quality of the acceleration signal (3-axis signal) acquired by the acceleration sensor 14 and the quality of the vibration information extracted by the preprocessing unit 31 are important.

[0044] For example, when the ear device 1 is not properly worn and the ear and the ear device 1 are not in close contact, the speech vibration is not sufficiently input to the acceleration sensor 14, and the ratio of the noise component included in the acceleration signal increases.

[0045] When sensing vibrations with acceleration in a certain direction, it has been found that high signal-to-noise vibration information can be obtained by appropriately weighting the three-axis signals respectively. Since the three-axis signals also contain noise components, if the three-axis signals acquired by the acceleration sensor 14 are directly combined to generate vibration information, the proportion of noise components contained in the vibration information will increase.

[0046] For example, if vibrations with acceleration only in the z-axis direction are input to the acceleration sensor 14, the acceleration signals in the x-axis direction and the y-axis direction do not contain vibration components but only noise components. If the three-axis signals are directly combined, the noise components contained in the vibration information will increase due to the acceleration signals in the x-axis direction and the y-axis direction, and the signal-to-noise ratio of the vibration information will decrease. In this case, if vibration information is generated using only the acceleration signal in the z-axis direction without using the acceleration signals in the x-axis direction and the y-axis direction, vibration information with a higher signal-to-noise ratio can be obtained than when generating vibration information using the three-axis signals.

[0047] When analyzing the speech vibrations actually detected by an acceleration sensor mounted on an earphone or the like, it has been found that there are individual differences and wearing errors in the main component direction of the speech vibrations. By determining the preprocessing parameters (weights) based on the average value of the main component direction of the speech vibrations for each individual, the detection rate of speech by an average wearer can be improved, but the detection rate of speech by a wearer deviating from the average will decrease.

[0048] In the first embodiment of the present technology, it is conceived by paying attention to the above points, and a technology is proposed that can improve the detection rate of speech by the wearer by recommending an appropriate wearing method to the wearer. Also, in the second embodiment, a technology is proposed that can improve the detection rate of speech by the wearer by calibrating the preprocessing parameters for each individual wearer. Hereinafter, the first embodiment and the second embodiment will be described in detail.

[0049] <2. First Embodiment (Example of Recommending an Appropriate Wearing Method)> FIG. 6 is a block diagram showing a functional configuration example of the information processing system according to the first embodiment.

[0050] The information processing system in FIG. 6 allows the wearer to manually calibrate the voice detection function. Specifically, the information processing system performs calibration by recommending an appropriate wearing method and the volume of speech to the wearer. For example, the information processing system recommends as an appropriate wearing method a method in which voice vibrations in a direction that matches the recommended direction (recommended direction) are input to the acceleration sensor 14. Note that FIG. 4 shows the configuration of the parts related to the recommendation of the wearing method among the configurations of the ear device 1 and the external terminal 2.

[0051] In FIG. 6, the same components as those in the configuration of FIG. 4 are denoted by the same reference numerals. Redundant descriptions will be omitted as appropriate. The ear device 1 in FIG. 6 is different from the ear device 1 in FIG. 4 in that an outer microphone 12A, an inner microphone 12B, a voice information analysis unit 51, and a recommendation determination unit 52 are provided. The voice information analysis unit 51 and the recommendation determination unit 52 shown in FIG. 6 are realized, for example, by executing a predetermined program by the DSP / CPU 15 in FIG. 2.

[0052] When the wearer U1 of the ear device 1 speaks, the sound waves of the speech are input to the outer microphone 12A by propagating through, for example, air, and are input to the inner microphone 12B by propagating through, for example, inside the head. The outer microphone 12A and the inner microphone 12B detect the input sound waves, acquire a voice signal, and supply the voice signal to the voice information analysis unit 51.

[0053] The voice information analysis unit 51 analyzes the voice signal supplied from the outer microphone 12A and the voice signal supplied from the inner microphone 12B. Specifically, the voice information analysis unit 51 analyzes the volume of the wearer's speech based on the voice signal acquired by the outer microphone 12A. In addition, the voice information analysis unit 51 analyzes the frequency characteristics of the voice signals acquired by the outer microphone 12A and the inner microphone 12B, respectively.

[0054] FIG. 7 is a diagram showing an example of the frequency characteristics of an audio signal. In FIG. 7, the horizontal axis represents frequency, and the vertical axis represents sound pressure.

[0055] In FIG. 7A, an example of the frequency characteristics of the audio signal acquired by the outer microphone 12A is shown. Since the sound wave radiated from the wearer's mouth is input through the propagation of air to the outer microphone 12A, the audio signal acquired by the outer microphone 12A becomes a broadband signal as shown in FIG. 7A.

[0056] In FIG. 7B, an example of the frequency characteristics of the audio signal acquired by the inner microphone 12B when the ear device 1 is properly worn is shown. When the ear device 1 is properly worn, since the inner microphone 12B is arranged in an acoustic condition sealed by the housing of the ear device 1 and the external auditory canal, in the inner microphone 12B, the energy of the sound wave input through the propagation of air becomes small, and the sound wave propagated through the head is also input. Therefore, the audio signal acquired by the inner microphone 12B becomes a narrowband signal as shown in FIG. 7B as compared with the audio signal acquired by the outer microphone 12A.

[0057] In FIG. 7C, an example of the frequency characteristics of the audio signal acquired by the inner microphone 12B when the ear device 1 is not properly worn is shown. When the ear device 1 is not properly worn, since the external auditory canal is not sealed by the housing of the ear device 1, it becomes easier for the sound wave to be input to the inner microphone 12B through the propagation of air. Therefore, the audio signal acquired by the inner microphone 12B becomes a signal having a band similar to the band of the audio signal acquired by the outer microphone 12A as shown in FIG. 7C.

[0058] Thus, depending on whether the ear device 1 is properly worn or not, the frequency characteristics of the audio signal acquired by the inner microphone 12B are different. The ear device 1 can determine whether the ear device 1 is properly worn by comparing the frequency characteristics of the audio signals acquired by the outer microphone 12A and the inner microphone 12, respectively.

[0059] Returning to FIG. 6, the voice information analysis unit 51 supplies the analysis result of the voice signal to the recommendation determination unit 52 as voice information.

[0060] The recommendation determination unit 52 is supplied with vibration information in a specific direction from the preprocessing unit 31 and the detection result of the absence of speech from the detector 32. The recommendation determination unit 52 recommends an appropriate wearing method and the volume of speech according to the detection result of the presence or absence of speech.

[0061] When speech is detected, the recommendation determination unit 52 presents, for example, the fact that the speech has been normally detected to the wearer via the external terminal 2.

[0062] When speech is not detected, the recommendation determination unit 52 functions as a determination unit that determines an appropriate wearing method of the ear device 1 and the volume of speech based on the vibration information in a specific direction and the analysis result of the voice signal in order to improve the speech detection rate. The recommendation determination unit 52 presents a guide indicating the determined wearing method of the ear device 1 and the volume of speech to the wearer via the external terminal 2. For example, the recommendation determination unit 52 presents a guide by means of a message, a graph, a voice, etc., prompting the wearer to change the wearing method of the ear device 1 or increase the volume of speech.

[0063] Note that a guide for recommending an appropriate wearing method and the volume of speech is created by the recommendation determination unit 52 or the external terminal 2 based on the vibration information in a specific direction, the detection result of the presence or absence of speech, and the analysis result of the voice signal, and is presented via the external terminal 2.

[0064] The external terminal 2 is constituted by a UI control unit 61. The UI control unit 61 controls a UI (User Interface) related to the calibration of the speech detection function. For example, when the start of calibration is instructed by a wearer of the ear device 1 through a predetermined operation on the UI, the UI control unit 61 controls the ear device 1 to start the calibration of the speech detection function.

[0065] Also, for example, during calibration, the UI control unit 61 presents a guide indicating an appropriate wearing method and the volume of speech to the wearer of the ear device 1 according to the detection result of the presence or absence of speech by the detector 32.

[0066] Next, with reference to the flowchart of FIG. 8, the processing performed by the information processing system having the configuration of FIG. 6 will be described. The processing in FIG. 8 is started, for example, when the wearer of the ear device 1 instructs the start of calibration of the speech detection function.

[0067] In step S1, the UI control unit 61 of the external terminal 2 presents a guide to prompt the wearer of the ear device 1 to speak. The wearer speaks at the volume desired to be detected as speech according to the guide.

[0068] In step S2, the acceleration sensor 14 of the ear device 1 detects speech vibration and acquires a three-axis signal. The preprocessing unit 31 of the ear device 1 extracts vibration information in a specific direction from the three-axis signal.

[0069] In step S3, the detector 32 of the ear device 1 performs a speech detection process based on the vibration information in the specific direction extracted from the three-axis signal.

[0070] In step S4, the recommendation determination unit 52 of the ear device 1 determines whether speech has been detected by the detector 32.

[0071] If it is determined in step S4 that speech has been detected, since it is considered that there is no problem with the wearing method of the ear device 1 or the volume of speech, in step S5, the recommendation determination unit 52 completes the calibration of the speech detection function. After the calibration of the speech detection function is completed, the processing ends.

[0072] On the other hand, if it is determined in step S4 that speech has not been detected, in step S6, the voice information analysis unit 51 of the ear device 1 analyzes the voice signals acquired by the outer microphone 12A and the inner microphone 12B, respectively.

[0073] In step S7, the recommendation determination unit 52 determines whether the volume of the speech is greater than a threshold value.

[0074] If it is determined in step S7 that the volume of the speech is greater than the threshold value, then in step S8, the recommendation determination unit 52 compares the frequency characteristics of the audio signal acquired by the outer microphone 12A with the frequency characteristics of the audio signal acquired by the inner microphone 12B.

[0075] In step S9, based on the comparison result of the frequency characteristics, the recommendation determination unit 52 identifies how the wearing method of the ear device 1 should be changed so that the speech can be detected, and recommends the appropriate wearing method to the wearer. When the volume of the speech is greater than the threshold value, since there is considered to be no problem with the volume of the speech, the recommendation determination unit 52 makes a recommendation regarding the wearing method.

[0076] On the other hand, if it is determined in step S7 that the volume of the speech is less than the threshold value, then in step S10, the recommendation determination unit 52 recommends an appropriate volume for the speech. When the volume of the speech is less than the threshold value, since there is considered to be a problem with the volume of the speech, the recommendation determination unit 52 makes a recommendation regarding the volume of the speech.

[0077] After an appropriate wearing method or the volume of the speech is recommended in step S9 or step S10, in step S11, the UI control unit 61 determines whether to perform the calibration again. For example, the UI control unit 61 displays a button asking whether to perform the calibration again. If the wearer of the ear device 1 selects to perform the calibration again, the UI control unit 61 determines to perform the calibration again.

[0078] If it is determined in step S11 that calibration is to be performed again, the process returns to step S1, and the subsequent processes are performed. By repeating the processes of steps S1 to S10, the wearer can improve the volume of speech and the wearing method, and increase the S / N of the three-axis signals acquired by the acceleration sensor.

[0079] On the other hand, if it is determined in step S11 that calibration is not to be performed again, the process ends.

[0080] As described above, in the information processing system according to the first embodiment of the present technology, based on the acceleration signal acquired by the acceleration sensor and the voice signal acquired by the microphone 12, the wearing method of the ear device 1 or the volume of speech when detecting the presence or absence of speech is determined. By such processing, the information processing system can discover the volume problem and the wearing problem that greatly affect speech detection, feedback to the wearer of the ear device 1, and recommend an appropriate volume of speech and a wearing method. By improving the volume of speech and the wearing method, the S / N of the three-axis signals acquired by the acceleration sensor 14 is increased, and the speech detection rate can be improved.

[0081] The information processing system separates and processes the volume problem and the wearing problem, such as determining whether the failure to detect speech is due to the volume of speech using the outer microphone 12A and determining whether it is due to the wearing method of the ear device 1 using the inner microphone 12B. Thereby, the information processing system can accurately grasp the volume problem and the wearing problem.

[0082] FIG. 9 is a diagram showing an example of a method of presenting a guide for recommending an appropriate wearing method and a volume of speech.

[0083] As shown in FIG. 9A, for example, a message for recommending an appropriate wearing method and a volume of speech is displayed in the upper part of the display of the external terminal 2, and a graph G1 is displayed below the message.

[0084] As shown in B of FIG. 9, for example, voices such as "Please raise the volume a little more" and "Please rotate the housing by 30 degrees" are output from the speaker of the external terminal 2. Note that instead of the speaker of the external terminal 2, voices for recommending an appropriate wearing method and the volume of speech may be output from the ear device 1.

[0085] FIG. 10 is a diagram showing an example of a graph for recommending an appropriate wearing method and the volume of speech.

[0086] As shown in A of FIG. 10, arrows indicating the direction of vibration actually input to the acceleration sensor 14 (the specific direction indicated by the vibration information generated by preprocessing) and the recommended direction are superimposed on, for example, an illustration of headphones and displayed in real time. In the example of A of FIG. 10, the direction of vibration actually input to the acceleration sensor 14 is indicated by a solid arrow, and the recommended direction is indicated by a dotted arrow.

[0087] The wearer of the ear device 1 rotates, for example, the ear device so that the solid arrow and the dotted arrow match while looking at the graph shown in A of FIG. 10.

[0088] As shown in B of FIG. 10, a bar graph indicating the speech probability (the probability that the vibration input to the acceleration sensor 14 is a vibration due to speech) is displayed in real time. Also, a threshold value serving as a criterion for detecting the presence or absence of speech is displayed.

[0089] The wearer of the ear device 1 improves the volume of speech and the wearing method of the ear device 1 so that the speech probability shown in the graph shown in B of FIG. 10 exceeds the threshold value.

[0090] As shown in C of FIG. 10, a line graph indicating the sound pressure level of speech at each time (time) is displayed in real time. Also, a band-shaped region indicating the sound pressure level recommended for detecting speech is displayed.

[0091] The wearer of the ear device 1 checks, for example, by looking at the line graph shown in C of FIG. 10 how much the sound pressure level (volume) of speech should be increased, and improves the volume of speech.

[0092] Note that the first embodiment of the present technology can also be applied to the ear device 1 not equipped with a microphone.

[0093] FIG. 11 is a diagram showing a configuration example of the ear device 1 not equipped with a microphone.

[0094] In FIG. 11, the same components as those in the configuration of FIG. 3 are denoted by the same reference numerals. Redundant descriptions will be omitted as appropriate. The ear device 1 in FIG. 11 is different from the ear device 1 in FIG. 3 in that the external microphone 12A and the internal microphone 12B are not provided.

[0095] When a microphone is not provided in the ear device 1, the processing flow of the information processing system is simplified, and only recommendations based on the three-axis signal acquired by the acceleration sensor 14 are made according to the detection result of the presence or absence of speech during calibration. For example, a graph as shown in A of FIG. 10 and a message regarding the wearing angle of the ear device 1 are displayed.

[0096] <3. Second Embodiment (Example of Optimizing Parameters for Speech Detection)> The information processing system according to the second embodiment automatically calibrates the speech detection function by adjusting and optimizing the preprocessing parameters for each individual wearer.

[0097] FIG. 12 is a diagram for explaining the flow of optimizing the preprocessing parameters.

[0098] First, the external terminal 2 prompts the wearer of the ear device 1 to speak. As shown on the right side of FIG. 12, when the wearer of the ear device 1 speaks, the vibration and sound wave generated by the speech are input into the ear device 1.

[0099] Next, the ear device 1 transmits the acceleration signal and the voice signal obtained by detecting the vibration and the sound wave to the external terminal 2 as shown on the left side of FIG. 12. Next, the external terminal 2 analyzes the acceleration signal and the voice signal transmitted from the ear device 1, and optimizes the parameters used for the preprocessing of speech detection based on the analysis results.

[0100] Finally, the external terminal 2 transmits the optimized parameters to the ear device 1 and updates the parameters used in the preprocessing by the preprocessing unit 31 of the ear device 1.

[0101] In this way, by the cooperation of the ear device 1 and the external terminal 2, the calibration of the speech detection function can be performed.

[0102] FIG. 13 is a block diagram showing a functional configuration example of the information processing system according to the second embodiment. Note that FIG. 13 shows the configuration of the parts related to the optimization of the parameters among the configurations of the ear device 1 and the external terminal 2.

[0103] In FIG. 13, the same configurations as those in FIG. 6 are denoted by the same reference numerals. Redundant descriptions will be omitted as appropriate. The ear device 1 in FIG. 13 is different from the ear device 1 in FIG. 6 in that a microphone 12 is provided instead of the outer microphone 12A and the inner microphone 12B. Further, the ear device in FIG. 13 is different from the ear device in FIG. 6 in that a transmission unit 101 is provided and the voice information analysis unit 51 and the recommendation determination unit 52 are not provided.

[0104] When the wearer U1 of the ear device 1 speaks, the sound wave generated by the speech is input to the microphone 12 by propagating through, for example, air. The microphone 12 detects the input sound wave, acquires a voice signal, and supplies the voice signal to the transmission unit 101.

[0105] The transmission unit 101 is supplied with the three-axis signals acquired by the acceleration sensor 14. The transmission unit 101 encodes the voice signal acquired by the microphone 12 and the three-axis signals acquired by the acceleration sensor 14 respectively to generate voice data and acceleration data. The transmission unit 101 transmits the voice data and the acceleration data to the external terminal 2.

[0106] The external terminal 2 in FIG. 13 is different from the external terminal 2 in FIG. 4 in that a receiving unit 111, a preprocessing unit 112, a detector 113, a voice information analysis unit 114, and a parameter optimization unit 115 are provided.

[0107] The receiving unit 111 receives the acceleration data and the voice data transmitted from the ear device 1, decodes the acceleration data and the voice signal to generate a three-axis signal and a voice signal. The receiving unit 111 supplies the three-axis signal to the preprocessing unit 112 and the parameter optimization unit 115, and supplies the voice signal to the voice information analysis unit 114.

[0108] The preprocessing unit 112 corresponds to the preprocessing unit 31 of the ear device 1, performs the same preprocessing as the preprocessing unit 31, and supplies vibration information in a specific direction to the detector 113.

[0109] The detector 113 corresponds to the detector 32 of the ear device 1, performs voice detection processing in the same manner as the detector 32, and supplies the detection result of the presence or absence of voice to the parameter optimization unit 115.

[0110] The voice information analysis unit 114 analyzes the voice signal supplied from the receiving unit 111. Specifically, the voice information analysis unit 114 analyzes the volume of the speaker's voice based on the voice signal and specifies the time (period) when the wearer is speaking. The voice information analysis unit 114 supplies the analysis result of the voice signal to the parameter optimization unit 115 as voice information.

[0111] Based on the three-axis signal supplied from the receiving unit 111 and the detection result of the presence or absence of speech by the detector 113, the parameter optimization unit 115 searches for parameters for preprocessing that maximize the speech probability. The three-axis signal used for parameter search is, for example, a signal at the time when the wearer is speaking, identified by the voice information analysis unit 114.

[0112] The parameter optimization unit 115 causes the preprocessing unit 112 to perform preprocessing using the searched parameters. Based on the detection result of the presence or absence of speech performed based on the vibration information in a specific direction extracted by the preprocessing, the parameter optimization unit 115 searches for parameters again that maximize the speech probability.

[0113] In this way, by repeating preprocessing, speech detection, and parameter search, the parameters are optimized. For example, when the speech probability becomes greater than a predetermined threshold, the parameter optimization unit 115 determines that the optimization of the parameters is complete, and causes the optimized parameters to be applied to the preprocessing in the preprocessing unit 31 of the ear device 1.

[0114] The parameter optimization unit 115 functions as a calibration unit that calibrates the speech detection function by optimizing the parameters for preprocessing.

[0115] Next, with reference to the flowchart of FIG. 14, the processing performed by the information processing system having the configuration of FIG. 13 will be described. The processing of FIG. 14 is started, for example, when the wearer of the ear device 1 instructs the start of calibration of the speech detection function.

[0116] In step S51, the UI control unit 61 of the external terminal 2 presents a guide to prompt the wearer of the ear device 1 to speak. The wearer speaks for a predetermined period according to the guide of the external terminal 2.

[0117] In step S52, the acceleration sensor 14 of the ear device 1 detects speech vibration and acquires a three-axis signal.

[0118] In step S53, the parameter optimization unit 115 of the external terminal 2 optimizes the parameters of the preprocessing and causes the optimized parameters to be applied to the preprocessing in the preprocessing unit 31 of the ear device 1.

[0119] As described above, in the external terminal 2 according to the second embodiment of the present technology, based on the acceleration signal acquired by the acceleration sensor 14 and the detection result of the presence or absence of speech by the speech detection function that detects the presence or absence of speech based on the acceleration signal, calibration of the speech detection function (optimization of the parameters of the preprocessing) is performed. By optimizing the parameters of the preprocessing, the S / N of the vibration information in a specific direction generated by the preprocessing unit 31 is increased, and it becomes possible to improve the speech detection rate.

[0120] By using the high-performance CPU mounted on the external terminal 2, the optimization of the preprocessing parameters can be performed at high speed.

[0121] When the processing memory of the ear device 1 is not abundant, if the ear device 1 itself tries to optimize the parameters, the wearer needs to keep talking while the parameters are being optimized. In the second embodiment of the present technology, since the parameters are optimized by the external terminal 2, the information processing system can optimize the parameters as long as the wearer talks for a predetermined period without having to keep talking.

[0122] Note that the optimization of the preprocessing parameters may be performed by the ear device 1.

[0123] FIG. 15 is a block diagram showing a functional configuration example of the information processing system when the ear device 1 performs the optimization of the preprocessing parameters.

[0124] In FIG. 15, the same components as those in the configuration of FIG. 6 are denoted by the same reference numerals. Redundant descriptions will be omitted as appropriate. The ear device 1 in FIG. 15 is different from the ear device 1 in FIG. 6 in that a microphone 12 is provided instead of the outer microphone 12A and the inner microphone 12B. Further, the ear device in FIG. 15 is different from the ear device in FIG. 6 in that a parameter optimization unit 151 is provided and a recommendation determination unit 52 is not provided.

[0125] The parameter optimization unit 151 is supplied with a three-axis signal acquired by the acceleration sensor 14, a detection result of the presence or absence of speech by the detector 32, and an analysis result of the speech signal by the speech information analysis unit 51.

[0126] The parameter optimization unit 115 searches for parameters of preprocessing such that the speech probability becomes maximum based on the three-axis signal, the detection result of the presence or absence of speech, and the speech information. Specifically, the parameter optimization unit 115 searches for optimal parameters by controlling the parameters used for the preprocessing of the preprocessing unit 31 while the wearer U1 of the ear device 1 is speaking.

[0127] For example, while the wearer U1 of the ear device 1 is speaking, the parameter optimization unit 151 causes the preprocessing unit 112 to perform preprocessing using each of a plurality of parameter candidates. The parameter optimization unit 151 compares the speech probabilities obtained based on the vibration information in a specific direction extracted by each preprocessing, and searches (selects) for the parameter that maximizes the speech probability. The parameter optimization unit 151 causes the optimized parameter to be applied to the preprocessing in the preprocessing unit 31.

[0128] During the search for the optimal parameters, the wearer U1 of the ear device 1 needs to keep speaking, but the difficulty of implementation is lower than when performed in cooperation with the external terminal 2. When the memory of the ear device 1 is abundant, the ear device 1 can also optimize the parameters of the preprocessing by temporarily storing the three-axis signal and the speech signal in the memory and using the stored three-axis signal and speech signal.

[0129] Note that although it is possible to implement the first and second embodiments of the present technology independently, by sequentially integrating an appropriate wearing method and volume recommendation for speech (first embodiment) with optimization of preprocessing parameters (second embodiment), it becomes possible to further improve the speech detection rate.

[0130] For example, first, an appropriate wearing method and volume recommendation for speech are performed. After the wearer of the ear device 1 improves the wearing method and volume of speech, optimization of the preprocessing parameters is performed.

[0131] In the above, an example where the ear device 1 or the external terminal 2 performs parameter optimization has been described, but it is also possible for the cloud to perform parameter optimization.

[0132] Also, in the above, an example where optimization of the parameters used in preprocessing is performed has been described, but optimization of the entire speech detection function realized by the preprocessing by the preprocessing unit 31 and the speech detection process by the detector 32 may be performed by the ear device 1, the external terminal 2, or the cloud. Optimization of the entire speech detection function includes, for example, adjustment of preprocessing parameters, re-learning of the learning model used in the speech detection process, and selection of the learning model used in the speech detection process from among a plurality of prepared learning model candidates.

[0133] FIG. 16 is a block diagram showing a functional configuration example of an information processing system when the cloud performs optimization of the entire speech detection function.

[0134] The information processing system in FIG. 16 is composed of the ear device 1, the external terminal 2, and the cloud 201. The external terminal 2 and the cloud 201 are connected via the network 202.

[0135] In FIG. 16, the same components as those in the configuration of FIG. 12 are denoted by the same reference numerals. Redundant descriptions will be omitted as appropriate. The ear device 1 in FIG. 16 is different from the ear device 1 in FIG. 12 in that the preprocessing unit 31 is not provided.

[0136] The detector 32 is supplied with the three-axis signal acquired by the acceleration sensor 14. The detector 32 inputs the three-axis signal into the learning model generated by machine learning to detect the presence or absence of the wearer's speech. Here, the learning model that takes the three-axis signal as input is a different model from the learning model that takes vibration information in a specific direction as input, and is a model that collectively executes the preprocessing of extracting the vibration information in a specific direction from the three-axis signal and the speech detection process of detecting the presence or absence of speech based on the vibration information.

[0137] The external terminal 2 in FIG. 16 is different from the external terminal 2 in FIG. 12 in that the communication unit 211 is provided and the receiving unit 111, the preprocessing unit 112, the detector 113, the voice information analysis unit 114, and the parameter optimization unit 115 are not provided.

[0138] The communication unit 211 receives the acceleration data and the voice data transmitted from the ear device 1, and transmits the acceleration data and the voice data to the cloud 201 via the network 202. Further, the communication unit 211 receives the re-trained learning model transmitted from the cloud 201 via the network 202, and transmits the re-trained learning model to the ear device 1.

[0139] The cloud 201 is composed of a receiving unit 221, a voice information analysis unit 222, and a re-learning unit 223.

[0140] The receiving unit 221 receives the acceleration data and the voice data transmitted from the external terminal 2, decrypts the acceleration data and the voice signal to generate a three-axis signal and a voice signal. The receiving unit 221 supplies the three-axis signal to the re-learning unit 223 and supplies the voice signal to the voice information analysis unit 222.

[0141] The voice information analysis unit 222 analyzes the voice signal supplied from the receiving unit 221. Specifically, the voice information analysis unit 114 analyzes the volume of the wearer's speech based on the voice signal and specifies the time (period) when the wearer is speaking. The voice information analysis unit 222 supplies the analysis result of the voice signal to the re-learning unit 223 as voice information.

[0142] Based on the three-axis signal supplied from the receiving unit 221, the relearning unit 223 performs relearning of the learning model used by the detector 32. The three-axis signal used for relearning is, for example, the signal at the time when the wearer is speaking, which is specified by the voice information analysis unit 222.

[0143] The relearning unit 223 transmits the relearned learning model to the detector 32 of the ear device 1 via the external terminal 2, and updates the learning model used for the speech detection process by the detector 32.

[0144] Next, with reference to the flowchart of FIG. 17, the process performed by the information processing system having the configuration of FIG. 16 will be described. The process of FIG. 17 is started, for example, when the wearer of the ear device 1 instructs the start of calibration of the speech detection function.

[0145] In step S101, the UI control unit 61 of the external terminal 2 presents a guide to prompt the wearer of the ear device 1 to speak. The wearer speaks for a predetermined period according to the guide of the external terminal 2.

[0146] In step S102, the acceleration sensor 14 of the ear device 1 detects speech vibration and acquires a three-axis signal.

[0147] In step S103, the transmitting unit 101 of the ear device 1 transmits the acceleration data obtained by encoding the three-axis signal to the cloud 201 via the external terminal 2.

[0148] In step S104, based on the three-axis signal obtained by decoding the acceleration data transmitted from the ear device 1, the relearning unit 223 performs relearning of the learning model used for the speech detection process by the detector 32 of the ear device 1.

[0149] In step S105, the relearning unit 223 transmits the relearned learning model to the detector 32 of the ear device 1 via the external terminal 2, and updates the learning model used for the speech detection process by the detector 32.

[0150] Note that, instead of a learning model that takes a three-axis signal as input, the relearning of a learning model that takes vibration information in a specific direction as input may be performed by the cloud 201. In this case, for example, a configuration corresponding to the preprocessing unit 31 is also provided in the cloud 201, and the relearning unit 223 performs relearning based on the vibration information in a specific direction extracted by the said configuration.

[0151] In the above, an example in which the wearer instructs the start of the calibration of the speech detection function has been described. However, the information processing system can also automatically perform the calibration of the speech detection function while the wearer is on a call. In this case, the wearer does not need to perform a special operation to instruct the start of the calibration, and the external terminal 2 does not need to present a guide to prompt the wearer to speak.

[0152] With reference to the flowchart of FIG. 18, the process of performing the calibration of the speech detection function during a call will be described.

[0153] In step S151, the external terminal 2 determines whether the wearer of the ear device 1 has started a call, and waits until the wearer starts a call.

[0154] If it is determined in step S151 that the wearer has started a call, then in step S152, the acceleration sensor 14 of the ear device 1 detects the speech vibration during the call and acquires a three-axis signal.

[0155] In step S153, the parameter optimization unit 115 (FIG. 12) of the external terminal 2 optimizes the parameters of the preprocessing, and causes the optimized parameters to be applied to the preprocessing in the preprocessing unit 31 of the ear device 1.

[0156] As described above, the information processing system of the present technology can calibrate the voice detection function in the background of a call without the wearer of the ear device 1 instructing the start of calibration. Note that the wearer of the ear device 1 can set in advance whether to perform calibration of the voice detection function in the background of a call.

[0157] <Regarding the computer> The above-described series of processes can be executed by hardware or by software. When the series of processes are executed by software, the program constituting the software is installed from a program recording medium into a computer incorporated in dedicated hardware or a general-purpose personal computer or the like.

[0158] FIG. 19 is a block diagram showing a configuration example of the hardware of a computer that executes the above-described series of processes by a program. The external terminal 2 and the cloud 201 are configured by an information processing apparatus having a configuration similar to the configuration of the computer shown in FIG. 10, for example.

[0159] The CPU 501, ROM (Read Only Memory) 502, and RAM 503 are interconnected by a bus 504.

[0160] An input / output interface 505 is further connected to the bus 504. An input unit 506 including a keyboard, a mouse, etc. and an output unit 507 including a display, a speaker, etc. are connected to the input / output interface 505. Further, a storage unit 508 including a hard disk, a non-volatile memory, etc., a communication unit 509 including a network interface, etc., and a drive 510 for driving a removable medium 511 are connected to the input / output interface 505.

[0161] In the computer configured as described above, for example, the CPU 501 loads and executes a program stored in the storage unit 508 into the RAM 503 via the input / output interface 505 and the bus 504, thereby performing the series of processes described above.

[0162] The program executed by the CPU 501 is recorded, for example, on a removable medium 511, or provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital broadcasting, and installed in the storage unit 508.

[0163] Note that the program executed by the computer may be a program in which processing is performed in time series in the order described in this specification, or a program in which processing is performed in parallel or at a necessary timing such as when a call is made.

[0164] Note that in this specification, the system means a collection of a plurality of components (devices, modules (parts), etc.), and it does not matter whether all the components are in the same housing. Therefore, a plurality of devices housed in separate housings and connected via a network, and one device in which a plurality of modules are housed in one housing are both systems.

[0165] Note that the effects described in this specification are merely examples and are not limited, and there may be other effects.

[0166] The embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible without departing from the gist of the present technology.

[0167] For example, the present technology can adopt a cloud computing configuration in which one function is shared and jointly processed by a plurality of devices via a network.

[0168] In addition, each step described in the above flowchart can be executed by one device or can be executed in cooperation by a plurality of devices.

[0169] Furthermore, when a plurality of processes are included in one step, the plurality of processes included in that one step can be executed by one device or can be executed in cooperation by a plurality of devices.

[0170] <Example of configuration combination> The present technology can also have the following configuration.

[0171] (1) A calibration unit that calibrates the speech detection function based on a sensor signal acquired by a sensor that detects a physical quantity related to the speech of the wearer of the device and a detection result of the presence or absence of the speech by the speech detection function based on the sensor signal. An information processing apparatus. (2) The device is worn on the ear of the wearer. The information processing apparatus according to (1) above. (3) The sensor is an acceleration sensor that detects vibrations generated by the speech. The information processing apparatus according to (1) or (2) above. (4) In the speech detection function, the presence or absence of the speech is detected by using a learning model. The calibration includes adjustment of parameters for preprocessing of the sensor signal. The preprocessing includes a process of generating information based on the sensor signal for input to the learning model. The information processing apparatus according to any one of (1) to (3) above. (5) The preprocessing generates information indicating vibrations in a specific direction for input to the learning model by weighting and synthesizing the three-axis sensor signals respectively. The information processing apparatus according to (4) above. (6) The calibration unit is provided in a device external to the device The information processing device according to (4) or (5) above (7) The calibration unit performs the calibration based on the sensor signal indicating the detection result of the physical quantity related to the speech made by the wearer according to the guide presented by the device external to the device, and the detection result of the presence or absence of the speech made by the wearer according to the guide by the speech detection function The information processing device according to (6) above (8) The calibration unit is provided in the device The information processing device according to (4) or (5) above (9) The calibration unit selects the parameter used for the preprocessing from among a plurality of candidates The information processing device according to (8) above (10) The calibration includes re-learning of the learning model used for detecting the presence or absence of the speech The information processing device according to any one of (1) to (3) above (11) The calibration includes selecting the learning model used for detecting the presence or absence of the speech from among a plurality of candidates The information processing device according to any one of (1) to (3) above (12) The calibration unit performs the calibration based on the sensor signal indicating the detection result of the physical quantity related to the speech of the wearer during a call, and the detection result of the presence or absence of the speech of the wearer during the call by the speech detection function The information processing device according to any one of (1) to (11) above (13) Based on the sensor signal, the calibration unit determines the wearing method of the device or the volume of the speech when detecting the presence or absence of the speech The information processing apparatus according to any one of (1) to (12) above. (14) The information processing apparatus performs calibration of the speech detection function based on a sensor signal acquired by a sensor that detects a physical quantity related to speech of the wearer of the device and a detection result of the presence or absence of the speech by the speech detection function that detects the presence or absence of the speech based on the sensor signal. Information processing method. (15) Causes a computer to perform calibration of the speech detection function based on a sensor signal acquired by a sensor that detects a physical quantity related to speech of the wearer of the device and a detection result of the presence or absence of the speech by the speech detection function that detects the presence or absence of the speech based on the sensor signal. Program for causing execution of processing. (16) Comprises a determination unit that determines the wearing method of the device or the volume of the speech when detecting the presence or absence of the speech based on a sensor signal acquired by a sensor that detects a physical quantity related to speech of the wearer of the device. Information processing apparatus. (17) The sensor is an acceleration sensor that detects vibration generated by the speech. The information processing apparatus according to (16) above. (18) The sensor is a microphone that detects sound waves generated by the speech. The information processing apparatus according to (16) or (17) above. (19) The device is worn on the ear of the wearer, the microphone is mounted on the inside and the outside of the housing of the device. The information processing apparatus according to (18) above. (20) Further comprises an analysis unit that analyzes the frequency characteristics of the sensor signal acquired by the microphone inside the housing and the frequency characteristics of the sensor signal acquired by the microphone outside the housing. The determination unit compares the frequency characteristics of the sensor signal acquired by the microphone inside the housing with the frequency characteristics of the sensor signal acquired by the microphone outside the housing, and determines the wearing method of the device. The information processing apparatus according to (19). (21) The determination unit presents a guide indicating the determined wearing method of the device or the volume of the utterance to the wearer. The information processing apparatus according to claim 16.

Description of reference numerals

[0172] 1 Ear device, 2 External terminal, 11 Driver, 12A Outer microphone, 12B Inner microphone, 13 Substrate, 14 Acceleration sensor, 15 CPU / DSP, 31 Preprocessing unit, 32 Detector, 33 Function execution unit, 51 Voice information analysis unit, 52 Recommendation determination unit, 61 UI control unit, 101 Transmission unit, 111 Reception unit, 112 Preprocessing unit, 113 Detector, 114 Voice information analysis unit, 115, 151 Parameter optimization unit, 201 Cloud, 211 Communication unit, 221 Reception unit, 222 Voice information analysis unit, 223 Relearning unit

Claims

1. A calibration unit that calibrates the speech detection function based on a sensor signal acquired by a sensor that detects a physical quantity related to speech of a wearer of the device, and a detection result of the presence or absence of the speech by the speech detection function based on the sensor signal. Information processing apparatus.

2. The device is worn on the ear of the wearer. The information processing apparatus according to claim 1.

3. The sensor is an acceleration sensor that detects vibrations generated by the speech. The information processing apparatus according to claim 1.

4. In the speech detection function, the presence or absence of speech is detected by using a learning model. The calibration includes adjustment of parameters for preprocessing of the sensor signal. The preprocessing includes a process of generating information based on the sensor signal for input to the learning model. The information processing apparatus according to claim 1.

5. The preprocessing generates information indicating vibrations in a specific direction for input to the learning model by weighting and synthesizing the three-axis sensor signals respectively. The information processing apparatus according to claim 4.

6. The calibration unit is provided in a device external to the device. The information processing apparatus according to claim 4.

7. The calibration unit calibrates based on the sensor signal indicating the detection result of the physical quantity related to the speech performed by the wearer according to a guide presented by a device external to the device, and the detection result of the presence or absence of the speech performed by the wearer according to the guide by the speech detection function. The information processing apparatus according to claim 6.

8. The calibration unit is provided in the device The information processing apparatus according to claim 4.

9. The calibration unit selects the parameter used in the preprocessing from among a plurality of candidates The information processing apparatus according to claim 8.

10. The calibration includes re-learning of a learning model used for detection of the presence or absence of speech The information processing apparatus according to claim 1.

11. The calibration includes selecting a learning model used for detection of the presence or absence of speech from among a plurality of candidates The information processing apparatus according to claim 1.

12. The calibration unit performs the calibration based on a sensor signal indicating a detection result of the physical quantity related to the speech of the wearer during a call and a detection result of the presence or absence of the speech of the wearer during the call by the speech detection function The information processing apparatus according to claim 1.

13. The calibration unit determines a wearing method of the device or a volume of the speech when detecting the presence or absence of the speech based on the sensor signal The information processing apparatus according to claim 1.

14. An information processing apparatus performs calibration of the speech detection function based on a sensor signal acquired by a sensor that detects a physical quantity related to the speech of a wearer of a device and a detection result of the presence or absence of the speech by the speech detection function that detects the presence or absence of the speech based on the sensor signal An information processing method.

15. To a computer Calibrate the voice detection function based on the sensor signal obtained by a sensor that detects a physical quantity related to the speech of the wearer of the device and the detection result of the presence or absence of the speech by the voice detection function that detects the presence or absence of the speech based on the sensor signal. A program for executing the process.

16. A determination unit that determines the wearing method of the device or the volume of the speech when detecting the presence or absence of the speech based on a sensor signal obtained by a sensor that detects a physical quantity related to the speech of the wearer of the device. An information processing apparatus.

17. The sensor is an acceleration sensor that detects vibrations generated by the speech. The information processing apparatus according to claim 16.

18. The sensor is a microphone that detects sound waves generated by the speech. The information processing apparatus according to claim 16.

19. The device is worn on the ear of the wearer. The microphone is mounted inside and outside the housing of the device. The information processing apparatus according to claim 18.

20. The apparatus further includes an analysis unit that analyzes the frequency characteristics of the sensor signal obtained by the microphone inside the housing and the frequency characteristics of the sensor signal obtained by the microphone outside the housing. The determination unit compares the frequency characteristics of the sensor signal obtained by the microphone inside the housing and the frequency characteristics of the sensor signal obtained by the microphone outside the housing, and determines the wearing method of the device. The information processing apparatus according to claim 19.

Citation Information

Patent Citations

  • Earphone, control arrangement of earphone, and control method of earphone

    JP2020150320A