Acoustic device, acoustic control method, and acoustic control program
By introducing a detection unit, a first determination unit and a second determination unit into the acoustic device, analyzing the user's voice and predetermined actions, the functional pause problem caused by misdetection in the prior art is solved, and better user experience and functional control are achieved.
Patent Information
- Application Number
- CN202380071275.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-11
- Filing Date
- 2023-08-28
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, when a user makes a sound in a non-interference manner, it is easy to detect it as a sound, resulting in the noise cancellation function being suspended and hinders the normal use of the user.
An acoustic device is designed, including a detection unit, a first determination unit and a second determination unit. During the content reproduction, the detection unit detects the user's voice, the first determination unit determines whether the control processing related to the content can be performed, and the second determination unit determines whether to perform the control processing by analyzing the user's predetermined action or the correlation between the sound and the content.
Achieve an excellent user experience, avoid unintentional function pauses, and ensure that users can easily talk to others when needed, without affecting the normal reproduction of the content.
Smart Images

Figure CN119998783A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an acoustic device, an acoustic control method, and an acoustic control program that perform predetermined processing when detecting an utterance of a user. Background Art
[0002] Some acoustic devices such as headphones with a noise cancellation function have a function of detecting a user's vocalization using a microphone, a vibration sensor, etc. and performing predetermined processing according to the detection result. For example, a certain acoustic device has a function of temporarily stopping the noise cancellation function when a user vocalizes in a manner that does not interfere with a conversation of a user using the device.
[0003] According to this function, when the user's voice is detected, not only the noise cancellation function but also the pause of the content being watched, the volume reduction or the change of the sound image positioning is possible. Thus, even when watching content, the user can immediately start a conversation with another person without being disturbed by the voice of the content.
[0004] Reference List
[0005] Patent Literature
[0006] Patent Document 1: Japanese Patent Application Publication No. 2011-97268 Summary of the invention
[0007] Technical issues
[0008] According to the related art, even a user wearing headphones or earphones can easily respond to a conversation when a voice is spoken by another person.
[0009] On the other hand, in the prior art, when a user sings a song along with music content reproduced from headphones or a user performs pronunciation practice according to language learning content, there is a possibility that it is detected as a vocalization. In this case, since the reproduction of the content is stopped at a timing that the user does not expect, there is a risk of hindering the user's use.
[0010] Thus, the present disclosure proposes an acoustic device, an acoustic control method, and an acoustic control program that can perform acoustic control with better user experience.
[0011] Solution to the problem
[0012] In order to solve the above problems, the acoustic device according to an embodiment of the present disclosure includes: a detection unit, which detects a user's voice during content reproduction; a first determination unit, which determines whether a content-related control process can be performed when the voice is detected by the detection unit; and a second determination unit, which determines whether to perform a content-related control process by analyzing whether the user has a predetermined action or the correlation between the voice and the content in the first determination process of the first determination unit. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 is a view showing an overview of an acoustic control process according to an embodiment.
[0014] Figure 2 is a block diagram showing an overview of an acoustic control process according to an embodiment.
[0015] Figure 3 : is a diagram showing an example of sensor positions of an acoustic device according to an embodiment.
[0016] Figure 4 is a diagram showing another example of the sensor position of the acoustic device according to the embodiment.
[0017] Figure 5 is a flowchart showing the procedure of an acoustic control process according to an embodiment.
[0018] Figure 6 : is a view showing a configuration example of the acoustic device according to the present embodiment.
[0019] Figure 7 is a block diagram showing an outline of an acoustic control process according to a first modification.
[0020] Figure 8 is a flowchart showing the procedure of the acoustic control process according to the first modification.
[0021] Fig. 9 is a block diagram showing an outline of an acoustic control process according to a second modification.
[0022] Fig.10 is a flowchart showing the procedure of the acoustic control process according to the second modification example.
[0023] Fig.11 is a block diagram showing an outline of an acoustic control process according to a third modification.
[0024] Fig.12 is a flowchart showing the procedure of the acoustic control process according to the third modification example.
[0025] Fig.13 is a diagram (1) for describing the configuration of a neural network according to a third modification.
[0026] Fig.14 is a diagram (2) for describing the configuration of a neural network according to a third modification.
[0027] Fig.15 is a block diagram showing an outline of an acoustic control process according to a fourth modification.
[0028] Fig.16 are views for describing a presentation process according to a fourth modification example.
[0029] Fig.17 is a flowchart showing the procedure of the acoustic control process according to the fourth modification.
[0030] Fig.18 is a block diagram showing an outline of an acoustic control process according to a fifth modification.
[0031] Fig.19 is a flowchart showing the procedure of the acoustic control process according to the fifth modification example.
[0032] Fig. 20 is a hardware configuration diagram showing an embodiment of a computer that realizes the functions of the acoustic device. DETAILED DESCRIPTION
[0033] Hereinafter, the embodiments of the present disclosure will be described in detail based on the drawings. In addition, in the following embodiments, the same reference numerals are attached to the same components, and repeated descriptions are omitted.
[0034] The present disclosure will be described in the order of the following items.
[0035] 1. Implementation Method
[0036] 1-1. Overview of Acoustic Control Processing According to Embodiment
[0037] 1-2. Procedure of Acoustic Control Processing According to Embodiment
[0038] 1-3. Configuration of Acoustic Device According to Embodiment
[0039] 1-4. Modifications of this embodiment
[0040] 1-4-1. Determination based on timing synchronization
[0041] 1-4-2. Determination of matching based on feature quantity
[0042] 1-4-3. Determination of input using voice and content
[0043] 1-4-4. User operation feedback
[0044] 1-4-5. Gesture feedback
[0045] 1-4-6. Determination based on content attributes and viewing patterns
[0046] 1-4-7. Mode of the function to be controlled
[0047] 2. Other Implementation Methods
[0048] 3. Effects of the Acoustic Device According to the Present Disclosure
[0049] 4. Hardware Configuration
[0050] (1. Implementation Method)
[0051] (1-1. Overview of Acoustic Control Processing According to the Present Embodiment)
[0052] First, refer to Figure 1 An outline of the acoustic control process according to the embodiment is described. Figure 1 is a view showing an outline of an acoustic control process according to an embodiment.
[0053] pass Figure 1 The acoustic device 100 shown in the figure performs the acoustic control processing according to the embodiment. The acoustic device 100 is, for example, a device having a sound output function, such as headphones or earphones. In addition, the acoustic device 100 includes a microphone, etc. to achieve a noise cancellation function. In this embodiment, it is assumed that the acoustic device 100 is a headphone and is worn by the user 10.
[0054] The acoustic device 100 outputs reproduced content, such as music. For example, the acoustic device 100 outputs a content reproduction device ( Figure 1 The content reproduction device (not shown) receives an acoustic signal of the content and outputs the received acoustic signal. The content reproduction device is, for example, a music reproduction player, a smart phone, etc. having a communication function.
[0055] In addition, the acoustic device 100 according to the present embodiment has a function of detecting whether the user 10 is uttering a sound during content reproduction and starting a predetermined process when the utterance is detected. For example, the utterance detection is achieved by utilizing a microphone with a beamforming function, a vibration sensor that detects the movement of the head and palate of the user 10, and the like. Specifically, the utterance detection process can be configured by utilizing a neural network. For example, the acoustic device 100 detects whether the user 10 is uttering a sound based on whether a numerical value (e.g., a value from 0 to 1) output by the neural network that represents the probability that the user 10 is currently uttering a sound satisfies a predetermined condition. The predetermined condition is, for example, that the state in which the output of the neural network is higher than a predetermined threshold value continues for a predetermined time.
[0056] Then, in Figure 1In the example shown in the upper part of , when the vocalization of the user 10 is detected, the acoustic device 100 pauses the content and converts the noise cancellation function into the external sound capturing function (step S10). The external sound capturing function is a function of outputting the external sound detected by the microphone and normally eliminated in the noise cancellation function from the output unit of the acoustic device 100. As described above, by stopping the content and capturing the external sound, the user 10 can have a conversation with another person without any trouble, without manually stopping the content or removing the acoustic device 100. It should be noted that in the following description, the function of temporarily switching to a mode suitable for conversation using the detection of vocalization as a trigger is referred to as a "conversation function." The conversation function is not limited to Figure 1 The modes shown in , and may be various modes such as temporarily reducing the reproduction volume of content or simply turning off the noise cancellation function instead of capturing external sounds, for example.
[0057] That is, the acoustic device 100 according to the present embodiment is a voice output device that detects the utterance of the user 10 who is a wearer using a microphone, a vibration sensor, etc., and performs some processing (in this example, a "conversation function") according to the result. According to the conversation function, the user 10 can enjoy an excellent user experience because it is possible to enjoy a short conversation with a nearby person while listening to music.
[0058] On the other hand, for example, in a case where the user 10 sings along with music content reproduced from the headphones or the user 10 performs pronunciation practice (referred to as shadowing, etc.) according to language learning content, the conversation function may operate against the intention of the user 10. Thus, conventionally, in the case of using a device having a conversation function, the user 10 needs to set the conversation function not to operate in the case of singing or performing pronunciation practice as described above.
[0059] Therefore, the acoustic device 100 according to the contents of the present disclosure appropriately implements both the conversation function and the sound emission detection through the processing described below. Specifically, in the case where it is detected that the user 10 is emitting sound during the content reproduction and the control processing of the conversation function is to be performed, the acoustic device 100 determines whether the user 10 is making a predetermined action (gesture). Then, in the case of performing the predetermined action, the acoustic device 100 performs control in a manner that does not operate the conversation function. As a result, in the case where the user 10 does not manually turn on or off the conversation function, the acoustic device 100 can turn on or off the conversation function at an appropriate timing according to the situation. As a result, therefore, the user 10 can enjoy the advantages of the original conversation function while emitting sound according to the content.
[0060] In the following, reference will be made to Figure 1 The following describes an overview of the processing according to the embodiment. Figure 1In the embodiment shown in the lower part of , it is assumed that the user 10 sings along with the content reproduced from the acoustic device 100 while wearing the acoustic device 100 .
[0061] Since the singing action of the user is detected as vocalization, the acoustic device 100 attempts to operate a control process of pausing the content. At this time, the acoustic device 100 determines whether the user 10 makes a predetermined gesture defined in advance (step S12).
[0062] The predefined predetermined gesture is a pattern of the user 10's action that is predetermined in a manner that does not operate the dialogue function. Specifically, an example of the gesture is an action in which the user 10 touches the touch sensor unit for operating the acoustic device 100. Figure 1 In the lower embodiment of , the user 10 sings a song while touching the touch sensor unit outside the right output unit (housing) of the acoustic device 100. Such an action corresponds to a predetermined gesture defined in advance.
[0063] In this case, the acoustic device 100 controls the activation of the conversation function (step S14). That is, in the case where the user 10 makes a sound while touching the touch sensor unit, the acoustic device 100 determines that the user 10 is in a state of intentionally controlling the conversation function, and does not activate the conversation function. On the other hand, in the case where the user 10 makes a sound without touching the touch sensor unit, the acoustic device 100 may activate the conversation function. As described above, the user 10 may control the activation of the conversation function based on a simple gesture or a natural gesture during singing a song. Therefore, the acoustic device 100 can provide acoustic control with high user satisfaction and a more excellent user experience.
[0064] Next, in Figure 2 The flow of acoustic control processing according to an embodiment is shown in the conceptual block diagram in . Figure 2 is a block diagram showing an overview of an acoustic control process according to an embodiment.
[0065] exist Figure 2 In the embodiment shown in FIG. 1 , the acoustic device 100 performs sound emission detection by using the sensor 20 included in the own device through the sound emission detection unit 22. The sensor 20 is, for example, a microphone, a vibration sensor, etc. It should be noted that the sound emission detection unit 22 is an embodiment of a processing unit that conceptually represents the information processing function of the acoustic device 100.
[0066] In addition, the content reproduction device 200 transmits the content 24 stored in the own device or acquired via the network to the acoustic device 100. At this time, when the acoustic device 100 detects the sound emission, the content reproduction device 200 controls the reproduction control unit 26 to pause the reproduction. Alternatively, even when the acoustic device 100 detects the sound emission, the reproduction control unit 26 performs control in such a manner that the reproduction is not paused when a specific gesture is detected at the same time.
[0067] Then, the output unit 28 of the acoustic device 100 outputs the voice of the content from the output unit 28 according to the result of the control of the reproduction control unit 26. The output unit 28 is a processing unit that actually outputs the voice in the acoustic device 100, and corresponds to, for example, a diaphragm or the like for voice output.
[0068] It should be noted that in Figure 2 In the embodiment of the present invention, there is a general configuration embodiment in which the upper part is on one side of the acoustic device 100 (e.g., earphone) and the lower part is on one side of the content reproduction device 200 (a host device connected to the earphone, e.g., a smart phone). That is, although the acoustic device 100 and the content reproduction device 200 are on the same side, the upper part is on the same side. Figure 2 , but the system configuration according to the embodiment is not limited to this embodiment. For example, the acoustic device 100 according to the present embodiment can perform reproduction control of content stored in the own device or content acquired via a network and perform sound emission detection through the own device. That is, in the case where the acoustic device 100 is a type of headphone that stores content in its own device, Figure 2 All components shown in FIG. 1 can be implemented in the acoustic device 100 as an earphone. Figure 2 , an example in which the sound emission detection unit 22 is installed inside the acoustic device 100 is shown, but the sound emission detection may be performed in a configuration in which the sensor signal is sent to the side of the host (content reproduction device 200) and the sound emission detection is performed on the side of the host. That is, the configuration described in the present embodiment is not limited, and any configuration may be used as long as the processing is achieved by the cooperation of the processing unit provided in any device. The same applies to the modified example described later.
[0069] Next, we will refer to Figure 3 and Figure 4 An example of the sensor position of the acoustic device 100 according to the embodiment is described. Figure 3 is a diagram showing an example of sensor positions of the acoustic device 100 according to the embodiment.
[0070] like Figure 3As shown, in the case where the acoustic device 100 is a headphone 100A, the touch sensor 30 for operation corresponds to the outside of the housing. Alternatively, the headphone 100A may include a touch sensor 32 located on an outer ring portion of the housing.
[0071] Next, Figure 4 2 shows an arrangement example of sensors in a case where the acoustic device 100 has another shape. Figure 4 is a diagram showing another example of sensor positions of the acoustic device 100 according to the embodiment.
[0072] As in Figure 4 As shown in FIG. 1 , in the case where the acoustic device 100 is a canal-type wireless headset 100B, the touch sensor 34 for operation corresponds to the external part attached to the auricle. Alternatively, the wireless headset 100B may include a touch sensor 36 on the outer ring part of the external part.
[0073] Notice, Figure 3 and Figure 4 The sensor positions shown in FIG. 1 are merely examples, and the acoustic device 100 may include sensors at different positions depending on the shape and purpose.
[0074] (1-2. Procedure of Acoustic Control Processing According to Embodiment)
[0075] Next, we will refer to Figure 5 A procedure of an acoustic control process according to an embodiment is described. Figure 5 is a flowchart showing the procedure of the acoustic control process according to the embodiment.
[0076] When reproducing content, the acoustic device 100 operates the sound emission detection function and waits for detection of sound emission (step S21). Then, the acoustic device 100 determines whether sound emission is detected (step S22). In the case where sound emission is not detected (step S22; No), the acoustic device 100 continues sound emission detection.
[0077] On the other hand, in the case where the utterance is detected (step S22; Yes), the acoustic device 100 determines whether the user is making a specific gesture (step S23).
[0078] In the case where the user does not make a specific gesture (step S23; No), the acoustic device 100 operates the conversation function and stops reproducing the content (step S24). On the other hand, in the case where the user makes a specific gesture (step S23; Yes), the acoustic device 100 does not operate the conversation function and ends the processing this time. At this time, in the case where the content reproduction continues, the acoustic device 100 may return the processing to step S21.
[0079] As referenced above Figures 1 to 5As described, the acoustic device 100 can predefine a specific gesture and prevent the processing operation corresponding to the sound emission detection (conversation function) when the user 10 performs the specific gesture. Therefore, when it is desired to make a sound according to the content, by performing a gesture with intention, the user 10 can make a sound according to the content without activating the processing corresponding to the sound emission detection. In addition, the user 10 can enjoy the benefits of the processing corresponding to the sound emission detection without making a gesture. Note that since it is a simple gesture of touching the touch sensor of the outer shell part and it is also a natural gesture for singing, the gesture of pressing the outer shell part with the hand is not too troublesome for the user 10. Thus, the user 10 can cause the processing according to the present embodiment to be performed by the natural action when making a sound.
[0080] By the way, a gesture of touching the touch sensor unit located on the earphone housing can be assigned to temporarily switch to the external sound capture mode. As described above, the external sound capture mode is a function that enables the user to check the surrounding sounds by collecting the surrounding sounds with a microphone and outputting the sounds from the earphone, and the external sound capture mode is used when the user desires to temporarily check the surrounding sounds in the noise cancellation earphone. Such earphones have the following specifications: when the user touches the touch sensor of the housing, it is temporarily switched from the noise cancellation mode to the external sound capture mode, and when the user releases his hand from the touch sensor, the mode returns to the noise cancellation mode again.
[0081] Here, the description of "not performing processing corresponding to sound emission detection when the touch sensor of the housing is touched (conversation function)" described in the embodiment may be combined with the description of the external sound capturing mode. That is, the acoustic device 100 may have a specification that the external sound capturing mode is temporarily switched to the external sound capturing mode and the conversation function is not operated when the user 10 touches the touch sensor of the housing. Although it seems that the two specifications conflict with each other, since the user 10 often wants to check his / her voice when performing sound emission according to the content, it may be convenient for the user 10 to switch to the external sound capturing mode.
[0082] Note that the acoustic device 100 may specify another gesture in order to avoid conflict with an existing gesture. For example, the acoustic device 100 may determine the number of fingers touching the touch sensor or whether the user 10 touches a different position from the touch sensor used to determine the transition to the external sound capturing mode. Alternatively, the acoustic device 100 may specify another gesture in addition to the finger touch sensor. Figure 3 and Figure 4 In addition to the positions shown in , a touch sensor dedicated to gestures according to the implementation may be added.
[0083] Note that as another example of a gesture for limiting the conversation function, the acoustic device 100 may detect “the user 10 brings the hand close to the housing.” For example, the acoustic device 100 may detect the gesture of the user 10 by using a sensor for detecting an object without physical contact, such as an infrared sensor.
[0084] In addition, the acoustic device 100 may detect a predetermined operation as an action of the user 10. That is, the acoustic device 100 may determine whether to execute the conversation function by analyzing whether the user 10 presses an operation unit (such as an operation button) to control the acoustic device 100 or the content. For example, the acoustic device 100 may use the user 10 pressing an operation button on the side surface of the acoustic device 100 or a touch panel of a smart phone, smart watch, etc. linked to the acoustic device 100 as a determination criterion for controlling the conversation function.
[0085] In addition, the acoustic device 100 can control the conversation function not only when the user 10 performs a gesture, but also within a predetermined time period after the gesture is performed or until the content being reproduced ends after the gesture is performed. Therefore, the user 10 does not need to continuously perform a specific gesture when performing a sound according to the content. In addition, on the contrary, the acoustic device 100 can perform a process corresponding to the sound detection only when the user 10 performs a specific gesture.
[0086] (1-3. Configuration of Acoustic Device According to Present Embodiment)
[0087] Next, we will refer to Figure 6 The configuration of the acoustic device 100 according to the embodiment is described. Figure 6 is a view showing a configuration example of the acoustic device 100 according to the present embodiment.
[0088] like Figure 6 As shown, the acoustic device 100 includes a communication unit 110, a storage unit 120, and a control unit 130. It is to be noted that the acoustic device 100 may include an input unit (e.g., a touch sensor unit) for receiving various operations from a user or the like who operates the acoustic device 100, and a display unit (e.g., a liquid crystal display) for displaying various information.
[0089] The communication unit 110 is implemented by, for example, a network interface card (NIC), a network interface controller, etc. The communication unit 110 is connected to the network N in a wired or wireless manner, and transmits and receives information to and from the content reproduction device 200 and the like via the network N. The network N is implemented by, for example, a wireless communication standard or system such as Bluetooth (registered trademark), the Internet, Wi-Fi (registered trademark), ultra-wideband (UWB), or low-power wide area (LPWA).
[0090] The storage unit 120 is implemented by, for example, a semiconductor memory element such as a random access memory (RAM) or a flash memory, or a storage device such as a hard disk or an optical disk. For example, the storage unit 120 stores data such as content reproduced by the acoustic device 100.
[0091] For example, when a program (such as the acoustic control program according to the present disclosure) stored in the acoustic device 100 is executed by a central processing unit (CPU), a micro processing unit (MPU), or the like having a random access memory (RAM) or the like as a work area, the control unit 130 is implemented. In addition, the control unit 130 is a controller and can be implemented by an integrated circuit such as an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA).
[0092] like Figure 6 As shown in FIG. 1 , the control unit 130 includes an acquisition unit 131, a reproduction control unit 132, a detection unit 133, a first determination unit 134, and a second determination unit, and implements or performs the functions and actions of the information processing described below. It should be noted that the internal configuration of the control unit 130 is not limited to Figure 6 The configuration shown in , and may be another configuration as long as it is a configuration that performs information processing described later. Figure 6 The acoustic device 100 includes the reproduction control unit 132, but the content reproduction apparatus 200 may include the reproduction control unit 132 as described above, etc. In this case, the information processing according to the embodiment is implemented by the cooperation of the acoustic device 100 and the content reproduction apparatus 200. That is, the control unit 130 collectively refers to a processing unit conceptually representing information processing, such as, Figure 2 , and the sound detection unit 22 and the reproduction control unit 26 shown in FIG. Similarly, the control unit 130 also corresponds to Figure 7 , Fig. 9 , Fig.11 , Fig.15 and Fig.18 Each processing unit shown in (described later).
[0093] The acquisition unit 131 acquires various information. For example, the acquisition unit 131 acquires data of content reproduced by the acoustic device 100 from the content reproduction device 200. In addition, the acquisition unit 131 acquires various types of setting information such as the definition of gestures for controlling the conversation function. For example, the acquisition unit 131 acquires various types of setting information by receiving various types of setting information provided from an administrator of the acoustic device 100 or the like via the network N.
[0094] The reproduction control unit 132 controls the reproduction of the content. For example, the reproduction control unit 132 reads the content data stored in the storage unit 120 and performs control in such a manner that a voice signal corresponding to the read data is output from the output unit of the acoustic device 100.
[0095] In addition, the reproduction control unit 132 controls the reproduction of the content in response to a request from the user 10. For example, the reproduction control unit 132 responds to a request from the user 10 to stop the reproduction, increase or decrease the volume, or view other content, and controls the reproduction mode. In addition, in the case where the detection unit 133 detects the utterance of the user 10 and the conversation function is activated, the reproduction control unit 132 performs control such as pausing the content or decreasing the volume.
[0096] The detection unit 133 detects events and actions (such as operations on the acoustic device 100) observed around the acoustic device 100 by using various sensors. For example, the detection unit 133 detects that the user 10 utters a sound during content reproduction by using a microphone or a vibration sensor.
[0097] In addition, by using a touch sensor or an infrared sensor, the detection unit 133 can detect that the user 10 makes a predetermined action for controlling the conversation function. For example, the detection unit 133 detects that the user 10 is making a predefined gesture, such as touching the outside of the acoustic device 100 during content reproduction.
[0098] In the case where the utterance is detected by the detection unit 133, the first determination unit 134 determines whether the control process related to the content can be performed. For example, the control process related to the content is various types of control useful for the user 10 to have a conversation with another person, such as pausing the content or activating the external sound capture function in the case of detecting the utterance of the user 10. That is, in the case of detecting the utterance of the user 10, the first determination unit 134 determines whether to activate the conversation function.
[0099] In the first determination process of the first determination unit 134, the second determination unit 135 determines whether to perform the control process related to the content by analyzing the presence or absence of the predetermined action of the user 10. That is, according to the situation of the user 10, the second determination unit 135 re-determines whether to actually activate the dialogue function to be activated by the first determination unit 134.
[0100] Specifically, the second determination unit 135 determines whether to perform a control process related to the content by analyzing whether the user 10 touches the touch sensor unit to control the acoustic device 100 as a predetermined action of the user 10. For example, Figure 1As shown, in the case where the user 10 sings while touching the outside of the housing of the earphone with a hand, the second determination unit 135 determines that the action is a gesture for controlling the conversation function, and performs control in a manner of not activating the conversation function.
[0101] In addition, the second determination unit 135 may determine whether to perform a content-related control process by analyzing whether the user 10 brings the hand close to a predetermined part of the acoustic device 100 as a predetermined action of the user 10. That is, in the case where the user 10 performs a predetermined action (such as holding the hand within a predetermined distance from the sensor unit of the acoustic device 100) even when the user 10 does not actually touch the acoustic device 100 with the hand, the second determination unit 135 may perform control in a manner that does not activate the dialogue function based on such an action.
[0102] In addition, the second determination unit 135 may determine whether to perform a control process related to the content by analyzing whether the user 10 presses the operation unit to control the acoustic device 100 or the content as a predetermined action of the user 10. For example, in the case where the user 10 presses a specific operation button included in the acoustic device 100 or an action of operating a content reproduction application, the second determination unit 135 may perform control in a manner not to activate the dialogue function based on such an action.
[0103] Note that such an action is not essential for the acoustic device 100, and may be an action of the content reproduction device 200 implemented by a smartphone or a smart watch. In addition, such an action does not necessarily involve the pressing of a button, and may be an action such as whether the user 10 is watching the script (such as lyrics) of the content displayed on the smartphone. In this case, the sensor corresponds to a sight line detection function and a face detection function by a camera mounted on the smartphone.
[0104] (1-4. Modifications of the Embodiments)
[0105] The information processing according to the above-described embodiment may be accompanied by various modifications. Modifications of the embodiment will be described below.
[0106] In the above-described embodiment, the gesture of the user 10 is used as an example of a trigger for the acoustic device 100 to control the conversation function. However, the acoustic device 100 may control the conversation function based on other determination criteria. Hereinafter, as a modified embodiment of the embodiment, an embodiment in which the acoustic device 100 controls the conversation function based on various determination criteria will be listed.
[0107] (1-4-1. Determination based on timing synchronization)
[0108] The acoustic device 100 may determine whether to control the conversation function based on synchronization between the timing of the utterance of the user 10 and the timing of the output of the voice signal included in the content. Figure 7 and subsequent figures describe this embodiment.
[0109] Figure 7 is a block diagram showing an outline of the acoustic control process according to the first modification. Figure 7 , an embodiment is shown in which the acoustic device 100 or the content reproducing apparatus 200 temporarily controls the conversation function by determining the synchronization between the utterance of the user 10 and the content. That is, in the case where the user 10 sings or performs shadowing along the reproduced content, it is assumed that the timing of the start of the voice contained in the content and the utterance of the user 10 substantially match. Thus, such synchronization is used as the determination criterion in the first modification.
[0110] It should be noted that it is technically difficult to distinguish between sounds according to the content and other sounds using only sensor signals from microphones, vibration sensors, etc. Thus, the acoustic device 100 is able to make the distinction by using information of the content being reproduced.
[0111] like Figure 7 As shown, the acoustic device 100 continues the sound utterance detection process by the sound utterance detection unit 40 during the content reproduction. In addition, the voice detection unit 42 of the content reproduction device 200 analyzes the content provided from the content reproduction device 200, and determines in real time whether the voice is included in the content reproduced at this time. Note that the voice detection unit 42 may include a neural network similar to the sound utterance detection unit 40.
[0112] Then, the synchronization determination unit 44 compares the timing of the output of the utterance detection unit 40 with the timing of the output of the voice detection unit 42. For example, in the case where it is determined that the two perform outputs at a timing within a predetermined threshold, the synchronization determination unit 44 determines that the user 10 utters a voice according to the voice included in the content. That is, the synchronization determination unit 44 determines whether the timing at which the output of the utterance detection unit 40 changes from the state of "absent" to "present" is substantially the same as the timing at which the output of the voice detection unit 42 changes from the state of "absent" to "present". Specifically, the synchronization determination unit 44 determines whether the timing of the two coincides within a prescribed time width.
[0113] Then, the reproduction control unit 46 controls the reproduction of the content based on the output of the utterance detection unit 40 and the output of the synchronization determination unit 44. For example, in the case where the output of the utterance detection unit 40 is "present" and the output of the synchronization determination unit 44 is "absent", the reproduction control unit 46 determines that the user 10 uttered a sound in accordance with the content instead of uttering a sound as a conversation, and stops the reproduced content. On the other hand, in the case where the output of the utterance detection unit 40 is "present" and the output of the synchronization determination unit 44 is "present", the reproduction control unit 46 determines that the user 10 uttered a sound in accordance with the content and performs control in a manner not to stop the reproduction of the content.
[0114] Then, the output unit 48 of the acoustic device 100 outputs the voice of the content from the output unit 48 according to the result of the control of the reproduction control unit 46 .
[0115] It should be noted that in the content reproduction apparatus 200 or the acoustic device 100, music data or the like of the content reproduced in the acoustic device 100 can be read in advance. Thus, the voice detection unit 42 can perform voice detection processing before the portion before the current reproduction position in the content. As a result, the acoustic device 100 can prevent processing delays and distribute processing loads.
[0116] Furthermore, the acoustic device 100 may acquire information of a timing at which a voice is included in content from meta information of music or from an external database accumulated in advance or the like.
[0117] In addition, the acoustic device 100 may change the parameters of the determination criteria in the synchronization determination unit 44 according to the type of content being reproduced. For example, in the case where the content being reproduced is music content, the acoustic device 100 may set the threshold value of the timing of the start of the voice included in the content and the start of the utterance (singing voice) of the user 10 to be extremely short. This is because it is assumed that for a song, the user 10 utters the voice almost simultaneously with the content. On the other hand, in the case where the reproduced content is content for language learning, the acoustic device 100 may assume that the utterance of the user is slightly delayed relative to the start of the voice included in the content, and the acoustic device 100 may set the threshold value for synchronization determination of the timing to be long compared to the case of music content. This is because, as a mode of language practice such as shadowing, it is assumed that the user 10 utters a sound slightly later than the reproduction of the content. Note that, Figure 7 Each of the processing units shown in FIG. 4 (such as the voice detection unit 42 and the synchronization determination unit 44) may be included in the Figure 2 The acoustic device 100 may be included in the content reproducing apparatus 200.
[0118] Next, we will refer to Figure 8 The procedure of the acoustic control process according to the first modification example is described. Figure 8is a flowchart showing the procedure of the acoustic control process according to the first modification.
[0119] When reproducing the content, the acoustic device 100 operates the sound emission detection function through the sensor and waits for the detection of the sound emission (step S31). That is, the acoustic device 100 performs the sound emission detection through the sensor (step S32). In addition, the acoustic device 100 continues to detect whether the content being reproduced includes the process of voice (step S33).
[0120] Then, the acoustic device 100 performs synchronization determination as to whether the timings of the two detected in step S32 and step S33 are consistent with each other (step S34). During this period, the acoustic device 100 determines whether the utterance of the user 10 is detected (step S35). In the case where the utterance is not detected (step S35; No), the acoustic device 100 continues the utterance detection.
[0121] On the other hand, in the case where the sound is detected (step S35; yes), the acoustic device 100 further determines whether the sound and the voice are synchronized with each other (step S36). In the case where the sound and the voice are not synchronized with each other (step S36; no), the acoustic device 100 operates the dialogue function and stops reproducing the content (step S37). On the other hand, in the case where the sound and the voice are synchronized with each other (step S36; yes), the acoustic device 100 does not operate the dialogue function and ends the current processing. At this time, in the case where the content reproduction continues, the acoustic device 100 may return the processing to step S31.
[0122] As described above, in the first determination process by the first determination unit 134, the second determination unit 135 according to the first modification example analyzes the correlation between the utterance and the content, and determines whether to perform the control process related to the content. For example, as the correlation between the utterance and the content, the second determination unit 135 determines whether to perform the control process related to the content based on the matching degree of the timing of the voice included in the content and the utterance.
[0123] As a result, the user 10 can control activation of the conversation function simply by speaking a word at substantially the same time as the voice included in the content without making a specific gesture, and can control the conversation function in a natural mode.
[0124] (1-4-2. Determination of Matching Based on Feature Quantity)
[0125] Next, as a second modification example, a process in which the acoustic device 100 determines whether to control the conversation function based on a match between a feature amount of a voice included in content and an utterance will be described.
[0126] That is, in the case where the user 10 utters a sound according to the reproduced content, the acoustic device 100 focuses on the similarity between the feature amount of the content and the feature amount of the user's utterance, and changes the control related to the content based on the feature amount. Fig. 9 and shown in the subsequent figures.
[0127] Fig. 9 is a block diagram showing an overview of the acoustic control process according to the second modification. Fig. 9 As shown, the acoustic device 100 continues the utterance detection process during content reproduction by the utterance detection unit 50. In addition, the first feature amount calculation unit 52 that calculates the feature amount related to the utterance analyzes the utterance of the user 10 and calculates its feature amount. In addition, the second feature amount calculation unit 54 analyzes the data of the content and calculates its feature amount.
[0128] That is, the acoustic device 100 calculates the feature quantity (i.e., sensor signal) and content (i.e., content signal) of each sound. The feature quantity is, for example, a musical element in the sound, such as the pitch, timbre, chord, rhythm, etc. of the sound. In addition, the feature quantity may include the content of the sound and the semantic content (text information) such as the lyrics of the content. These feature quantities can be calculated based on known methods of speech or music analysis, analysis and recognition of text of speech, etc. Note that as a feature quantity for information processing, only a single feature quantity may be used, or multiple features may be combined.
[0129] The feature amount matching determination unit 56 determines the degree of matching between the output of the first feature amount calculation unit 52 and the output of the second feature amount calculation unit 54. That is, the feature amount matching determination unit 56 determines how much the feature amount of the sensor signal acquired by the sensor of the acoustic device 100 is similar to the feature amount acquired from the signal of the content itself. Then, in the case where the degree of matching is equal to or greater than a predetermined threshold, the feature amount matching determination unit 56 determines that the feature amounts match.
[0130] As an embodiment, in the case where the feature quantity is "pitch", the feature quantity matching determination unit 56 calculates the pitch of the sensor signal and the content signal at a specific time interval, and determines that the two match when the difference is equal to or less than a predetermined threshold. As described above, the acoustic device 100 determines that the probability that the user 10 sings along with the content is high based on the match between the pitch of the user 10's vocalization and the pitch of the voice of the content.
[0131] Note that the user 10 may sing at different octaves, or may sing at intervals that are in harmony with the voice of the content. Thus, the acoustic device 100 may adjust the determination criteria such as the score so that the feature quantity match determination unit 56 may determine two "matches" even if the pitch of the feature quantity match determination unit 56 is different for the octave or has a chord relationship (e.g., major third or minor third).
[0132] Furthermore, in the case where the feature quantity is “pitch” or “chord length”, the feature quantity match determination unit 56 determines whether the pitch or chord length of the sensor signal (the utterance of the user 10) matches the pitch or chord length of the content. As an embodiment, in the case where the content is music of C major, the feature quantity match determination unit 56 may determine that the feature quantity matches when the utterance of the user 10 includes the scale of C major, and may determine that the feature quantity does not match when it includes an interval not included in the scale of C major.
[0133] In addition, for example, in the case where the feature quantity is "rhythm", the feature quantity matching determination unit 56 determines whether the beat of the sensor signal (such as beats per minute (BPM)) matches the rhythm of the content. Specifically, the feature quantity matching determination unit 56 analyzes the temporal changes in the sound pressure level of the sensor signal and the sound pressure level of the content signal, detects the timing of the intensity, the way of beating, etc., and determines whether these match.
[0134] In addition, when the feature amount is content, the feature amount matching judgment unit 56 judges, for example, whether the content sent by the user 10 matches the lyrics information of the content.
[0135] It should be noted that the feature quantity matching determination unit 56 may analyze the feature quantity such as the pitch, content, tone, chord length or rhythm of the content signal in real time based on the reproduced content signal, or may acquire the feature quantity from metadata included in the content. In addition, the method of determining whether the feature quantities of the sensor signal and the content signal match based on these feature quantities may be a rule-based algorithm or a method using a neural network, etc.
[0136] Then, the reproduction control unit 58 controls the reproduction of the content based on the output of the utterance detection unit 50 and the output of the feature amount matching determination unit 56. For example, in the case where the output of the utterance detection unit 50 is "present" and the output of the feature amount matching determination unit 56 is "mismatched", the reproduction control unit 58 determines that the user 10 uttered the utterance as a conversation instead of uttering the utterance in accordance with the content, and stops reproducing the content. On the other hand, in the case where the output of the utterance detection unit 50 is "present" and the output of the feature amount matching determination unit 56 is "matched", the reproduction control unit 58 determines that the user 10 uttered the utterance in accordance with the content, and performs control in a manner that does not stop the reproduced content.
[0137] Then, the output unit 60 of the acoustic device 100 outputs the voice of the content from the output unit 60 according to the result of the control of the reproduction control unit 58 .
[0138] It should be noted that the control of the content of the determination based on the matching of the feature amounts described in the second modification example may be performed in conjunction with the method of determining the timing synchronization described in the first modification example.
[0139] Next, we will refer to Fig.10 The procedure of the acoustic control process according to the second modification example is described. Fig.10 is a flowchart showing the procedure of the acoustic control process according to the second modification example.
[0140] When reproducing content, the acoustic device 100 operates a sound emission detection function and waits for detection of sound emission (step S41). That is, the acoustic device 100 performs sound emission detection by a sensor (step S42). In addition, the acoustic device 100 calculates a feature amount of a signal observed by the sensor (step S43).
[0141] In parallel, the acoustic device 100 calculates the feature amount in the content being reproduced (step S44). Then, the acoustic device 100 determines a match between the feature amount acquired from the sensor at predetermined time intervals and the content, for example (step S45).
[0142] During this time, the acoustic device 100 determines whether the utterance of the user 10 is detected (step S46). In the case where the utterance is not detected (step S46; No), the acoustic device 100 continues the utterance detection.
[0143] On the other hand, in the case where the utterance is detected (step S46; yes), the acoustic device 100 further determines whether the feature amount of the utterance matches the content (step S47). In the case where the feature amount does not match (step S47; no), the acoustic device 100 operates the dialogue function and stops reproducing the content (step S48). On the other hand, in the case where the feature amount matches (step S47; yes), the acoustic device 100 does not operate the dialogue function and ends the current processing. At this time, in the case where the content reproduction continues, the acoustic device 100 may return the processing to step S41.
[0144] As described above, the second determination unit 135 according to the second modification example determines whether to perform content-related control processing based on the degree of matching between the feature amount of the voice included in the content and the utterance as the correlation between the utterance and the content. For example, the second determination unit 135 determines whether to perform content-related control processing based on the degree of matching in at least one element of pitch, tonality, chord, and rhythm between the utterance included in the content as the feature amount and the voice.
[0145] Alternatively, the second determination unit 135 may determine whether to perform content-related control processing based on the degree of matching between the uttered content and the content of the voice included in the content as the feature amount. The matching of the content means, for example, the matching of the content uttered by the user 10 with the lyrics or script in the content, the matching of the semantic content of the two, etc.
[0146] As a result, the user 10 can control activation of a conversation function based on natural movements, such as practicing a song while listening to the content or singing a song while viewing the lyrics.
[0147] (1-4-3. Determination of input using voice and content)
[0148] Next, as a third modification example, a process will be described in which the acoustic device 100 processes both a signal related to utterance and a signal related to content as inputs to a machine learning model and determines whether to control a conversation function based on an output from the model.
[0149] In the embodiments and the like, an example has been described in which the acoustic device 100 detects utterance based on the sensor signal in the utterance detection unit 22. On the other hand, the acoustic device 100 according to the third modification distinguishes between utterance according to the content and other utterances because it has a configuration in which the utterance detection unit 62 uses both the sensor signal and the content signal as inputs. Fig.11 and subsequent figures describe this.
[0150] Fig.11 is a block diagram showing an overview of the acoustic control process according to the third modification. Fig.13 As shown, the acoustic device 100 continues the sound emission detection process by the sound emission detection unit 62 during content reproduction. In addition, the acoustic device 100 acquires information related to the content from the content reproduction apparatus 200 and inputs the acquired information to the sound emission detection unit 62.
[0151] For example, the utterance detection unit 62 may be configured by utilizing a machine learning model such as a neural network. For example, the acoustic device 100 prepares a large amount of combined data of the sensor signal and the content signal when the utterance is made according to the content, and the combined data of the sensor signal when the utterance is not made according to the content and the content signal as learning data. By learning such data, when a combination of the sensor signal and the content signal is newly input, the acoustic device 100 can distinguish whether the sensor signal is emitted according to the content, and acquire a model capable of detecting only the utterance that is initially detected and not emitted according to the content. Therefore, the utterance detection unit 62 can react only to the user's utterance of the element that is initially desired to be detected among various noises such as the utterance of the user, the ambient sound other than the utterance of the user, and the utterance of another person. For example, the acoustic device 100 acquires a model that outputs "1" only when it is discerned by machine learning that the sensor signal is not emitted according to the content and is initially detected, and outputs "0" in other cases (including the utterance according to the content).
[0152] The acoustic device 100 transmits the discrimination result by the sound utterance detection unit 62 to the reproduction control unit 64. For example, in the case where the output of the sound utterance detection unit 62 is a result of distinguishing that "the sensor signal (sound utterance) is not a sound utterance according to the content, but a sound utterance initially detected (conversation, etc.)", the reproduction control unit 64 determines that the user 10 does not utter a sound according to the content, and controls to stop the reproduced content.
[0153] Then, the output unit 66 of the acoustic device 100 outputs the voice of the content from the output unit 66 according to the result of the control of the reproduction control unit 64 .
[0154] Next, we will refer to Fig.12 The procedure of the acoustic control process according to the third modification example is described. Fig.12 is a flowchart showing the procedure of the acoustic control process according to the third modification example.
[0155] When reproducing content, the acoustic device 100 operates the sound emission detection function and waits for detection of sound emission (step S51). At this time, the acoustic device 100 detects sound emission by inputting both the sensor signal and the content signal to a neural network or the like (step S52).
[0156] Then, the acoustic device 100 determines whether a sound emission is detected (step S53). In the case where a sound emission is not detected (step S53; No), the acoustic device 100 continues the sound emission detection.
[0157] On the other hand, in the case where a normal utterance that is not an utterance according to the content is detected (step S53; Yes), the acoustic device 100 operates the conversation function and stops reproducing the content (step S54).
[0158] Here, an embodiment of a neural network according to a third modification will be described. Fig.13 is a diagram (1) for describing the configuration of a neural network according to a third modification.
[0159] Fig.13 The model 70 shown in FIG. 7 conceptually illustrates Fig.11 The structure of the neural network in the utterance detection unit 62 is shown in FIG. For example, the model 70 is a convolutional neural network including a plurality of convolutional layers and pooling layers. Fig.13In the embodiment of the invention, in the model 70, the time waveforms of the sensor signal 72 and the content signal 74 are bundled in the channel direction for input. In the model 70, the input signal passes through a plurality of convolutional layers and pooling layers, and a scalar value indicating the probability of the user 10 speaking (the provability of not generating a sound according to the content and not generating a normal sound according to the content) is output. According to the structure of the model 70, since the time waveforms of both the sensor signal 72 and the content signal 74 are input, the convolutional layer and the pooling layer can capture the correlation between the two and determine whether the sound is a sound according to the content or not a normal sound.
[0160] In addition, neural networks can have different Fig.13 The structure in Fig.14 Describe this. Fig.14 is a diagram (2) for describing the configuration of a neural network according to a third modification.
[0161] exist Fig.14 In the model 80 shown in , a plurality of convolution layers and pooling layers are applied to each of the sensor signal 82 and the content signal 84. Then, the model 80 has a structure in which the results of the sensor signal 82 and the content signal 84 to which the plurality of convolution layers and pooling layers are applied are combined in the middle part of the network. Since the features of the upper-level concept such as frequency or time change are extracted in the middle part of the network, according to this structure, by combining the two in the middle part of the network, by utilizing this upper-level concept, the similarity between the sensor signal 82 and the content signal 84 can be considered.
[0162] Note that although Fig.13 and Fig.14 In the model shown in , a time waveform is used as an input, but the structure of the model is not limited thereto. For example, the model can extract frequency information of each time segment in advance by short-time Fourier transform or the like and use the frequency information as an input. With this structure, conversion to frequency information can be performed in advance. Thus, it is possible to reduce the information to be learned by the neural network, and it is possible to perform learning with a relatively small amount of learning data. In addition, it is also effective to use a recursive neural network having a loop structure therein as a neural network.
[0163] As described above, the second determination unit 135 according to the second variant can determine the degree of matching between the feature quantity of the voice included in the content and the utterance by using a machine learning model, and the machine learning model is learned to learn the correlation between the feature quantity of the utterance and the content. Therefore, the acoustic device 100 can accurately determine whether the utterance is a conversation, even for utterances that are difficult to detect using only sensor signals, etc.
[0164] (1-4-4. Feedback from user operations)
[0165] Next, as a fourth modification example, a process in which the acoustic device 100 changes the determination criterion based on feedback from the user 10 will be described.
[0166] In the embodiments and the like, an embodiment has been described in which the acoustic device 100 controls the conversation function by determining whether the utterance of the user 10 is a conversation. In this process, when the user 10 speaks according to the content, for example, in the case where the user forgets the predetermined gesture or in the case where the utterance detection algorithm does not operate as expected, there is a possibility that the function operates in a mode that the user 10 does not expect. At this time, the user 10 provides feedback of the operating mode to the acoustic device 100, and the acoustic device 100 adjusts its behavior in response to the result. Thereby, the acoustic device 100 can further improve the user experience. Reference will be made to Fig.15 and subsequent figures describe this.
[0167] Fig.15 is a block diagram showing an outline of the acoustic control process according to the fourth modification. Fig.15 In the embodiment of FIG. 1 , it is assumed that the user 10 performs feedback by using a smartphone as an example of the content reproduction device 200 .
[0168] like Fig.15 As shown, the acoustic device 100 performs sound emission detection during content reproduction by the sound emission detection unit 210. In addition, the reproduction control unit 212 operates functions such as stopping content and capturing external sound based on the result of the sound emission detection.
[0169] At this time, the content reproducing apparatus 200 stops the content reproduction, and presents a question about whether the operation is correct from the presentation unit 216 to the user 10. The presentation unit 216 is, for example, a display screen of the content reproducing apparatus 200 connected to the acoustic device 100.
[0170] exist Fig.16 A display embodiment of the presentation unit 216 is shown in FIG. Fig.16 2 is a view for describing the presentation processing according to the fourth modification. Display embodiment 220 is an embodiment of a message displayed on the screen of the content reproduction device 200. As shown in display embodiment 220, the content reproduction device 200 displays whether the user 10 actually utters a sound when the sound utterance detection function is activated as a question to be presented to the user 10. In response to this, for example, the user 10 selects one of three options of "Yes", "No", and "Invalidate the detection of this content".
[0171] Note that the answer options are not limited to Fig.16, and may include, for example, "invalid for a specific period of time", "invalid until the power of the device is turned off", "permanently invalid", etc. In addition, although the question of the presentation unit 216 to the user 10 may be displayed each time a utterance is detected, frequent display may make the user 10 feel uncomfortable. Therefore, the content reproduction device 200 may calculate the confidence (probability) when discerning whether the utterance of the user 10 is a utterance according to the content or a utterance different from the content, and ask questions only when the calculation result is unclear. For example, the confidence of the discrimination is calculated based on the utterance detection processing described in each modification example, etc.
[0172] return Fig.15 , will continue to describe. The answer acquisition unit 218 acquires an answer from the user 10. Note that the answer acquisition unit 218 may be a processing unit that acquires an answer via a touch panel or the like included in the content reproduction device 200, or may be a processing unit that acquires an answer via a touch sensor or the like included in the acoustic device 100. At this time, in the case where the answer of the user 10 is "invalidate the detection of the content", even when an output indicating that a sound is detected is obtained from the sound detection unit 210, the reproduction control unit 212 does not stop reproducing the content until the reproduced content ends. It should be noted that in addition to invalidating the processing, the change in behavior according to the result of the answer of the user 10 may include making it difficult to detect the sound by increasing the threshold of the sound detection. Specifically, the acoustic device 100 may change the parameters of the algorithm of the neural network used for detection or change the algorithm itself.
[0173] Then, the output unit 214 of the acoustic device 100 outputs the voice of the content from the output unit 214 according to the result of the control of the reproduction control unit 212 .
[0174] Next, we will refer to Fig.17 The procedure of the acoustic control process according to the fourth modification example is described. Fig.17 is a flowchart showing the procedure of the acoustic control process according to the fourth modification.
[0175] When reproducing content, the acoustic device 100 operates the sound emission detection function and waits for detection of sound emission (step S61). Then, the acoustic device 100 determines whether sound emission is detected (step S62). In the case where sound emission is not detected (step S62; No), the acoustic device 100 continues sound emission detection.
[0176] On the other hand, in the case where the utterance is detected (step S62; Yes), the acoustic device 100 operates the dialogue function and stops reproducing the content (step S63). Subsequently, the acoustic device 100 displays a question to the user 10 to inquire whether it is appropriate to operate the dialogue function (step S64).
[0177] The acoustic device 100 obtains the user 10's reply to the utterance (step S65). In the case where the answer is "yes", the acoustic device 100 determines that such utterance detection is appropriate and ends the processing. On the other hand, in the case where the answer is "no", the acoustic device 100 determines that such utterance detection is inappropriate, reflects the result in the determination criteria of the algorithm related to the utterance detection (such as a neural network) (step S66), and ends the processing. In this case, the acoustic device 100 may perform recovery processing, such as resuming the reproduction of the content. In addition, in the case where the answer is "invalidate the detection of the content", the acoustic device 100 controls the function of invalidating the detection until the end of the current content (step S67), and ends the processing.
[0178] (1-4-5. Feedback through gestures)
[0179] It is to be noted that feedback from the user 10 may be performed not by screen operation of the user 10 but by a gesture of the user 10 or the like. For example, in a case where the conversation function is unintentionally operated in the acoustic device 100, the user 10 may give feedback indicating that the immediately preceding operation is inappropriate for the acoustic device 100 by making a gesture such as shaking the head. Fig.18 and subsequent figures describe this.
[0180] Fig.18 is a block diagram showing an overview of the acoustic control process according to the fifth modification. Fig.18 In the embodiment of the present invention, the acoustic device 100 includes an acceleration sensor 232 as an example of the sensor 230 .
[0181] The acoustic device 100 performs sound emission detection during content reproduction by the sound emission detection unit 234. Furthermore, the reproduction control unit 238 operates functions such as stopping content and capturing external sound based on a result of the sound emission detection.
[0182] Here, in the case where the operated function is inappropriate, the user 10 performs a predetermined gesture such as shaking his / her head sideways. The acoustic device 100 determines that the user 10 shakes his / her head by using the head shake detection unit 236, which obtains an output indicating that the acceleration sensor 232 detects an acceleration equal to or greater than a threshold.
[0183] The reproduction control unit 238 which obtains the output indicating that the user 10 shakes his head through the head shaking detection unit 236 redetermines that the operation of the dialogue function is inappropriate, resumes the reproduction of the content, and performs control in such a manner that the dialogue function is not operated until the current content ends.
[0184] Then, the output unit 240 of the acoustic device 100 outputs the voice of the content from the output unit 240 according to the result of the control of the reproduction control unit 238 .
[0185] Next, we will refer to Fig.19 The procedure of the acoustic control process according to the fifth modification example is described. Fig.19 is a flowchart showing the procedure of the acoustic control process according to the fifth modification example.
[0186] When reproducing content, the acoustic device 100 operates the sound emission detection function and waits for detection of sound emission (step S71). Then, the acoustic device 100 determines whether sound emission is detected (step S72). In the case where sound emission is not detected (step S72; No), the acoustic device 100 continues sound emission detection.
[0187] On the other hand, in the case where the utterance is detected (step S72; Yes), the acoustic device 100 operates the conversation function and stops reproducing the content (step S73). Subsequently, the acoustic device 100 determines whether the head shaking performed by the user 10 is detected (step S74).
[0188] In the case where the head shaking is not detected (step S74; No), in the case where the answer is "Yes", the acoustic device 100 determines that such sound emission detection is appropriate and ends the processing. On the other hand, in the case where the head shaking is detected (step S74; Yes), the acoustic device 100 determines that the user 10 has operated the undesired sound emission detection, invalidates the detection until the current content ends (step S75), and ends the processing.
[0189] As described above, in the case of performing the control process related to the content, the second determination unit 135 according to the fourth modification or the fifth modification acquires information related to the appropriateness / inappropriateness of the control process related to the content from the user 10, and determines the mode of the control process related to the content based on the acquired information. For example, as the mode of the control process related to the content, the second determination unit 135 resumes the reproduction of the paused content, adjusts the determination standard of the control process, and performs control in such a manner that the control process is not performed until the reproduced content ends.
[0190] The second determination unit 135 obtains information related to the appropriateness / inappropriateness of the control process related to the content based on the user operation on the device that controls the reproduction of the content. Specifically, the second determination unit 135 obtains feedback from the user 10 based on the touch operation of the user 10 on the screen of the smart phone, the operation on the touch sensor of the acoustic device 100, etc.
[0191] Furthermore, the second determination unit 135 may acquire information related to appropriateness / inappropriateness of control processing related to content by detecting a predetermined action of the user 10. Specifically, the second determination unit 135 acquires feedback from the user 10 based on a processing gesture such as a head shaking action of the user 10.
[0192] (1-4-6. Determination based on content attributes and viewing patterns)
[0193] Not only the processing based on the algorithm based on the utterance detection as described above, but also the control processing related to the content of the acoustic device 100 can be performed based on various conditions and usage patterns.
[0194] For example, the acoustic device 100 can automatically turn off the conversation function under the condition that the possibility of the user 10 uttering according to the content is considered high, and turn on the function only under the condition that the possibility is not considered high. As a result, the acoustic device 100 can improve the usability of the function.
[0195] As an embodiment, the acoustic device 100 may automatically turn on / off a function of performing processing corresponding to sound emission detection according to the type of content being reproduced. Note that the type of content includes all types of information for classifying content, such as the type of content, the attributes of content, and the file format of content.
[0196] For example, the acoustic device 100 determines the type of the content based on metadata of the reproduced content, etc. Specifically, the acoustic device 100 distinguishes the reproduced content according to each type such as "music (with singing voice)", "music (without singing voice)", "music (for karaoke)", "TV", "broadcast", "video", "comedy", and "language learning".
[0197] Then, in the case where the attribute corresponds to the predetermined condition, the acoustic device 100 automatically turns off the dialogue function. For example, the acoustic device 100 automatically turns off the dialogue function of the content having the attribute assuming that the user 10 makes a sound or makes a voice unintentionally according to the content. Specifically, in the case of reproducing the content having the attribute "comedy", the acoustic device 100 may automatically turn off the dialogue function because the user 10 may laugh according to the content. Note that the attribute for which the function is stopped may be set by the administrator of the acoustic device 100, or may be arbitrarily set by the user 10.
[0198] In addition, the acoustic device 100 can automatically turn on and off the function according to the type of application or device that reproduces the content. For example, in the case where the reproduced application is dedicated to music appreciation, the acoustic device 100 is set not to turn off the conversation function. On the other hand, in the case where the reproduced application is dedicated to communication (for example, assuming that the viewer can listen to songs based on moving images, etc.), the acoustic device 100 can perform settings in a manner that turns off the conversation function. Note that for which application the function is stopped can be set by the administrator of the acoustic device 100, or can be arbitrarily set by the user 10.
[0199] In addition, some content reproduction applications have a function of displaying a script (such as lyrics) of the content. Then, the possibility that the user 10 sings along with the content while using the lyrics display function is high. Thus, in the case of detecting a usage pattern of displaying lyrics in the application, the acoustic device 100 can perform control to turn off the dialogue function while continuously displaying the lyrics.
[0200] Alternatively, the acoustic device 100 may control the on / off of the conversation function according to the type of device that reproduces the content. For example, in the case where the content reproduction device 200 is a smart phone, the acoustic device 100 turns on the conversation function under the assumption that there is a high possibility of speaking to another person due to the user 10 going out or moving. On the other hand, in the case where the content reproduction device 200 is a fixed personal computer (PC) or the like, the acoustic device 100 may turn off the conversation function under the assumption that the user 10 does not go out or move and is unlikely to speak to another person. In addition, the acoustic device 100 may control the on / off of the conversation function based on the daily usage pattern. For example, for a user who mainly reproduces music content in a smart phone and mainly reproduces comedy or language learning content in a PC, the acoustic device 100 may improve the availability of the function by turning on / off the conversation function according to the device. Note that in which device the function is stopped may be set by the administrator of the acoustic device 100, or may be arbitrarily set by the user 10.
[0201] In addition, as other conditions, the acoustic device 100 may control the conversation function based on time, the user's location, the user's behavior, the surrounding noise level, etc. For example, the user's location is a situation such as whether the user is at home, at work, in a vehicle, etc., and is determined based on, for example, location information acquired by the acoustic device 100. The user's behavior is a situation of the user, such as the user's sitting, standing, walking, running, riding a bicycle, riding a train, etc., and is estimated by, for example, a vibration sensor or an acceleration sensor of the acoustic device 100. For example, when the user 10 is on a train or when the surrounding noise level is high, the user 10 is less likely to talk to people around while wearing headphones. Thus, the acoustic device 100 may perform control to turn off the conversation function.
[0202] In addition, the acoustic device 100 can detect whether there are people around the user 10, and when it is estimated that there is no person around the user 10, the conversation function can be turned off. The detection of the surrounding people can be performed by, for example, a motion sensor, a camera, etc. included in the acoustic device 100, a smart phone, etc., or can be realized by detecting the presence of a device used by the surrounding people via a public network. Note that the user 10 can arbitrarily set the on / off of the conversation function based on such conditions.
[0203] As described above, in the first determination process by the first determination unit 134 , the second determination unit 135 according to the acoustic device 100 may determine whether to perform a content-related control process based on information related to the content or information related to an application controlling the content.
[0204] For example, the second determination unit 135 may determine whether to perform the content-related control process based on the type of content viewed by the user 10. Alternatively, the second determination unit 135 may determine whether to perform the content-related control process based on the operation mode of the user 10 viewing the content with respect to the application or the content being displayed in the application.
[0205] As described above, the acoustic device 100 can further improve usability related to functions by controlling the functions based on various conditions.
[0206] (1-4-7. Mode of the function to be controlled)
[0207] In addition, the functions controlled by the acoustic device 100 are not limited to the pause of the content and the capture of the external sound, and may include various modes. For example, in the case where the type of content is "language learning" content, the acoustic device 100 may perform control in a manner that makes it easier to listen to the content and voice of the user 10 by monoaural reproduction or performs voice recognition of the uttered content of the user 10 and displays it on the screen of the smartphone.
[0208] (2. Other Implementation Methods)
[0209] In addition to each of the above-described embodiments, the processing according to each of the above-described embodiments may be performed in various different forms.
[0210] In addition, in each of the processes described in the above embodiments, all or part of the processes described to be automatically performed may be performed manually, or all or part of the processes described to be manually performed may be performed automatically by a known method. In addition, unless otherwise specified, the processing procedures, specific names, and information including various data and parameters shown in the above documents or drawings may be arbitrarily changed. For example, the various information shown in each drawing is not limited to the information shown.
[0211] In addition, each component of each of the devices shown is a functional concept and does not need to be physically configured in the manner shown. That is, the specific form of allocation / integration of each device is not limited to the form shown in the drawings, and all or part of it can be functionally or physically allocated / integrated in any unit according to various loads and usage conditions. For example, Figure 6The first determination unit 134 and the second determination unit 135 shown in the figure are not necessarily internal configurations of the acoustic device 100. That is, the sound emission detection information or gesture detection information by touch or the like acquired by the acoustic device 100 may be transmitted from the acoustic device 100 to the reproduction control device, and processing such as determination may be performed on the side of the reproduction control device.
[0212] Furthermore, the above-described embodiments and modifications may be arbitrarily combined within a range in which the processing contents do not conflict with each other.
[0213] Furthermore, the effects described in this specification are merely embodiments and not limitations, and there may be another effect.
[0214] (3. Effects of the Acoustic Device According to the Present Disclosure)
[0215] As described above, the acoustic device according to the present disclosure (acoustic device 100 in the embodiment) includes a detection unit (detection unit 133 in the embodiment), a first determination unit (first determination unit 134 in the embodiment), and a second determination unit (second determination unit 135 in the embodiment). The detection unit detects that the user utters a sound during content reproduction. In the case where the sound is detected by the detection unit, the first determination unit determines whether a control process related to the content can be performed. In the first determination process performed by the first determination unit, the second determination unit determines whether to perform a control process related to the content by analyzing whether the user has a predetermined action or the correlation between the sound and the content.
[0216] As described above, the acoustic device according to the present disclosure detects the user's utterance, and determines whether to execute the control of the content of the control associated with the detection according to the user's mode. Therefore, the acoustic device can perform flexible processing, for example, preventing the user's unintentional function from being operated, or activating the function according to the user's utterance. Thus, an acoustic control with a better user experience can be performed.
[0217] Furthermore, the second determination unit determines whether to perform a control process related to the content by analyzing whether the user touches the touch sensor unit to control the acoustic device as a predetermined action by the user.
[0218] For example, the second determination unit may determine whether to perform the control process related to the content by analyzing whether the user brings the hand close to a predetermined part of the acoustic device as a predetermined action of the user.
[0219] Alternatively, the second determination unit may determine whether to perform the control process related to the content by analyzing whether the user presses the operation unit to control the acoustic device or the content as a predetermined action of the user.
[0220] As described above, the acoustic device determines whether to perform content control processing according to the user's gesture, operation, etc. Therefore, the acoustic device can control a function based on the user's natural movement.
[0221] Furthermore, the second determination unit determines whether to perform the control process related to the content based on the degree of matching of the voice contained in the content with the timing of the utterance as the correlation between the utterance and the content.
[0222] For example, the second determination unit determines whether to perform the control process related to the content based on the degree of matching between the feature amount of the voice included in the content and the utterance as the correlation between the utterance and the content.
[0223] Specifically, the second determination unit determines whether to perform the control process related to the content based on the degree of matching in at least one element of pitch, tone, chord length, and rhythm between the utterance included in the content and the voice as the feature amount.
[0224] Alternatively, the second determination unit determines whether to perform the content-related control process based on a degree of matching between the content of the utterance and the content of the voice included in the content as the feature amount.
[0225] In addition, the second determination unit may determine the degree of matching between the feature amount of the voice included in the content and the utterance by using a machine learning model that is learned to learn the correlation between the feature amount of the utterance and the content.
[0226] As described above, the acoustic device determines whether to perform content control processing based on the correlation between the utterance and the content. Therefore, the acoustic device can perform control corresponding to the user's intention without causing the user to take any triggering action, such as performing a specific gesture.
[0227] In addition, in the case of executing the control processing related to the content, the second determination unit acquires information related to the appropriateness / inappropriateness of the control processing related to the content from the user, and determines a mode of the control processing related to the content based on the acquired information.
[0228] For example, the second determination unit acquires information related to appropriateness / inappropriateness of control processing related to the content based on a user operation on a device that controls reproduction of the content.
[0229] Alternatively, the second determination unit may acquire information related to appropriateness / inappropriateness of the control process related to the content by detecting a predetermined action of the user.
[0230] As described above, the acoustic device adjusts the determination of the control process based on the feedback received from the user. As a result, the acoustic device can perform a control process that more accurately reflects the user's intention.
[0231] In addition, the acoustic device may have the following configuration. That is, the detection unit detects that the user utters a sound during content reproduction. In the case where the utterance is detected by the detection unit, the first determination unit determines whether a control process related to the content can be performed. In the first determination process performed by the first determination unit, the second determination unit may determine whether to perform a control process related to the content based on information related to the content or information related to an application that controls the content.
[0232] For example, the second determination unit may determine whether to perform the content-related control process based on the type of content viewed by the user.
[0233] Alternatively, the second determination unit may determine whether to perform the control process related to the content based on the operation mode of the application by the user 10 who is viewing the content or the content displayed in the application.
[0234] As described above, the acoustic device determines whether to perform content control processing based on the characteristics and attributes of the content, the usage pattern of the application that reproduces the content, etc. As a result, since the acoustic device can perform control processing suitable for the user's usage, it is possible to provide the user with control processing that is highly usable for the user.
[0235] (4. Hardware Configuration)
[0236] For example, by having Fig. 20 The computer 1000 configured in the manner shown in embodies information equipment such as the acoustic device 100 according to each of the above-described embodiments. Hereinafter, the acoustic device 100 according to the embodiment will be described as an example. Fig. 20 1 is a hardware configuration diagram showing an embodiment of a computer 1000 that realizes the functions of the acoustic device 100. The computer 1000 includes a CPU 1100, a RAM 1200, a read-only memory (ROM) 1300, a hard disk drive (HDD) 1400, a communication interface 1500, and an input / output interface 1600. Each unit of the computer 1000 is connected by a bus 1050.
[0237] The CPU 1100 operates based on the program stored in the ROM 1300 or the HDD 1400, and controls each unit. For example, the CPU 1100 expands the program stored in the ROM 1300 or the HDD 1400 into the RAM 1200, and executes processing corresponding to various programs.
[0238] The ROM 1300 stores a boot program such as a basic input output system (BIOS) executed by the CPU 1100 during activation of the computer 1000 , a program depending on the hardware of the computer 1000 , and the like.
[0239] The HDD 1400 is a computer-readable recording medium that non-temporarily records a program executed by the CPU 1100 , data used by the program, and the like. More specifically, the HDD 1400 is a recording medium that records an acoustic control program according to the present disclosure, which is an example of the program data 1450 .
[0240] The communication interface 1500 is an interface for connecting the computer 1000 to an external network 1550 such as the Internet. For example, the CPU 1100 receives data from another device via the communication interface 1500 or transmits data generated by the CPU 1100 to another device.
[0241] The input / output interface 1600 is an interface for connecting the input / output device 1650 and the computer 1000. For example, the CPU 1100 receives data from an input device such as a keyboard or a mouse via the input / output interface 1600. In addition, the CPU 1100 transmits data to an output device such as a display, a speaker, or a printer via the input / output interface 1600. In addition, the input / output interface 1600 can be used as a medium interface for reading a program recorded on a predetermined recording medium (medium), etc. For example, the medium is an optical recording medium such as a digital versatile disk (DVD) or a phase change rewritable disk (PD), a magneto-optical recording medium such as a magneto-optical disk (MO), a magnetic tape medium, a magnetic recording medium, a semiconductor memory, etc.
[0242] For example, in the case where the computer 1000 is used as the acoustic device 100 according to the embodiment, the CPU 1100 of the computer 1000 implements the functions of the control unit 130 and the like by executing the acoustic control program loaded on the RAM 1200. In addition, the HDD 1400 stores the acoustic control program according to the present disclosure and the data in the storage unit 120. Note that the CPU 1100 reads the program data 1450 from the HDD 1400 and executes the program data 1450. However, in another embodiment, these programs may be acquired from another device via the external network 1550.
[0243] It should be noted that the present technology can also have the following configurations.
[0244] (1) An acoustic device comprising:
[0245] a detection unit for detecting that a user utters a sound during reproduction of the content;
[0246] a first determination unit that determines whether a control process related to the content can be performed when the detection unit detects the utterance; and
[0247] The second determination unit determines whether to execute the control process related to the content by analyzing whether there is a correlation between the user's predetermined action or utterance and the content in the first determination process of the first determination unit.
[0248] (2) The acoustic device according to (1), wherein
[0249] The second determination unit
[0250] Whether to perform a control process related to the content is determined by analyzing whether the user touches the touch sensor unit to control the acoustic device as a predetermined action by the user.
[0251] (3) The acoustic device according to (1) or (2), wherein
[0252] The second determination unit
[0253] Whether to execute the control process related to the content is determined by analyzing whether the user brings the hand close to a predetermined part of the acoustic device as a predetermined action of the user.
[0254] (4) The acoustic device according to any one of (1) to (3), wherein
[0255] The second determination unit
[0256] Whether to perform a control process related to the content is determined by analyzing whether the user presses the operation unit to control the acoustic device or the content as a predetermined action of the user.
[0257] (5) The acoustic device according to any one of (1) to (4), wherein
[0258] The second determination unit
[0259] Whether to perform content-related control processing is determined based on the degree of matching between the voice contained in the content and the timing of the utterance, as the correlation between the utterance and the content.
[0260] (6) The acoustic device according to any one of (1) to (5), wherein
[0261] The second determination unit
[0262] Whether to execute the control process related to the content is determined based on the degree of matching of the feature amount between the voice included in the content and the utterance as the correlation between the utterance and the content.
[0263] (7) The acoustic device according to (6), wherein
[0264] The second determination unit
[0265] Whether to execute a control process related to the content is determined based on a degree of matching in at least one of elements of pitch, tone, chord length, and rhythm between an utterance contained in the content and the voice as a feature amount.
[0266] (8) The acoustic device according to (6) or (7), wherein
[0267] The second determination unit
[0268] Whether to execute the control process related to the content is determined based on the degree of matching between the content of the utterance and the content of the voice contained in the content as a feature amount.
[0269] (9) The acoustic device according to any one of (6) to (8), wherein
[0270] The second determination unit
[0271] The degree of matching of feature quantities between speech and utterance included in the content is determined by using a machine learning model, and the machine learning model is learned to learn the correlation of feature quantities between the utterance and the content.
[0272] (10) The acoustic device according to any one of (1) to (9), wherein
[0273] The second determination unit
[0274] In the case of executing the content-related control process, information related to the appropriateness / inappropriateness of the content-related control process is acquired from the user, and a mode of the content-related control process is determined based on the acquired information.
[0275] (11) The acoustic device according to (10), wherein
[0276] The second determination unit
[0277] Based on a user operation on a device that controls reproduction of content, information related to appropriateness / inappropriateness of control processing related to the content is acquired.
[0278] (12) The acoustic device according to (10) or (11), wherein
[0279] The second determination unit
[0280] By detecting a predetermined action of the user, information related to appropriateness / inappropriateness of control processing related to the content is acquired.
[0281] (13) An acoustic device comprising:
[0282] a detection unit for detecting that a user utters a sound during reproduction of the content;
[0283] a first determination unit that determines whether a control process related to the content can be performed when the detection unit detects the utterance; and
[0284] The second determination unit determines whether to perform a control process related to the content based on the information related to the content or the information related to the application that controls the content in the first determination process of the first determination unit.
[0285] (14) The acoustic device according to (13), wherein
[0286] The second determination unit
[0287] Whether to execute content-related control processing is determined based on the type of content viewed by the user.
[0288] (15) The acoustic device according to (13) or (14), wherein
[0289] The second determination unit
[0290] Whether to execute control processing related to the content is determined based on the operation pattern of the application by the user who is viewing the content or the content being displayed in the application.
[0291] (16) An acoustic control method comprising:
[0292] detecting a user uttering a sound during reproduction of content;
[0293] determining whether content-related control processing can be performed if a sound is detected; and
[0294] In the process of determining whether the control process related to the content can be executed, determining whether the control process related to the content can be executed by analyzing whether there is a correlation between a predetermined action or utterance of the user and the content,
[0295] The detection and determination are performed by a computer.
[0296] (17) Acoustic Control Program
[0297] Computer, used as
[0298] a detection unit, detecting a user utterance during content reproduction,
[0299] a first determination unit that determines whether a control process related to the content can be performed when the detection unit detects the utterance; and
[0300] The second determination unit determines whether to execute the control process related to the content by analyzing whether there is a correlation between the user's predetermined action or utterance and the content in the first determination process of the first determination unit.
[0301] (18) An acoustic control method comprising:
[0302] detecting a user uttering a sound during reproduction of content;
[0303] determining whether content-related control processing can be performed if a sound is detected; and
[0304] In the process of determining whether the control process related to the content can be executed, determining whether to execute the control process related to the content is based on the information related to the content or the information related to the application controlling the content,
[0305] The detection and determination are performed by a computer.
[0306] (19) Acoustic Control Program
[0307] Computer, used as
[0308] a detection unit, detecting a user utterance during content reproduction,
[0309] a first determination unit that determines whether a control process related to the content can be performed when the detection unit detects the utterance; and
[0310] The second determination unit determines whether to perform a control process related to the content based on the information related to the content or the information related to the application that controls the content in the first determination process of the first determination unit.
[0311] Reference Numbers List
[0312] 10 Users
[0313] 100 Acoustic Devices
[0314] 110 Communication unit
[0315] 120 storage units
[0316] 130 Control unit
[0317] 131 Get Unit
[0318] 132 Reproduction control unit
[0319] 133 Detection Unit
[0320] 134 First Determination Unit
[0321] 135 Second Determination Unit
[0322] 200 Content reproduction device
Claims
1. An acoustic device, comprising: a detection unit to detect a user utterance during reproduction of the content; a first determination unit that determines whether control processing related to the content can be performed when the sound is detected by the detection unit; as well as A second determination unit determines whether to perform a control process related to the content by analyzing whether the user has a predetermined action or a correlation between the utterance and the content in the first determination process of the first determination unit.
2. The acoustic device according to claim 1, wherein: The second determining unit Whether to execute the control process related to the content is determined by analyzing whether the user touches a touch sensor unit for controlling an acoustic device as a predetermined action of the user.
3. The acoustic device according to claim 1, wherein: The second determining unit Whether to execute the control process related to the content is determined by analyzing whether the user brings a hand close to a predetermined part of the acoustic device as a predetermined action of the user.
4. The acoustic device according to claim 1, wherein: The second determining unit Whether to execute the control process related to the content is determined by analyzing whether the user presses an operation unit for controlling the acoustic device or the content as a predetermined action of the user.
5. The acoustic device according to claim 1, wherein: The second determining unit Whether to execute the control process related to the content is determined based on a degree of matching between a voice included in the content and the timing of the utterance as the correlation between the utterance and the content.
6. The acoustic device according to claim 1, wherein: The second determining unit Whether to execute the control process related to the content is determined based on a degree of matching of feature amounts between a voice included in the content and the utterance as the correlation between the utterance and the content.
7. The acoustic device according to claim 6, wherein: The second determining unit Whether to execute the control process related to the content is determined based on a degree of matching between the utterance and the voice included in the content in at least one element of pitch, timbre, chord, and rhythm as a feature amount.
8. The acoustic device according to claim 6, wherein: The second determining unit Whether to execute the control process related to the content is determined based on a degree of matching between the content of the utterance and the content of the voice included in the content as a feature amount.
9. The acoustic device according to claim 6, wherein: The second determining unit A degree of matching in the feature amount between the speech included in the content and the utterance is determined by using a machine learning model that has learned a correlation in the feature amount between the utterance and the content.
10. The acoustic device according to claim 1, wherein: The second determining unit In the case where the control process related to the content is executed, information related to appropriateness / inappropriateness of the control process related to the content is acquired from the user, and a mode of the control process related to the content is determined based on the acquired information.
11. The acoustic device according to claim 10, wherein: The second determining unit Information on appropriateness / inappropriateness of the control processing related to the content is acquired based on a user operation of a device that controls reproduction of the content.
12. The acoustic device according to claim 10, wherein: The second determining unit Information on appropriateness / inappropriateness of the control process related to the content is acquired by detecting a predetermined action of the user.
13. An acoustic device comprising: a detection unit to detect a user utterance during reproduction of the content; a first determination unit that determines whether control processing related to the content can be performed when the sound is detected by the detection unit; as well as A second determination unit determines whether to execute the control process related to the content based on the information related to the content or the information related to the application controlling the content in the first determination process of the first determination unit.
14. The acoustic device according to claim 13, wherein: The second determining unit Whether to execute the control process related to the content is determined based on the type of content viewed by the user.
15. The acoustic device according to claim 13, wherein: The second determining unit Whether to execute the control process related to the content is determined based on an operation mode of an application or content displayed in the application while the user is viewing the content.
16. An acoustic control method, comprising: detecting user utterances during content reproduction; determining whether control processing related to the content can be performed when the utterance is detected; as well as In the process of determining whether the control process related to the content can be executed, determining whether to execute the control process related to the content is performed by analyzing whether the user has a predetermined action or a correlation between the utterance and the content, The detection and determination are performed by a computer.
17. An acoustic control program, Computer, used as a detection unit, detecting a user utterance during content reproduction, a first determination unit that determines whether control processing related to the content can be performed when the utterance is detected by the detection unit; and A second determination unit determines whether to perform a control process related to the content by analyzing whether the user has a predetermined action or a correlation between the utterance and the content in the first determination process of the first determination unit.
18. An acoustic control method, comprising: detecting user utterances during content reproduction; determining whether control processing related to the content can be performed when the utterance is detected; as well as In the process of determining whether the control process related to the content can be executed, determining whether to execute the control process related to the content is based on information related to the content or information related to an application that controls the content, The detection and determination are performed by a computer.
19. An acoustic control program, Computer, used as a detection unit, detecting a user utterance during content reproduction, a first determination unit that determines whether control processing related to the content can be performed when the utterance is detected by the detection unit; and A second determination unit determines whether to execute the control process related to the content based on the information related to the content or the information related to the application controlling the content in the first determination process of the first determination unit.
Citation Information
Patent Citations
Playback device, headphone, and playback method
JP2011097268A