Interactive control methods, devices, equipment, storage media, and program products for electronic devices
Patent Information
- Application Number
- CN202510270070.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2026-09-11
AI Technical Summary
[0005]本公开实施例提供一种电子设备的交互控制方法、装置、设备、存储介质及程序产品,以解决出现交互错误概率较高的问题
[0012] Fifthly, embodiments of this disclosure provide a computer program product, including a computer program, and an interactive control method for an electronic device that implements the first aspect and various possible designs of the first aspect when the computer program is executed by a processor.
Smart Images

Figure CN122741827A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of headphone technology, and more particularly to an interactive control method, apparatus, device, storage medium, and program product for an electronic device. Background Technology
[0002] In the digital age, voice interaction technology is widely used in various electronic devices, with small electronic devices such as headphones becoming an important carrier.
[0003] In related technologies, wearable devices such as headphones can receive the user's voice and interact with the terminal device using the user's voice.
[0004] However, in public places, due to the complex environment, other people's voices may be mistakenly collected by electronic devices and used to interact with terminal devices, resulting in a high probability of interaction errors. Summary of the Invention
[0005] This disclosure provides an interactive control method, apparatus, device, storage medium, and program product for electronic devices to address the problem of a high probability of interactive errors.
[0006] In a first aspect, embodiments of this disclosure provide an interactive control method for an electronic device. The electronic device includes a processing unit, a bone conduction sensor, and a microphone. The method is applied to the processing unit and includes: acquiring an interaction state with a terminal device; receiving a voice signal through the microphone; if a bone conduction signal is received through the bone conduction sensor, determining whether the voice signal was emitted by the wearer based at least on the bone conduction signal; if the voice signal was emitted by the wearer, performing interactive processing with the terminal device based on the interaction state according to the voice signal.
[0007] Secondly, embodiments of this disclosure provide an interactive control device for an electronic device. The electronic device includes a processing unit, a bone conduction sensor, and a microphone. The device is applied to the processing unit and includes: a status acquisition module for acquiring the interaction status with a terminal device; a signal receiving module for receiving voice signals through the microphone; a sound source determination module for determining, at least based on the bone conduction signal, whether the voice signal was emitted by the wearer if a bone conduction signal is received through the bone conduction sensor; and an interaction processing module for performing interaction processing with the terminal device based on the interaction status according to the voice signal if it was emitted by the wearer.
[0008] Thirdly, embodiments of this disclosure provide an electronic device, including: a processing unit, a storage unit, a bone acoustic sensor, and a microphone;
[0009] The storage unit stores the instructions executed by the computer;
[0010] The processing unit executes computer execution instructions stored in the storage unit, causing the processing unit to combine the bone conduction sensor and microphone to execute the interactive control method of the first aspect and various possible electronic devices described above.
[0011] Fourthly, embodiments of this disclosure provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the interactive control method of an electronic device as described in the first aspect and various possible designs of the first aspect.
[0012] Fifthly, embodiments of this disclosure provide a computer program product, including a computer program, and an interactive control method for an electronic device that implements the first aspect and various possible designs of the first aspect when the computer program is executed by a processor.
[0013] The interactive control method, apparatus, device, storage medium, and program product for electronic devices provided in this embodiment obtain the interaction status with the terminal device and determine whether the voice signal is emitted by the wearer by combining the bone conduction signal. Interaction with the terminal device is only performed when the voice signal is emitted by the wearer, thereby reducing erroneous interactions between the electronic device and the terminal device. Attached Figure Description
[0014] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 A schematic diagram of a scenario for the interactive control method of the electronic device provided in this application;
[0016] Figure 2 A flowchart illustrating the interactive control method for an electronic device provided in an embodiment of this application;
[0017] Figure 3 This is a schematic diagram of the structure of an interactive control device for an electronic device provided in an embodiment of this application. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0019] In today's digital age, voice interaction technology has been widely and deeply applied in various electronic devices. Small electronic devices, especially headphones, have become an important carrier of this technology.
[0020] In related technological fields, wearable devices such as headphones have demonstrated powerful voice interaction capabilities. They can accurately receive voice information emitted by users and use it as a bridge to achieve efficient interaction between users and terminal devices. However, in public places such as subways, buses, and shopping malls, due to the complex environment and large flow of people, the voices of other people may be mistakenly collected by electronic devices and used to interact with terminal devices, resulting in a high probability of interaction errors.
[0021] To address the aforementioned technical problems, the inventors propose the following technical concept: by adding a bone conduction sensor to the electronic device, interaction with the terminal device is only initiated based on the voice signal received by the bone conduction sensor and the determination that it is a voice signal emitted by the wearer, and only then based on the voice signal and the existing interaction state.
[0022] This application is applied to scenarios involving the interactive control of electronic devices. It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with relevant laws, regulations and standards, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0023] Figure 1 A schematic diagram illustrating a scenario for the interactive control method of the electronic device provided in this application. For example... Figure 1 The scenario includes: electronic device 100 and terminal device 200. Electronic device 100 includes processing unit 101, bone acoustic sensor 102, microphone 103, wear detection sensor 104 and speaker 105.
[0024] In the specific implementation process, the electronic device 100 may include headphones such as separate headphones and over-ear headphones, or virtual reality devices, extended reality devices, etc.
[0025] Terminal device 200 includes devices such as mobile phones, computers, and laptops that can input and output data.
[0026] Electronic device 100 and terminal device 200 can be connected via wired or wireless means.
[0027] The processing unit 101 may include a processor for a VPU (Voice Pick-Up Unit) unit, or it may include any type of processor. The processing unit 101 can be connected to the bone conduction sensor 102, microphone 103, wear detection sensor 104, and speaker 105 via wired or wireless connection.
[0028] The bone acoustic sensor 102 may include any type of bone acoustic sensor such as an acceleration sensor, a vibration sensor, a piezoelectric bone acoustic sensor, a capacitive bone acoustic sensor, or an electromagnetic bone acoustic sensor.
[0029] Microphone 103 may include directional microphones, omnidirectional microphones, windproof omnidirectional microphones, etc.
[0030] The wearable detection sensor 104 may include optical sensors, pressure sensors, inertial sensors, etc.
[0031] Speaker 105 can be any type of speaker used for audio output.
[0032] Among them, the bone acoustic sensor 102 is used to transmit sound waves by utilizing the principle of human bone vibration. The chip inside the bone acoustic sensor converts the vibration signal generated by the sound wave into a bone acoustic signal.
[0033] Microphone 103 is used to collect the sound of a user speaking through the air.
[0034] Wear detection sensor 104 is used to detect whether electronic device 100 is being worn.
[0035] Speaker 105 is used to convert received audio information into electrical signals and output sound.
[0036] The electronic device 100 may further include a storage unit 106, which stores computer-executable instructions. The processing unit executes the computer-executable instructions stored in the storage unit, causing the processing unit to perform the interactive control method of the electronic device described in any of the embodiments below, in conjunction with the bone conduction sensor 102 and the microphone 103. The storage unit may include a read-only memory, a random access memory, etc.
[0037] It is understood that the scenarios illustrated in the embodiments of this application do not constitute a specific limitation on the interactive control method of electronic devices. In other feasible embodiments of this application, the above scenarios may include more or fewer components than illustrated, or combine some components, or split some components, or arrange different components, which can be determined according to the actual application scenario and are not limited here. Figure 1 The scenario shown can be implemented by hardware, software, or a combination of both.
[0038] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0039] Figure 2 This is a flowchart illustrating the interactive control method for an electronic device provided in an embodiment of this application. The execution entity of this embodiment may be... Figure 1 The processing unit 101 in this embodiment is not particularly limited in this respect. Figure 2 As shown, the method includes steps S201 to S204.
[0040] S201: Obtain the interaction status with the terminal device.
[0041] In this step, the pre-stored interaction state can be read from a preset location, or the interaction state with the terminal device can be obtained based on the interaction record with the terminal device. It can also include sending interaction state query information to the terminal device to obtain the interaction state reply from the terminal device.
[0042] For example, upon receiving a response from the voice assistant sent by the terminal device, the interaction state is "outputting the voice assistant's response." Alternatively, if the electronic device is not currently outputting any audio and has not sent a voice assistant wake-up command to the terminal device, the interaction state is "voice assistant waiting to be woken up." Yet another example is if the electronic device is currently playing audio; in this case, "playing audio" is defined as the interaction state.
[0043] S202: Receives voice signals via microphone.
[0044] This step may include receiving electrical signals output from a microphone to obtain a speech signal. It may also include receiving speech signals from multiple microphones and selecting the speech signal with the best reception.
[0045] S203: If a bone acoustic signal is received through the bone acoustic sensor, then at least based on the bone acoustic signal, determine whether the voice signal was emitted by the wearer.
[0046] In this step, the process of receiving bone conduction signals is similar to that of receiving speech signals, and will not be described again here. At least based on the bone conduction signals, it is determined whether the speech signal was emitted by the wearer. This can include confirming that the speech signal was emitted by the wearer upon receiving the bone conduction signals, or confirming that the bone conduction signals are received and their decibel or intensity exceeds a preset decibel or intensity threshold. It can also include calculating the similarity between the bone conduction signals and the speech signals; if the similarity is greater than a preset similarity threshold, then it is determined that the speech signal was emitted by the wearer. Before calculating the similarity between the bone conduction signals and the speech signals, noise reduction can be performed on both the bone conduction signals and the speech signals.
[0047] S204: If the voice signal is issued by the wearer, then the interaction with the terminal device will be processed based on the voice signal and the interaction status.
[0048] In this step, when the voice signal is emitted by the wearer, different interaction methods are used to interact with the terminal device for different interaction states, including but not limited to sending commands to the terminal device to activate the voice assistant, interrupt the voice assistant, change the volume of the terminal device, and send control commands for specific operations to the terminal device.
[0049] As can be seen from the description of the above embodiments, the embodiments of this disclosure obtain the interaction status with the terminal device and determine whether the voice signal is emitted by the wearer by combining the bone conduction signal. Only when the voice signal is emitted by the wearer will the interaction with the terminal device be performed, thereby reducing erroneous interactions between the electronic device and the terminal device.
[0050] In one possible implementation, in step S204 above, if the voice signal is issued by the wearer, the user interacts with the terminal device based on the interaction state according to the voice signal, including steps S204A1 and S204A2.
[0051] S204A1: If the voice signal is issued by the wearer and the interaction state is the voice assistant waiting to be woken up, then determine whether the voice signal contains the wake-up voice.
[0052] In this step, the voice assistant's pending wake-up state indicates that the voice assistant on the terminal device has not been started, there has been no interaction with the voice assistant on the terminal device, or no voice assistant response has been played. Determining whether the voice signal contains a wake-up voice includes inputting the voice signal into a speech recognition model, obtaining the speech recognition result output by the speech recognition model, and searching for the wake-up voice in the speech recognition result. Alternatively, it may include inputting the voice signal into a large language model and obtaining the large language model's determination result regarding whether it contains a wake-up voice.
[0053] S204A2: If a wake-up voice is included, then wake up the voice assistant.
[0054] In this step, waking up the voice assistant may include sending a wake-up command to the voice assistant or sending an instruction to the terminal device to open the voice assistant.
[0055] As can be seen from the description of the above embodiments, the embodiments of this disclosure reduce false wake-ups of the voice assistant by only waking it up when it is determined that the voice belongs to the wearer, even when the voice assistant is not woken up, and by continuously receiving the voice input from the user.
[0056] In one possible implementation, in step S204 above, if the voice signal is issued by the wearer, the user interacts with the terminal device based on the interaction state according to the voice signal, including step S204B.
[0057] S204B: If the message is sent by the wearer and the interaction status is outputting a response from the voice assistant, then stop outputting the voice assistant's response and send the voice signal to the terminal device so that the terminal device can input the voice signal into the voice assistant, allowing the voice assistant to generate a new response based on the voice signal.
[0058] In this step, stopping the output of the voice assistant's response includes sending an interrupt command to the terminal device, causing the terminal device to input a stop output command to the voice assistant, thereby stopping the voice assistant from outputting content and putting the voice assistant into listening mode. The electronic device is currently outputting the voice assistant's response, which may include the electronic device receiving a response message containing a voice assistant response identifier sent by the terminal device; the electronic device is outputting the response message during this step. Sending the voice signal to the terminal device may include converting the voice signal into text and then sending it to the terminal device, or it may include sending the voice signal directly to the terminal device, or it may include converting the voice signal into a format required by a preset protocol and then sending it to the terminal device.
[0059] In this context, the voice assistant is in listening mode, indicating that it is ready to receive new voice signals and generate new responses. The electronic device will continuously collect voice signals through the microphone and send them to the terminal device, so that the terminal device can input the voice signals into the voice assistant.
[0060] As can be seen from the description of the above embodiments, the present disclosure embodiments interrupt the output of the voice assistant only when it is determined that the voice signal is emitted by the wearer, thereby avoiding the situation where the voice assistant is interrupted due to environmental noise or voices from non-wearers.
[0061] In one possible implementation, in step S204 above, if the voice signal is issued by the wearer, then interactive processing is performed with the terminal device based on the interaction state according to the voice signal, including:
[0062] S204C: If the voice signal is emitted by the wearer, the interaction state is playing audio, and the duration of the voice signal is greater than the first preset duration, then the output volume of the audio is reduced.
[0063] In this step, the electronic device is outputting audio other than the voice assistant's response, such as music or audio from a video. Once it is determined that the audio voice signal is being emitted by the wearer, the volume is reduced.
[0064] The method of lowering the volume can include lowering the volume of the electronic device itself, or it can include sending a volume-lowering command to the terminal device, causing the terminal device to lower the volume in response to the received volume-lowering command. Lowering the volume can include lowering the volume by a preset percentage, or it can include lowering the volume to a preset level. The first preset duration is, for example, 4 seconds, 5 seconds, 6 seconds, etc.
[0065] As can be seen from the description of the above embodiments, the embodiments of this disclosure reduce the volume when the wearer's voice signal is received during audio output, thereby preventing interference with the user's communication with others. At the same time, it can reduce the interference of the output audio on the microphone's received voice signal and the interference of the output audio on the bone conduction signal.
[0066] In one possible implementation, if step S204C is initiated by the wearer, the interaction state is that audio is playing, and the duration of the voice signal is greater than the first preset duration, then after reducing the output volume of the audio, the system further includes steps S220A, S220B, or S220C.
[0067] S220A: If no new bone acoustic signal is received by the bone acoustic sensor within the second preset time after the output volume is reduced, the output volume will be restored to the level before it was reduced.
[0068] In this step, restoring the output volume to its previous level may include increasing the volume of the electronic device itself to its previous level, or it may include sending a volume restoration command to the terminal device so that the terminal device restores the volume to its previous level.
[0069] The second preset duration is, for example, 12 seconds, 13 seconds, 15 seconds, etc.
[0070] For example, if the volume is 50 before step S204C and 20 after step S204C, and no new bone acoustic signal is received by the bone acoustic sensor within a second preset time period (15 seconds) after step S204C, then the volume will be restored to 50. As another example, if the volume is 80% before step S204C and 30% after step S204C, and no new bone acoustic signal is received by the bone acoustic sensor within a second preset time period (14 seconds) after step S204C, then the volume will be restored to 80%.
[0071] S220B: If no new voice signal is received through the microphone within a second preset time after the output volume is reduced, the output volume will be restored to the level before it was reduced.
[0072] This step is similar to step S220A above, and will not be repeated here.
[0073] S220C: If a new voice signal is received through the microphone and a new bone conduction signal is received through the bone conduction sensor within a second preset time after the output volume is reduced, the system determines whether the new voice signal was emitted by the wearer based on the new bone conduction signal and the new voice signal. If it was not emitted by the wearer, the output volume is restored to its previous level.
[0074] In this step, based on the new bone acoustic signal and the new voice signal, it is determined whether the new voice signal was emitted by the wearer. This is similar to step S203 and will not be repeated here.
[0075] As can be seen from the description of the above embodiments, the present disclosure embodiments restore the output volume within a second preset time after reducing the output volume if no new bone conduction signal or new voice signal is received, or if the received new bone conduction signal or new voice signal is not emitted by the wearer. This eliminates the need for the user to actively adjust the volume and reduces the number of steps the user needs to take to change the volume.
[0076] In one possible implementation, if step S204C is initiated by the wearer, the interaction state is that audio is playing, and the duration of the voice signal is greater than the first preset duration, then after reducing the output volume of the audio, step S230 is further included.
[0077] S230: If a user inputs a volume adjustment command within a second preset time after the output volume is reduced, the target volume corresponding to the volume adjustment command is used to determine the new output volume.
[0078] In this step, the process of receiving volume adjustment commands input by the user can include receiving clicks or slides from the user on the volume adjustment button, or it can include receiving a volume adjustment command sent by the terminal device after the user triggers the volume adjustment button on the terminal device. The volume adjustment command can include a volume value, i.e., the target volume.
[0079] As can be seen from the description of the above embodiments, the embodiments of this disclosure use the volume adjusted by the user as the actual output volume used subsequently after the user changes the volume, thereby avoiding the automatic adjustment of the volume back to the level before the volume was reduced, and thus avoiding negative impacts on the user's operation.
[0080] In one possible implementation, in step S203 above, at least it is determined whether the voice signal was emitted by the wearer based on the bone conduction signal, including step S2031.
[0081] S2031: Based on bone conduction signals and voice signals, determine whether the voice signal was emitted by the wearer.
[0082] This step includes calculating the similarity between the bone conduction signal and the speech signal. If the similarity is greater than a preset similarity threshold, it is determined whether the speech signal was emitted by the wearer. Alternatively, it may include calculating the similarity between the bone conduction signal and the speech signal. If the similarity is greater than a preset similarity threshold, and the time difference between the start time of the bone conduction signal and the start time of the speech signal is less than a preset duration threshold, and the time difference between the end time of the bone conduction signal and the end time of the speech signal is also less than a preset duration threshold, it is determined whether the speech signal was emitted by the wearer.
[0083] As can be seen from the description of the above embodiments, the embodiments of this disclosure increase the accuracy of judgment by combining bone acoustic signals and speech signals.
[0084] In one possible implementation, the above embodiments include steps S240 to S242.
[0085] S240: Based on the voice signal, determine whether the voice signal was emitted by the wearer and obtain the first judgment result.
[0086] In this step, PVAD (Personalized Voice Activity Detection) can be used to identify the voice signal and determine whether the voice signal was emitted by the wearer, thus obtaining the first judgment result.
[0087] S241: Based on the bone conduction signal and the voice signal, determine whether the voice signal was emitted by the wearer, and obtain the second judgment result.
[0088] This step is similar to step S202 above, and will not be repeated here.
[0089] S242: If the first judgment result is negative and the second judgment result is positive, then interactive processing is performed with the terminal device based on the voice signal and the interactive state.
[0090] The interaction with the terminal device in this step is similar to step S204 above, and will not be repeated here.
[0091] S243: If the first judgment result is yes and the second judgment result is no, then no interaction processing with the terminal device will be performed based on the voice signal and the interaction state.
[0092] The interaction with the terminal device in this step is similar to step S204 above, and will not be repeated here.
[0093] In one possible implementation, if both judgment results are yes, then interaction with the terminal device is also performed based on the voice signal.
[0094] As can be seen from the description of the above embodiments, the embodiments of this disclosure, in addition to using voice signals to determine whether they are emitted by the wearer, also combine voice signals and bone acoustic signals to determine whether the voice signals are emitted by the wearer, and take the judgment result of combining bone acoustic signals as the main basis to achieve the effect of increasing the accuracy of judgment.
[0095] In one possible implementation, step S2031 above, determining whether the voice signal was emitted by the wearer based on the bone acoustic signal and the voice signal, includes steps S20311 and S20312.
[0096] S20311: Calculate the correlation between bone acoustic signals and speech signals.
[0097] This step includes calculating the correlation between single frames of the bone acoustic signal and the speech signal, and using the correlation between single frames to calculate the overall correlation between the bone acoustic signal and the speech signal.
[0098] S20312: If the correlation is greater than the preset correlation threshold, then the voice signal is determined to be emitted by the wearer; otherwise, the voice signal is determined not to be emitted by the wearer.
[0099] In this step, the correlation threshold can be preset by staff based on experimental data or empirical parameters.
[0100] As can be seen from the description of the above embodiments, the embodiments of this disclosure calculate the correlation between bone acoustic signals and voice signals. Only when the correlation is greater than a preset correlation threshold is it determined that the voice signal was emitted by the wearer, so as to facilitate subsequent interactive processing based on whether the voice signal was emitted by the wearer.
[0101] In one possible implementation, after receiving the voice signal via the microphone in step S202, the method further includes step S250.
[0102] S250: If no bone acoustic signal is received through the bone acoustic sensor, the speech signal is ignored.
[0103] In this step, failure to receive bone acoustic signals through the bone acoustic sensor can include either no bone acoustic signals being received from the sensor at all, or the intensity of the bone acoustic signals received from the sensor being lower than a preset intensity threshold. Ignoring voice signals can include not processing the voice signals or performing a process to determine whether the voice signals were spoken by the wearer.
[0104] As can be seen from the description of the above embodiments, the embodiments of this disclosure achieve the avoidance of environmental noise or voice interference from non-wearers by ignoring the voice signal when no bone conduction signal is received and not interacting with the terminal device.
[0105] In one possible implementation, the electronic device further includes a wear detection sensor. The method provided in any of the above embodiments further includes step S260.
[0106] S260: If a wearing signal is received from the wearing detection sensor, then interactive processing is performed with the terminal device based on the voice signal and the interaction state.
[0107] In this step, the wearing signal can include a signal sent by the wearing detection sensor when it detects proximity to an object, a signal sent after detecting infrared light, or a signal generated by the wearing detection sensor when it detects pressure. The process of interacting with the terminal device based on the voice signal and the interaction state is similar to step S204 above, and will not be described again here.
[0108] As can be seen from the description of the above embodiments, the embodiments of this disclosure interact with the terminal device when a wearing signal is received from the wearing detection sensor, thereby avoiding interaction with the terminal device when the device is not worn.
[0109] Figure 3 This is a schematic diagram of the structure of an interactive control device for an electronic device provided in an embodiment of this application. The interactive control device for the electronic device is applied to the processing unit of the electronic device. For example... Figure 3 As shown, the interactive control device 300 of the electronic device includes: a status acquisition module 301, a signal receiving module 302, a sound source judgment module 303, and an interactive processing module 304.
[0110] The status acquisition module 301 is used to acquire the interaction status with the terminal device;
[0111] Signal receiving module 302 is used to receive voice signals through a microphone;
[0112] The sound source determination module 303 is used to determine whether the voice signal is emitted by the wearer if a bone sound signal is received through the bone sound sensor.
[0113] The interaction processing module 304 is used to perform interaction processing with the terminal device based on the interaction state according to the voice signal if it is issued by the wearer.
[0114] In one possible implementation, the interaction processing module 304 is specifically used to determine whether the voice signal contains wake-up voice if it is issued by the wearer and the interaction state is the voice assistant waiting to be woken up; if it contains wake-up voice, then wake up the voice assistant.
[0115] In one possible implementation, the interaction processing module 304 is specifically used to stop outputting the voice assistant's reply if the message is sent by the wearer and the interaction state is outputting the voice assistant's reply, and send the voice signal to the terminal device so that the terminal device inputs the voice signal into the voice assistant, so that the voice assistant generates a new reply based on the voice signal.
[0116] In one possible implementation, the interaction processing module 304 is specifically used to reduce the output volume of the audio if the voice signal is emitted by the wearer, the interaction state is playing audio, and the duration of the voice signal is greater than a first preset duration.
[0117] In one possible implementation, the interactive control device 300 of the electronic device further includes a volume restoration module 305.
[0118] The volume restoration module 305 is used to restore the output volume to its previous level if no new bone acoustic signal is received through the bone acoustic sensor within a second preset time period after the output volume is reduced; or, if no new voice signal is received through the microphone within a second preset time period after the output volume is reduced, restore the output volume to its previous level; or, if a new voice signal is received through both the microphone and the bone acoustic sensor within a second preset time period after the output volume is reduced, determine whether the new voice signal was emitted by the wearer based on the new bone acoustic signal and the new voice signal; if it was not emitted by the wearer, restore the output volume to its previous level.
[0119] In one possible implementation, the interactive control device 300 of the electronic device further includes a volume fixing module 306.
[0120] The volume fixing module 306 is used to determine the new output volume by using the target volume corresponding to the volume adjustment command if a user inputs a volume adjustment command within a second preset time after the output volume is reduced.
[0121] In one possible implementation, the sound source determination module 303 is specifically used to determine whether the voice signal is emitted by the wearer based on the bone conduction signal and the voice signal.
[0122] In one possible implementation, the sound source determination module 303 is specifically used to determine whether the voice signal is emitted by the wearer based on the voice signal, and obtain a first determination result; determine whether the voice signal is emitted by the wearer based on the bone conduction signal and the voice signal, and obtain a second determination result; if the first determination result is negative and the second determination result is positive, then the interaction state is changed according to the voice signal to obtain an updated interaction state; if the first determination result is positive and the second determination result is negative, then no interaction processing is performed with the terminal device based on the interaction state according to the voice signal.
[0123] In one possible implementation, the sound source determination module 303 is specifically used to calculate the correlation between the bone acoustic signal and the speech signal; if the correlation is greater than a preset correlation threshold, it is determined that the speech signal was emitted by the wearer, otherwise it is determined that the speech signal was not emitted by the wearer.
[0124] In one possible implementation, the interactive control device 300 of the electronic device further includes a voice ignore module 307.
[0125] The voice ignoring module 307 is used to ignore the voice signal if no bone acoustic signal is received through the bone acoustic sensor.
[0126] In one possible implementation, the electronic device further includes a wear detection sensor; the interaction processing module 304 is also configured to perform interaction processing with the terminal device based on the interaction state according to the voice signal if a wear signal sent by the wear detection sensor is received.
[0127] The apparatus provided in this embodiment can be used to execute the technical solutions of the above method embodiments. Its implementation principle and technical effects are similar, and will not be described again here.
[0128] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable storage medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device, or installed from a storage unit. When the computer program is executed by a processing device, it performs the functions defined above in the methods of the embodiments of this application.
[0129] It should be noted that the computer-readable storage medium described above in this application can be a computer-readable signal medium, a computer storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage unit, a magnetic storage unit, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable storage medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0130] The aforementioned computer-readable storage medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0131] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the method shown in the above embodiments.
[0132] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof. These programming languages include object-oriented programming languages—such as Java, Smalltalk, and C++—and conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0133] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0134] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the units are not necessarily limiting of the module itself; for example, a status acquisition module can also be described as an "interactive status acquisition module".
[0135] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0136] This application also provides a computer-readable storage medium storing computer-executable instructions. When a processing unit executes the computer-executable instructions, it implements the technical solution of the interactive control method of the electronic device in any of the above embodiments. Its implementation principle and beneficial effects are similar to those of the interactive control method of the electronic device. Please refer to the implementation principle and beneficial effects of the interactive control method of the electronic device. It will not be repeated here.
[0137] In the context of this application, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0138] This application also provides a computer program product, including a computer program. When the computer program is executed by a processing unit, it implements the technical solution of the interactive control method of the electronic device in any of the above embodiments. Its implementation principle and beneficial effects are similar to those of the interactive control method of the electronic device. Please refer to the implementation principle and beneficial effects of the interactive control method of the electronic device. It will not be repeated here.
[0139] In a first aspect, according to one or more embodiments of the present disclosure, an interactive control method for an electronic device is provided. The electronic device includes a processing unit, a bone conduction sensor, and a microphone. The method is applied to the processing unit and includes: acquiring an interaction state with a terminal device; receiving a voice signal through the microphone; if a bone conduction signal is received through the bone conduction sensor, determining whether the voice signal was emitted by the wearer based at least on the bone conduction signal; if the voice signal was emitted by the wearer, performing interactive processing with the terminal device based on the interaction state according to the voice signal.
[0140] According to one or more embodiments of this disclosure, if the voice signal is issued by the wearer, the interaction with the terminal device is performed based on the interaction state, including: if the voice signal is issued by the wearer and the interaction state is a voice assistant waiting to be woken up, then it is determined whether the voice signal contains a wake-up voice; if it contains a wake-up voice, then the voice assistant is woken up.
[0141] According to one or more embodiments of this disclosure, if the voice signal is issued by the wearer, interactive processing is performed with the terminal device based on the interaction state, including: if the voice signal is issued by the wearer and the interaction state is outputting a response from the voice assistant, then the outputting of the voice assistant's response is stopped, and the voice signal is sent to the terminal device so that the terminal device inputs the voice signal into the voice assistant, and the voice assistant generates a new response based on the voice signal.
[0142] According to one or more embodiments of this disclosure, if the voice signal is issued by the wearer, interactive processing is performed with the terminal device based on the interactive state, including: if the voice signal is issued by the wearer, the interactive state is playing audio, and the duration of the voice signal is greater than a first preset duration, then the output volume of the audio is reduced.
[0143] According to one or more embodiments of this disclosure, after reducing the audio output volume if the voice signal is emitted by the wearer, the interaction state is audio playback, and the duration of the voice signal is greater than a first preset duration, the method further includes: if no new bone acoustic signal is received by the bone acoustic sensor within a second preset duration after the output volume is reduced, then the output volume is restored to the level before the reduction; or, if no new voice signal is received by the microphone within a second preset duration after the output volume is reduced, then the output volume is restored to the level before the reduction; or, if a new voice signal is received by both the microphone and the bone acoustic sensor within a second preset duration after the output volume is reduced, then the method determines whether the new voice signal was emitted by the wearer based on the new bone acoustic signal and the new voice signal; if it was not emitted by the wearer, then the output volume is restored to the level before the reduction.
[0144] According to one or more embodiments of this disclosure, if the voice signal is issued by the wearer, the interaction state is playing audio, and the duration of the voice signal is greater than a first preset duration, after reducing the output volume of the audio, the method further includes: if a user inputs a volume adjustment command within a second preset duration after reducing the output volume, then the target volume corresponding to the volume adjustment command is used to determine the new output volume.
[0145] According to one or more embodiments of this disclosure, determining whether a voice signal is emitted by the wearer is based at least on the bone acoustic signal, including: determining whether the voice signal is emitted by the wearer based on the bone acoustic signal and the voice signal.
[0146] According to one or more embodiments of this disclosure, the method includes: determining whether the voice signal is emitted by the wearer based on the voice signal, and obtaining a first determination result; determining whether the voice signal is emitted by the wearer based on the bone conduction signal and the voice signal, and obtaining a second determination result; if the first determination result is negative and the second determination result is positive, then changing the interaction state based on the voice signal to obtain an updated interaction state; if the first determination result is positive and the second determination result is negative, then not performing interaction processing with the terminal device based on the interaction state according to the voice signal.
[0147] According to one or more embodiments of this disclosure, determining whether a voice signal is emitted by the wearer based on bone acoustic signals and voice signals includes: calculating the correlation between bone acoustic signals and voice signals; if the correlation is greater than a preset correlation threshold, then determining that the voice signal is emitted by the wearer, otherwise determining that the voice signal is not emitted by the wearer.
[0148] According to one or more embodiments of this disclosure, after receiving a voice signal via a microphone, the method further includes: ignoring the voice signal if no bone acoustic signal is received via a bone acoustic sensor.
[0149] According to one or more embodiments of this disclosure, an electronic device further includes a wear detection sensor; the method further includes: if a wear signal sent by the wear detection sensor is received, then performing interactive processing with a terminal device based on an interaction state according to a voice signal.
[0150] Secondly, according to one or more embodiments of this disclosure, an interactive control device for an electronic device is provided. The electronic device includes a processing unit, a bone conduction sensor, and a microphone. The device is applied to the processing unit and includes: a status acquisition module for acquiring the interaction status with a terminal device; a signal receiving module for receiving a voice signal through the microphone; a sound source determination module for determining, at least based on the bone conduction signal, whether the voice signal was emitted by the wearer if a bone conduction signal is received through the bone conduction sensor; and an interaction processing module for performing interaction processing with the terminal device based on the interaction status according to the voice signal if it was emitted by the wearer.
[0151] According to one or more embodiments of this disclosure, the interaction processing module is specifically used to determine whether the voice signal contains wake-up voice if it is issued by the wearer and the interaction state is a voice assistant waiting to be woken up; if it contains wake-up voice, then wake up the voice assistant.
[0152] According to one or more embodiments of this disclosure, the interaction processing module is specifically configured to, if the voice signal is issued by the wearer and the interaction state is outputting a response from the voice assistant, stop outputting the response from the voice assistant and send the voice signal to the terminal device, so that the terminal device inputs the voice signal into the voice assistant, and the voice assistant generates a new response based on the voice signal.
[0153] According to one or more embodiments of this disclosure, the interaction processing module is specifically used to reduce the output volume of the audio if the voice signal is emitted by the wearer, the interaction state is playing audio, and the duration of the voice signal is greater than a first preset duration.
[0154] According to one or more embodiments of the present disclosure, the interactive control device for an electronic device further includes a volume recovery module 305.
[0155] The volume restoration module is used to restore the output volume to its previous level if no new bone acoustic signal is received through the bone acoustic sensor within a second preset time period after the output volume has been reduced; or, if no new voice signal is received through the microphone within a second preset time period after the output volume has been reduced, the output volume is restored to its previous level; or, if a new voice signal is received through both the microphone and the bone acoustic sensor within a second preset time period after the output volume has been reduced, the module determines whether the new voice signal was emitted by the wearer based on the new bone acoustic signal and the new voice signal; if it was not emitted by the wearer, the output volume is restored to its previous level.
[0156] According to one or more embodiments of the present disclosure, the interactive control device for an electronic device further includes: a volume fixing module.
[0157] The volume fixing module is used to determine the new output volume by using the target volume corresponding to the volume adjustment command if a user inputs a volume adjustment command within a second preset time after the output volume is reduced.
[0158] According to one or more embodiments of this disclosure, the sound source determination module is specifically used to determine whether the voice signal is emitted by the wearer based on the bone conduction signal and the voice signal.
[0159] According to one or more embodiments of this disclosure, the sound source determination module is specifically used to determine whether the voice signal is emitted by the wearer based on the voice signal, and obtain a first determination result; determine whether the voice signal is emitted by the wearer based on the bone conduction signal and the voice signal, and obtain a second determination result; if the first determination result is negative and the second determination result is positive, then the interaction state is changed according to the voice signal to obtain an updated interaction state; if the first determination result is positive and the second determination result is negative, then no interaction processing is performed with the terminal device based on the interaction state according to the voice signal.
[0160] According to one or more embodiments of this disclosure, the sound source determination module is specifically used to calculate the correlation between the bone acoustic signal and the speech signal; if the correlation is greater than a preset correlation threshold, it is determined that the speech signal was emitted by the wearer, otherwise it is determined that the speech signal was not emitted by the wearer.
[0161] According to one or more embodiments of the present disclosure, the interactive control device for an electronic device further includes: a voice ignore module.
[0162] The voice ignore module is used to ignore the voice signal if no bone acoustic signal is received through the bone acoustic sensor.
[0163] According to one or more embodiments of this disclosure, the electronic device further includes a wear detection sensor; the interaction processing module is further configured to, if it receives a wear signal sent by the wear detection sensor, perform interaction processing with the terminal device based on the interaction state according to the voice signal.
[0164] Thirdly, according to one or more embodiments of the present disclosure, an electronic device is provided, including: a processing unit, a storage unit, a bone acoustic sensor, and a microphone;
[0165] The storage unit stores the instructions executed by the computer;
[0166] The processing unit executes computer execution instructions stored in the storage unit, causing the processing unit to combine the bone conduction sensor and microphone to execute the interactive control method of the first aspect and various possible electronic devices described above.
[0167] According to one or more embodiments of this disclosure, the electronic device further includes a wear detection sensor.
[0168] Fourthly, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided that stores computer-executable instructions, which, when executed by a processor, implement the interactive control method of an electronic device as described in the first aspect and various possible designs of the first aspect.
[0169] Fifthly, according to one or more embodiments of this disclosure, a computer program product is provided, including a computer program, which, when executed by a processor, implements the interactive control method of an electronic device as described in the first aspect above and various possible designs of the first aspect.
[0170] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0171] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0172] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. An interactive control method for an electronic device, characterized in that, The electronic device includes a processing unit, a bone conduction sensor, and a microphone. The method is applied to the processing unit and includes: Obtain the interaction status with the terminal device; Voice signals are received through the microphone; If a bone acoustic signal is received through the bone acoustic sensor, then at least based on the bone acoustic signal, it is determined whether the voice signal was emitted by the wearer; If the voice signal is issued by the wearer, then the user interacts with the terminal device based on the interaction state according to the voice signal.
2. The method according to claim 1, characterized in that, If the voice signal is emitted by the wearer, then the user interacts with the terminal device based on the voice signal and the interaction state, including: If the voice signal is issued by the wearer and the interaction state is a voice assistant waiting to be woken up, then determine whether the voice signal contains a wake-up voice. If the wake-up voice is included, then the voice assistant is activated.
3. The method according to claim 1, characterized in that, If the voice signal is emitted by the wearer, then the user interacts with the terminal device based on the voice signal and the interaction state, including: If the message is sent by the wearer and the interaction status is outputting a response from the voice assistant, then the output of the voice assistant's response is stopped, and the voice signal is sent to the terminal device so that the terminal device inputs the voice signal into the voice assistant, allowing the voice assistant to generate a new response based on the voice signal.
4. The method according to claim 1, characterized in that, If the voice signal is emitted by the wearer, then the user interacts with the terminal device based on the voice signal and the interaction state, including: If the message is sent by the wearer, the interaction state is that audio is playing, and the duration of the voice signal is greater than a first preset duration, then the output volume of the audio is reduced.
5. The method according to claim 4, characterized in that, After reducing the audio output volume when the interaction is initiated by the wearer, the interaction state is audio playback, and the duration of the voice signal is greater than a first preset duration, the method further includes: If no new bone acoustic signal is received by the bone acoustic sensor within the second preset time period after the output volume is reduced, the output volume is restored to its previous level; or, If no new voice signal is received through the microphone within the second preset time period after the output volume is reduced, the output volume is restored to its previous level; or, If a new voice signal is received through the microphone and a new bone conduction signal is received through the bone conduction sensor within a second preset time after the output volume is reduced, then based on the new bone conduction signal and the new voice signal, it is determined whether the new voice signal was emitted by the wearer; if it was not emitted by the wearer, then the output volume is restored to the level before it was reduced.
6. The method according to claim 4, characterized in that, After reducing the audio output volume when the interaction is initiated by the wearer, the interaction state is audio playback, and the duration of the voice signal is greater than a first preset duration, the method further includes: If a user inputs a volume adjustment command within a second preset time period after the output volume is reduced, the target volume corresponding to the volume adjustment command is used to determine the new output volume.
7. The method according to claim 1, characterized in that, The method of determining whether the voice signal is emitted by the wearer based at least on the bone conduction signal includes: Based on the bone conduction signal and the voice signal, determine whether the voice signal was emitted by the wearer.
8. The method according to claim 7, characterized in that, The method includes: Based on the voice signal, determine whether the voice signal was emitted by the wearer, and obtain a first determination result; Based on the bone acoustic signal and the voice signal, it is determined whether the voice signal was emitted by the wearer, and a second determination result is obtained; If the first judgment result is negative and the second judgment result is positive, then the interaction state is changed according to the voice signal to obtain an updated interaction state; If the first judgment result is yes and the second judgment result is no, then no interaction processing with the terminal device will be performed based on the voice signal and the interaction state.
9. The method according to claim 7 or 8, characterized in that, The step of determining whether the voice signal is emitted by the wearer based on the bone conduction signal and the voice signal includes: Calculate the correlation between the bone acoustic signal and the speech signal; If the correlation is greater than a preset correlation threshold, then the voice signal is determined to have been emitted by the wearer; otherwise, the voice signal is determined not to have been emitted by the wearer.
10. The method according to any one of claims 1 to 7, characterized in that, After receiving the voice signal through the microphone, the method further includes: If no bone acoustic signal is received through the bone acoustic sensor, the speech signal is ignored.
11. The method according to any one of claims 1 to 7, characterized in that, The electronic device further includes a wear detection sensor; the method further includes: If a wearing signal is received from the wearing detection sensor, then interactive processing is performed with the terminal device based on the voice signal and the interaction state.
12. An interactive control device for an electronic device, characterized in that, The electronic device includes a processing unit, a bone conduction sensor, and a microphone. The device is applied to the processing unit and includes: The status acquisition module is used to acquire the interaction status with the terminal device; A signal receiving module is used to receive voice signals through the microphone; The sound source determination module is used to determine, at least based on the bone acoustic signal, whether the voice signal is emitted by the wearer if a bone acoustic signal is received through the bone acoustic sensor. An interaction processing module is used to perform interaction processing with the terminal device based on the interaction state according to the voice signal if it is issued by the wearer.
13. An electronic device, characterized in that, include: Processing unit, storage unit, bone conduction sensor and microphone; The storage unit stores computer-executed instructions; The processing unit executes the computer execution instructions stored in the storage unit, causing the processing unit to perform the interactive control method of the electronic device as described in any one of claims 1 to 11 in conjunction with the bone acoustic sensor and the microphone.
14. The electronic device according to claim 13, characterized in that, It also includes wearing detection sensors.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, implement the interactive control method of the electronic device as described in any one of claims 1 to 11.
16. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method for interactive control of an electronic device as described in any one of claims 1 to 11.