Voice control system and method, voice kit, bone conduction and voice processing device
By collecting voice input through a bone conduction device and combining it with semantic analysis by a voice processing device, the problem of microphones having difficulty recognizing voice commands in noisy environments is solved, and accurate voice control in noisy environments is achieved.
Patent Information
- Application Number
- CN201911378410.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-12-27
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2039-12-27
AI Technical Summary
When existing intelligent voice devices use microphones to receive voice commands, they are unable to block voice commands from non-device users and have difficulty accurately recognizing voice commands when the ambient noise is high.
The bone conduction device is used as the voice collection entrance. The user's voice input is collected through the bone conduction sensor and sent to the voice processing device for semantic analysis and command issuance, thereby achieving accurate control of various devices.
It achieves clear sound restoration in noisy environments, shields the noise impact on others, and improves the accuracy of voice command recognition and the intelligence of device use.
Smart Images

Figure CN113053371B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular to a voice control system and method, as well as a corresponding voice kit, bone conduction and voice processing device. Background Art
[0002] With the popularization and development of smart technology, voice control of various devices has become a standard. For example, in existing technologies, voice control can be achieved through smart speakers that serve as control nodes in the home or appliances that have voice interaction functions.
[0003] Existing intelligent voice devices typically use microphones to receive voice commands. However, using microphones to receive voice commands cannot block voice commands from non-device users, and it is difficult to accurately recognize voice commands in noisy environments.
[0004] To this end, a reliable and accurate voice control solution is needed. Summary of the Invention
[0005] In order to solve at least one of the above problems, the present invention proposes a solution that uses a bone conduction device as a voice collection entrance and sends it to a voice processing device, which then performs semantic analysis and issues corresponding commands locally or in the cloud, thereby facilitating accurate control of various devices.
[0006] According to a first aspect of the present invention, a voice control system is provided, comprising a voice suite and a server communicating with the voice suite. The voice suite includes: a bone conduction device for collecting a user's voice input based on bone conduction and transmitting the collected voice input to a voice processing device; a voice processing device for receiving the voice input collected by the bone conduction device and uploading the voice input to a server. The server is configured to perform semantic recognition on the voice input sent by the voice processing device to generate and issue an operation command for a target device operation corresponding to the recognized semantics.
[0007] According to a second aspect of the present invention, a voice suite is provided, comprising: a bone conduction device for collecting voice input based on bone conduction and transmitting the collected voice input to a voice processing device; and a voice processing device including a communication unit communicatively coupled to the bone conduction device, the voice processing device receiving voice data collected by the bone conduction device via the communication unit to implement semantic recognition of the voice input and target device operation corresponding to the recognized semantics.
[0008] According to a third aspect of the present invention, a bone conduction device is provided, comprising: a bone conduction sensor for collecting a user's voice input via bone conduction; a bone conduction speaker for transmitting content received from a voice processing device and / or a target device into the user's ear canal via bone conduction; and a communication module for transmitting the collected voice input to the voice processing device, so that the voice processing device can perform semantic recognition of the voice input and target device operation corresponding to the recognized semantics.
[0009] According to a fourth aspect of the present invention, a speech processing device is provided, comprising: a communication unit for receiving speech data collected by a bone conduction device; and a networking unit for uploading the speech data received from a user by the bone conduction device to a server, wherein the server and / or the speech processing device performs semantic recognition on the speech input to generate and issue an operation command for a target device operation corresponding to the recognized semantics.
[0010] According to a fifth aspect of the present invention, a voice control method is proposed, comprising: a bone conduction device collecting voice input; the bone conduction device sending the voice input to a voice processing device; and the voice processing device implementing semantic recognition of the voice input and generating corresponding target device operation commands.
[0011] This invention utilizes the principle of bone conduction of sound waves and uses a bone conduction sensor to effectively solve the problem that the microphone is susceptible to interference when receiving signals transmitted through the air, ensuring that the device can only be awakened by the user. At the same time, it enhances the accuracy of voice command recognition, thereby improving the intelligent voice operation experience of the device user. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The above and other objects, features and advantages of the present disclosure will become more apparent through a more detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings, wherein like reference numerals generally represent like components in the exemplary embodiments of the present disclosure.
[0013] Figure 1 FIG2 shows a schematic diagram of the composition of a bone conduction device according to an embodiment of the present invention.
[0014] Figure 2 An example of wearing a bone conduction device is shown.
[0015] Figure 3 An example of the structure of a bone conduction device according to the present invention is shown.
[0016] Figure 4 FIG. 4 is a schematic diagram showing the composition of a voice suite according to an embodiment of the present invention.
[0017] Figure 5An example of collecting voice input by the voice suite of the present invention is shown.
[0018] Figure 6 A schematic diagram showing the composition of a voice control system according to an embodiment of the present invention is shown.
[0019] Figure 7 A schematic flow chart of a voice control method according to an embodiment of the present invention is shown.
[0020] Figure 8 An example of a voice control processing flow of the present invention is shown.
[0021] Figure 9 The figure shows a working diagram of the intelligent voice wearable device of the present invention. DETAILED DESCRIPTION
[0022] The preferred embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although preferred embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments described herein. Rather, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.
[0023] As mentioned above, existing intelligent voice devices use microphones to receive voice commands, but they often fail to block voice commands from non-users and struggle to accurately recognize them in noisy environments. To address this, a solution has been proposed that uses a bone conduction device as a voice acquisition entry point and transmits it via local communication to a voice processing device. The latter then performs semantic analysis and issues corresponding commands locally or in the cloud, thereby facilitating accurate control of various devices.
[0024] Bone conduction is a method of sound transmission that converts sound into mechanical vibrations of varying frequencies, transmitting them through the skull, bony labyrinth, inner ear lymphatic system, augumentary organ, auditory nerve, and auditory center. Compared to classic sound transmission methods that generate sound waves through a diaphragm, bone conduction eliminates many of the necessary steps, enabling clear sound reproduction in noisy environments. Furthermore, the sound waves are prevented from dispersing through the air and potentially disturbing others.
[0025] Bone conduction technology is divided into bone conduction speaker technology and bone conduction microphone technology. Bone conduction speaker technology is used to hear sound. Air conduction speakers convert electrical signals into sound waves (vibration signals) and transmit them to the auditory nerve. Bone conduction speakers, on the other hand, transmit these converted electrical signals directly through the bones to the auditory nerve. Bone conduction microphone technology is used to collect sound. While air conduction transmits sound waves through the air to the microphone, bone conduction transmits them directly through the bones. A bone conduction microphone is a non-acoustic sensor, also referred to as a bone conduction sensor. When a person speaks, the vibrations of the vocal cords are transmitted to the larynx and skull. The bone conduction sensor collects these vibration signals and converts them into electrical signals to detect speech. Background noise is less likely to affect this type of non-acoustic sensor, so bone conduction audio shields the noise at the source, making it particularly suitable for voice communication in noisy environments.
[0026] Figure 1 FIG. 1 shows a schematic diagram of the composition of a bone conduction device according to an embodiment of the present invention. Figure 1 As shown, the bone conduction device 100 may include a bone conduction sensor 110 and a communication module 120. The bone conduction sensor 110 is used to collect the user's voice input via bone conduction. The communication module 120 is used to transmit the collected voice input to a voice processing device, which then performs semantic recognition of the voice input and performs target device operations corresponding to the recognized semantics. As described below, the target device may be the voice processing device itself, another smart device, or a traditional household appliance.
[0027] The bone conduction device can be combined with the speech processing device to form a speech kit (as shown below Figure 4 As shown in the figure, "kit" refers to a set of devices that work together to achieve a specific function. In the present invention, the bone conduction device needs to be worn directly by the user because it collects vibration signals directly from the larynx and skull. In one embodiment, the bone conduction device can be implemented as a standalone bone conduction headset. Figure 2 An example of wearing a bone conduction device is shown. As shown, the user's vocal cords vibrate in the throat to produce sound. The sound propagates outward through the air along the solid line and internally through the bone along the dotted line. At this point, the bone conduction device uses bone conduction sensor 110 to collect the vibration signal, convert it into an electrical signal, and transmit it to communication module 120. The voice processing device aggregates the voice information collected by the bone conduction device and, through local and / or cloud-based semantic processing functions, implements semantic recognition of the voice input and the corresponding target device operation.
[0028] In different implementations, the bone conduction device and voice processing device in the kit can interact to varying degrees. In one embodiment, the bone conduction device and voice processing device can be arranged within the same physical device. For example, the voice kit can be implemented as a wearable smart device, such as a smart Bluetooth headset, a smart VR / AR helmet, or smart glasses. In this case, the bone conduction device and voice processing device can be implemented as different functional units of the smart device, and signals can be transmitted via a communication bus within the device (which can be regarded as the communication module 120 in a special case). In one embodiment, the bone conduction device and voice processing device can also be a kit connected via wired or wireless communication. For example, a bone conduction headset connected via Bluetooth or other short-range communication and a voice processing device implemented as a smart speaker or smartphone, or a wired bone conduction headset and a body-worn processing box, etc. In another embodiment, the bone conduction device and voice processing device can also be a detachable kit. Depending on the usage scenario, the two devices can be combined into a single device or separated into two independent devices for use when needed.
[0029] In the case of cloud semantic recognition and command issuance, the above-mentioned voice suite can also be combined with the server to form a voice control system (as shown below Figure 5 The server can be a local server that communicates with the voice suite via short-range communication, or a remote server (e.g., a server cluster) that can remotely communicate with the voice processing devices in multiple voice suites to provide cloud-based semantic recognition, command generation, and delivery functions.
[0030] return Figure 1 A bone conduction sensor is a transducer that converts sound into electrical signals. As shown in the figure, bone conduction sensor 110 collects vibration signals emitted from the user's throat and conducted through the bones, converts them into electrical signals containing the user's voice information, and transmits them to communication module 120. Communication module 120 then transmits the electrical signals containing the user information as the user's voice input data to the voice processing device for subsequent semantic recognition of the voice input and target device operation corresponding to the recognized semantics. Thus, by retaining only the most basic semantic collection and communication functions, bone conduction device 100 achieves low power consumption and miniaturization, making it easier for users to wear.
[0031] In order to achieve miniaturization and low power consumption, the communication module 110 can be a low-power short-range communication module. Here, "short-range communication" refers to short-range wireless communication with a communication distance generally within a few hundred meters. In one embodiment, the communication module 110 can be a Bluetooth (BT) communication module that communicates with the voice processing device based on Bluetooth technology, for example, a communication module based on the Bluetooth Mesh solution. In another embodiment, the communication module 110 can be an infrared (IR) communication module that communicates with the voice processing device based on infrared technology, for example, a high-speed infrared transmission module. In one embodiment, the communication module 110 can be a Zigbee communication module that communicates with the voice processing device based on Zigbee technology. In other embodiments, the communication module 110 can also use a combination of BT and IR. It should be understood that the communication module 110 of the present invention can also be implemented by, for example, new low-power short-range communication technologies developed in the future, so as to create conditions for its convenient wearing through the miniaturization and low power consumption of the device 100 itself.
[0032] In other embodiments, the bone conduction device 100 may also include a WiFi communication module, which consumes relatively high power and generally requires more processing power, to communicate with the voice processing device via a local area network. Of course, the WiFi communication module may also be used for short-range communication in some embodiments.
[0033] Typically, the bone conduction device 100 of the present invention also includes a bone conduction speaker for voice outputting the content received by the communication module from the voice processing device. The inclusion of the speaker enables further voice interaction with the user. The content voice output by the speaker includes at least one of the following: a statement of the execution command; and interactive content, such as interaction with the user, to further capture missing semantic elements. Voice interaction will be described in detail below with reference to the description of the voice control system. In a broader application, the bone conduction device 100 can be implemented as a bone conduction headset, at least a device with a bone conduction sound playback function. In this case, the bone conduction speaker can also be used to output the results of the command execution. For example, when a user listens to music through bone conduction headphones, they can control the playback through voice, and the playback will proceed accordingly. For example, upon receiving a voice command "skip this song," the next song will be played as the result of the command execution.
[0034] In one embodiment, device 100 further includes a power supply module, including but not limited to: a wireless charging component; a battery pack; and a USB port. Because device 100's voice collection and transmission functions require minimal energy, and the energy required for everyday use, such as music playback, is also relatively low, device 100 consumes relatively little power, making it suitable for a power supply structure that does not require a power cord. This significantly enhances the portability and flexibility of device 100.
[0035] To reduce power consumption and avoid misoperation, the bone conduction device 100 of the present invention preferably also includes a remote wakeup function. Here, "remote wakeup" refers to the ability to wake up a voice device using a specific voice wakeup word. For example, the commercially available Tmall Genie can be woken up using the wakeup word "Tmall Genie." Specifically, the device 100 may also include a wakeup module for identifying the wakeup word from the user's voice input. After the wakeup module identifies the wakeup word, the communication module 120 can then transmit the collected voice input to the voice processing device. For example, if the wakeup word is used solely for wakeup and does not include other commands, the communication module 120 can receive and transmit the user's voice input after the wakeup word. In other words, the intelligent voice interaction function of the bone conduction device 100 is activated by the wakeup word, and only then does voice input transmission to the voice processing device begin. If the wakeup word also includes a command, the communication module 120 can receive and transmit the user's wakeup word itself and the subsequent voice input. Because the wake-up module can be implemented using a limited, miniaturized, low-power DSP (digital signal processing) circuit, the addition of far-field wake-up functionality does not substantially impact the miniaturization and low-power characteristics of device 100. If bone conduction device 100 does not include a wake-up module, the voice input transmission function from the bone conduction device to the voice processing device can be always enabled, and the voice processing device can use its equipped wake-up module to recognize the corresponding wake-up word and enable the voice interaction function.
[0036] Figure 3 FIG. 1 shows an example of the composition of a bone conduction device according to the present invention. Figure 3 As shown, it is implemented as Figure 2 The bone conduction device of the bone conduction earphone 300 may include a bone conduction sensor 310, a communication module implemented as a Bluetooth and / or infrared short-range communication module 320, a battery 330, and a bone conduction speaker 340. The smart voice sticker 300 may also have an attachment structure suitable for attachment to any suitable attachment surface (such as Figure 3 In other embodiments, the communication module 320 may also include a Zigbee communication module.
[0037] Specifically, the bone conduction sensor 310 converts the vibration caused by the received user voice into an electrical signal, and sends the above electrical signal carrying the user voice information to the BT / IR module 320. The BT / IR module 320 sends the above user's voice input data to the voice processing device, so as to use the voice processing device and the cloud to realize semantic recognition and generate corresponding operation commands. Subsequently, the BT / IR module 320 can also obtain the content data that the cloud expects the bone conduction headset 300 to output to the user through the voice processing device, such as data for further interaction with the user or reporting operation results, or call or music data in call or music playback scenarios. The BT / IR module 320 can send the above electrical signal including cloud content information to the bone conduction speaker 340, and the latter converts the electrical signal into bone vibration that can understand speech, so that the user can hear the above content.
[0038] In different embodiments, TTS (speech synthesis) can be implemented by different entities. For example, the cloud can directly send data via TTS, or the voice processing device or the bone conduction headset 300 can include the above-mentioned TTS module. In one embodiment, for the sake of transmission efficiency, low power consumption and miniaturization, it is preferred that the voice processing device performs speech synthesis based on the content sent from the cloud, and then transmits the signal containing the above-mentioned speech synthesis to the bone conduction headset 300, and the BT / IR module 320 transmits the above-mentioned information in the form of an electrical signal for the speaker to directly perform electrical vibration conversion. It should be understood that in other embodiments, Figure 3 The bone conduction device shown can also be implemented as other devices that include bone conduction and information transmission and reception functions, such as a smart helmet.
[0039] In addition, the bone conduction device may also include other sensor devices (described in detail below) for collecting scene or action information, and the bone conduction device turns on or off the voice input collection function based on the collected scene or action information.
[0040] As mentioned above, the bone conduction device of the present invention can be combined with a voice processing device to obtain a voice suite for realizing the voice collection and networking functions required for local operation. Figure 4 FIG. 1 shows a schematic diagram of the composition of a voice suite according to an embodiment of the present invention. Figure 4 As shown, the voice suite 400 may include the above combined Figure 1-3 The bone conduction device 410 and the speech processing device 420 are described.
[0041] The speech processing device 420 includes a communication unit 421 that is communicatively connected to the bone conduction device 410. The speech processing device 420 receives speech data collected by the bone conduction device via its communication unit 421 to perform semantic recognition of the speech input and target device operations corresponding to the recognized semantics. The communication unit 421 can be adapted to communicate with the communication module 411 of the small bone conduction device 410 using a communication mode corresponding to that used by the communication module, such as low-power, short-range communication. In one embodiment, the communication unit 421 is adapted to communicate with the corresponding communication module 411 using Bluetooth, Zigbee, and / or infrared technology.
[0042] Specifically, the aforementioned semantic recognition and generation and issuance of operational commands can be implemented in the cloud. To this end, the voice processing device 420 may include a networking unit 422 for uploading the user's voice data received from the bone conduction device 410 to a cloud-based server. The networking unit 422 is, for example, a module that accesses the internet using WiFi and / or mobile communication technologies such as 4G and 5G. The server can perform semantic recognition of the voice input to generate and issue operational commands for target device operations corresponding to the recognized semantics.
[0043] In some embodiments, the generation and issuance of the above-mentioned semantic recognition and operation commands can be implemented locally. Therefore, the server may include: a local server in close communication with the voice suite, the local server being used to perform semantic recognition on at least part of the voice input, generate operation commands for target device operations corresponding to the recognition semantics, and issue the operation commands. For example, the local server may be a smart speaker used as a home intelligent processing terminal. Thus, the processing speed of voice commands is improved by the local server. In one embodiment, the above-mentioned local server can be connected to a cloud server, and the whole constitutes the "server" in the present invention.
[0044] Based on different application scenarios, the kit may include other devices in addition to the bone conduction device 410 and the voice processing device 420, such as multiple smart voice stickers arranged in different areas, each of which is communicatively connected to one or more of the miniaturized devices within communication range. For example, in a home scenario, different voice stickers can be placed in different rooms (e.g., the living room, bedroom, bathroom, and kitchen). These multiple voice stickers 410 can be connected to a voice processing device 420 within the Bluetooth communication range. The only voice processing device 420 in the kit can then connect to the cloud, thereby achieving more comprehensive kit control.
[0045] According to different control scenarios, the control of the target device by the user's voice input can be achieved based on different approaches. For example, in different embodiments, the target device can directly receive the operation command issued from the server and perform the operation corresponding to the operation command; and / or the voice processing device 420 receives the operation command issued by the server via its networking unit 422, and issues the operation command to the target device. The target device that directly receives the operation command issued from the server and executes it may include a smart home appliance that is connected to the network. The target device that obtains the operation command via the voice processing device 520 may include a smart home appliance, and may also include a traditional home appliance, for example, via a device in the kit that has an infrared control function for controlling traditional home appliances.
[0046] For example, all the smart appliances running in the home are connected to the control server. At this time, the operation commands for the smart appliances (for example, lowering the temperature of the refrigerator compartment) can be directly issued by the control server. For traditional appliances that need to be controlled using corresponding infrared codes, the server can generate corresponding operation commands (for example, turning off the air conditioner) based on semantic recognition, look up the infrared operation code of the air conditioner, and send the above command directly to a voice sticker device used as a universal infrared remote control. In other embodiments, the above infrared operation code can also be implemented locally, for example, at a voice processing device.
[0047] If the bone conduction device 410 includes a speaker for voice output, the voice processing device 420 may also use its networking unit 422 to receive user interaction content from the server, and the communication unit 421 is further configured to transmit the user interaction content to the bone conduction device for voice output. The interaction content may include confirmation of device operation (e.g., "The light is on"), acquisition of required semantic elements (e.g., upon recognizing the user's voice input of "turn on the light" and if more than one light is within range, further inquiring "Which light should I turn on"), or a combination of the two (e.g., "The TV is on, which channel should I watch?").
[0048] In one embodiment, the speech processing device 400 itself can have simple speech recognition and command generation and issuance capabilities. To this end, the speech processing device 420 can include: a speech recognition unit for performing semantic recognition on the speech input; and an operation command generation unit for generating an operation command for the target device operation corresponding to the recognized semantics. This enables the kit of the present invention to not only understand complex semantics by connecting to a cloud server, but also quickly respond to simple input.
[0049] In one embodiment, the voice processing device 420 itself can be a smart speaker connected to a cloud server, or other device that also has voice collection capabilities. Here, the bone conduction device 410 can be used as an intelligent assistant to help the smart speaker collect voice in, for example, noisy environments.
[0050] In certain embodiments, the voice interaction function of the voice suite 400 can be activated based on the wake-up module's recognition of a wake-up word. In one embodiment, the bone conduction device 410 includes a wake-up module that transmits the collected voice input to the voice processing device 420 after the wake-up module recognizes the wake-up word from the voice input. In other words, the bone conduction device 410 can only begin transmitting the voice collected via bone conduction to the voice processing device 420 after the voice interaction function is activated, rather than continuously collecting and transmitting the user's voice, thereby avoiding unnecessary power consumption of the bone conduction device. Alternatively or in addition, the voice processing device 420 includes a wake-up module that uploads the received voice input to the server after the wake-up module recognizes the wake-up word from the voice input. In other words, the voice processing device 420 can only begin transmitting the voice collected via bone conduction to the server after the voice interaction function is activated, rather than continuously collecting and transmitting the user's voice, thereby avoiding unnecessary power consumption.
[0051] Typically, the voice processing device 420 itself may also include a voice collection device, such as a built-in microphone or a wirelessly connected voice tag. Figure 5 An example of collecting voice input by the voice suite of the present invention is shown.
[0052] and Figure 4 similar, Figure 5 The illustrated voice suite 500 may include a bone conduction device 510 and a voice processing device 520. Frame 510 shows the collection and transmission of vibrations conducted by the bone conduction device 510, as indicated by the dashed line. A communication module 511 is in communication with a communication unit 521 of the voice processing device 520. The voice processing device 520 receives voice data collected by the bone conduction device via its communication unit 521 to perform semantic recognition of the voice input and target device operations corresponding to the recognized semantics. Similarly, the voice processing device 520 may include a networking unit 522 for uploading voice data received from the user via the bone conduction device 510 to a cloud-based server for semantic recognition and the generation and issuance of operational commands.
[0053] Different from Figure 4 , Figure 5The illustrated voice processing device 520 also includes a microphone (MIC) 523. MIC 523 can be used to collect signals transmitted through the air by the user's voice as a second voice input. Voice processing device 520 can upload the second voice input to a server via its networking unit 522. The server can then perform semantic recognition on the second voice input to generate and issue a second operation command for a target device operation corresponding to the recognized semantics.
[0054] In other words, the speech processing device also has a speech collecting device (for example, Figure 5 In the case of MIC 523 shown in FIG5 , bone conduction data can be collected simultaneously through bone conduction by bone conduction device 510 and air conduction by the microphone. Therefore, by comparing the voice input collected through bone conduction with the second voice input collected through air conduction, information can be obtained from multiple levels, facilitating more accurate identification of user intent and providing more appropriate feedback. In certain embodiments, the server can generate and issue the operation command and / or the second operation command based on the comparison of the voice input and the second voice input.
[0055] Based on the comparison between the first voice input and the second voice input, current environment information can be generated, a multi-person interaction scenario can be determined, and / or irrelevant information can be filtered out of the second voice input. The above processing can be performed locally by the voice processing device 520 or uploaded to the cloud for execution, and the server can generate the second operation command based on the above processing.
[0056] Because bone conduction voice input is more accurate and can avoid erroneous voice capture from users other than the intended user, the voice processing device 520 can activate voice control based solely on the wake-up word recognized from the bone conduction voice input. In other words, the wake-up word captured by the MIC 523 must match the wake-up word received by the communication unit 521 to activate voice control.
[0057] As an alternative and supplement, in the subsequent voice interaction process, if it is found through comparison that the second voice input collected by air conduction includes the voice input waveform collected by bone conduction and other waveforms, it can be judged that the current environmental information is noisy, and thereby, for example, the confidence of the second voice input can be reduced.
[0058] In addition, if the comparison finds that the second voice input collected by air conduction includes the voice input waveform collected by bone conduction and the input waveform from other users, it is possible to determine the scenario in which multiple people participate in the voice interaction, and, for example, start a script to deal with multi-person interaction in the background, so as to facilitate giving more accurate feedback.
[0059] Furthermore, irrelevant information in the second voice input, such as background music, television or chatting sounds, can be discovered through comparison, and the second voice input can be filtered out of irrelevant information to facilitate the generation of the second operation command based on the processed second voice input.
[0060] Furthermore, the above kit can be combined with a server to implement a voice control system. Figure 6 FIG. 1 shows a schematic diagram of the composition of a voice control system according to an embodiment of the present invention. Figure 6 As shown, the system 600 may include multiple voice suites 610 and a server 620 as described above. Here, the server 620 may refer to a server group that provides a specific function, for example, a server group that provides a cloud voice interaction service.
[0061] At least part of the voice suite 610 may include a bone conduction device and a voice processing device. Other voice suites may include a voice processing device and other voice interaction devices, such as voice stickers. The bone conduction device and voice stickers, which serve as voice collection portals, may communicate with the voice processing device via short-range, low-power communication means (e.g., BT as shown).
[0062] The voice suite 610 can be connected to the server 620 via the networking function of the voice processing device (e.g., a WiFi module). The server 620 can perform semantic recognition on the voice input uploaded by the voice processing device to generate and issue an operation command for the target device operation corresponding to the recognized semantics.
[0063] In one embodiment, the server itself can perform all operations such as semantic recognition, operation command generation and issuance. Therefore, the server 620 may include: a semantic processing server for performing semantic recognition on the uploaded voice input; a command generation server for generating operation commands for target recognition operations based on the recognized semantics; and a command issuance server for issuing operation commands.
[0064] In another embodiment, the server 620 can be used only for semantic recognition, or for generating and issuing operation commands for some target devices. At least the control of the target device of the present invention can be achieved by an external server. This is especially applicable to situations where a service provider of a certain brand provides remote control functions for its own smart devices. Therefore, the server 620 may include: a semantic processing server for performing semantic recognition on the uploaded voice input, and the server sends the recognized semantics to an external server, wherein the external server generates a command generation server for an operation command for a target recognition operation based on the recognized semantics, and a command issuing server for issuing the operation command.
[0065] In one embodiment, the server 620 can pre-acquire at least one of the following local device configuration information. This local device configuration information may include: distribution and device information for the bone conduction device, the voice processing device, and / or at least a portion of the target device itself; and the corresponding relationship between at least two of the bone conduction device, the voice processing device, and the target device. Thus, the server 620 can also automatically complete missing semantic elements in the recognition semantics for executing operations on the target device based on the local device configuration information.
[0066] Since voice input collection requires the device to monitor sound information, such as keeping the sound collection device (e.g., bone conduction sensor) and remote wake-up module active, as well as the modules for local analysis or remote upload, this function consumes considerable power. Therefore, to reasonably control power consumption, the bone conduction device, or even the voice processing device, can be enabled only in specific scenarios or actions to collect voice input.
[0067] In one embodiment, the bone conduction device can enable or disable a voice input collection function based on scene information. Alternatively or in addition, the voice processing device can enable or disable a second voice input collection function based on scene information. In different implementation scenarios, the voice collection functions of the bone conduction device and the voice processing device can be enabled or disabled independently or in conjunction with each other, and the present invention is not limited thereto. Here, the scene information includes at least one of the following: scene information determined based on signals collected by sensors on the bone conduction device; scene information determined based on signals collected by sensors on the voice processing device; scene information determined based on an associated function on the voice processing device; and scene information determined based on a comparison of the voice input and the second voice input.
[0068] Specifically, the bone conduction sensor configured on the bone conduction device can determine that the user has started wearing the bone conduction sensor and speaking (i.e., obtains wearing scenario information) when it first receives voice vibrations, and accordingly activates its own voice input collection function. In other embodiments, the bone conduction device can also be provided with other sensors, such as a motion sensor (e.g., an accelerometer), a temperature sensor, or an infrared sensor, and determine the scenario (e.g., the wearing scenario) based on the signals collected by the above sensors, and activate the voice input collection function accordingly. For example, the accelerometer can determine the wearing scenario and activate the voice input collection function by identifying the wearing action, the temperature sensor can identify the human body temperature, and the infrared sensor can determine the wearing scenario and activate the voice input collection function by identifying the ear.
[0069] Additionally, a voice processing device can also be used to obtain scene information. For example, a voice processing device can determine scene information based on its own microphone or other sensors, and accordingly determine whether to enable or disable its own or the bone conduction device's voice collection function. Furthermore, because voice processing devices (e.g., smartphones or smart speakers) have greater processing power and more functionality, it is preferable to be able to obtain scene information through methods other than sensing, such as scene information determined by a connected function. For example, if a user is using the bus query and arrival reminder functions of a map app installed on their smartphone, the smartphone acting as a voice processing device can determine that the user is in a noisy environment while on public transportation. In this case, the voice collection function of the bone conduction device can be activated separately, which is more likely to accurately collect voice in noisy environments. Furthermore, if the user activates the running GPS recording function on their smartphone, the voice collection function of the bone conduction device can also be activated separately, which is more likely to accurately collect voice while running.
[0070] Furthermore, when both the bone conduction device and the voice processing device's voice input collection functions are enabled, the scene information can be determined based on a comparison of the collected voice input and the second voice input, and the voice input collection function of one or both devices can be disabled accordingly. For example, if the voice input and the second voice input are almost identical, it can be determined that the voice processing device's microphone is capable of collecting the voice input well. In this case, the bone conduction device's collection function can be disabled to avoid unnecessary power consumption.
[0071] Furthermore, the voice input collection function of the bone conduction device can be activated or deactivated based on an operation on the device itself, such as when the device is worn, or based on a specific movement of the device. For example, the voice collection function can be activated or deactivated based on a head movement (e.g., shaking the device side to side) or a hand movement (e.g., tapping a specific location on the device).
[0072] The aforementioned scene- or action-based activation can be combined with remote control of the target device. For example, in a noisy or sporty scene, the bone conduction device's voice input collection function can be activated, and the voice input can be uploaded to the voice processing device and processed by the server, enabling remote control of the target device.
[0073] The present invention can also be implemented as a voice processing device, comprising: a communication unit for receiving voice data collected by the bone conduction device; and a networking unit for uploading the voice data from the user received from the bone conduction device to a server, wherein the server and / or the voice processing device performs semantic recognition on the voice input to generate and issue an operation command for a target device operation corresponding to the recognized semantics.
[0074] Furthermore, the voice processing device may include: a voice collection device for collecting a second voice input to implement semantic recognition of the second voice input and a second target device operation corresponding to the recognized semantics. The voice processing device may initiate a voice control operation based on a wake-up word recognized from the voice input collected by bone conduction.
[0075] In addition, the present invention can also be implemented as a voice control method. Figure 7 A schematic flow chart of a voice control method according to an embodiment of the present invention is shown. The method can be implemented by the bone conduction device, kit, and system described above.
[0076] In step S710, a bone conduction device (e.g., a bone conduction headset) collects voice input. In step S720, the bone conduction device transmits the voice input to a voice processing device. In one embodiment, this transmission may be via short-range communication, such as infrared, Bluetooth, and / or Zigbee.
[0077] In step S730, the speech processing device performs semantic recognition of the speech input and generates a corresponding target device operation command. In different embodiments, the operations for semantic recognition and command generation can be performed by different objects, for example, the speech processing device itself, a local server, a remote server, or any combination thereof.
[0078] Therefore, in one embodiment, in step S730, the speech processing device may send the speech input to the server, and use the server to perform semantic recognition of the speech input to obtain an operation command for the target device operation corresponding to the recognized semantics. In another embodiment, in step S730, the server may include a local server, which may perform semantic recognition on at least a portion of the speech input, generate an operation command for the target device operation corresponding to the recognized semantics, and issue the operation command.
[0079] If the voice processing device includes a voice acquisition device, the method may further include the voice processing device acquiring a second voice input; the voice processing device performing semantic recognition of the second voice input and generating a second operation command corresponding to the target device. Furthermore, a more appropriate operation command may be generated by comprehensively considering the second voice input acquired by air conduction and the voice input acquired by bone conduction.
[0080] Figure 8An example of a voice control processing flow for the present invention is shown. As shown, during the local collection and upload phase, the bone conduction device in the voice suite (and, in some embodiments, the voice collection module in the voice processing device) monitors voice command input from the user or smart device. The voice processing device performs preliminary processing on the collected voice, such as ASR (Analog Speech Recognition) and transmits the captured voice commands to the cloud. During the cloud processing phase, the server performs subsequent processing on the captured voice commands, such as NIP (Natural Voice Processing) and NIU (Natural Voice Understanding), and based on the processing results, performs command parsing and text-to-text (TTS) output. During the local processing phase, the parsed commands can be transmitted directly to the target device for execution (e.g., the smart device directly executes the parsed commands from the cloud), or the target device can execute the commands after conversion by the voice processing device or bone conduction device (e.g., infrared commands issued by the bone conduction device to traditional home appliances). Additionally, if audio output is available, the bone conduction device can output voice through its built-in bone conduction speaker.
[0081] In a specific application scenario, the voice kit of the present invention can be implemented as an intelligent voice wearable device. In this case, the intelligent voice wearable device combines the functions of both the bone conduction device and the voice processing device, and can serve as the target device for executing instructions. Figure 9 The following figure shows the operation of the intelligent voice wearable device of the present invention. As shown in the figure, a user wearing a bone conduction sensor can issue voice commands. Here, the voice commands can include a wake-up word and an operation instruction. The voice signal is transmitted to the device through the bone conduction sensor to wake the device. The device receives the voice signal received by the bone conduction sensor and microphone and uploads it to the cloud engine for recognition. The cloud engine converts the voice recognition results into device control instructions and returns them to the device, which then executes the instructions.
[0082] Bone conduction is not affected by background noise or wind noise, and does not receive any voices from people other than the device operator. Air conduction is easily affected by background noise and wind noise, and receives the voices of people other than the device operator indiscriminately. Since the device is only awakened when the bone conduction sensor transmits a wake-up word signal (that is, when the device user speaks the wake-up word) to the device, the microphone signal alone cannot wake the device, thus preventing non-device users from waking up the device. During the operation command acquisition phase, bone conduction and air conduction acquisition are performed simultaneously. This ensures that when the microphone is affected by environmental noise and the voice quality is significantly degraded, the bone conduction sensor can enhance the voice signal quality, thereby ensuring the accuracy of voice command recognition. The bone conduction function can also be applied to other similar scenarios where voice quality needs to be enhanced, such as the uplink voice call quality when making a call with headphones.
[0083] The bone conduction device, voice control kit, and voice control system and method according to the present invention have been described in detail above with reference to the accompanying drawings. This solution utilizes the principle of bone conduction of sound waves, using a bone conduction sensor to effectively address the issue of microphones being susceptible to interference when receiving signals transmitted through air. This ensures that the device can only be activated by the user, while also enhancing the accuracy of voice command recognition, thereby improving the user's experience with intelligent voice control.
[0084] In addition, the method according to the present invention may also be implemented as a computer program or a computer program product, which includes computer program code instructions for executing the above steps defined in the above method of the present invention.
[0085] Alternatively, the present invention can also be implemented as a non-transitory machine-readable storage medium (or computer-readable storage medium, or machine-readable storage medium) on which executable code (or computer program, or computer instruction code) is stored. When the executable code (or computer program, or computer instruction code) is executed by a processor of an electronic device (or computing device, server, etc.), the processor executes the various steps of the above-mentioned method according to the present invention.
[0086] Those skilled in the art will further appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the disclosure herein may be implemented as electronic hardware, computer software, or combinations of both.
[0087] The flowcharts and block diagrams in the accompanying drawings show the possible implementation architecture, functions and operations of the systems and methods according to multiple embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of code, and the part of the module, program segment or code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0088] While various embodiments of the present invention have been described above, the foregoing description is intended to be illustrative, non-exhaustive, and not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A voice control system, comprising a voice suite and a server communicating with the voice suite, wherein: The voice suite includes: a bone conduction device that collects a user's voice input based on bone conduction and sends the collected voice input to a voice processing device; A voice processing device is used to receive the voice input collected by the bone conduction device and upload the voice input to the server. The server is used to perform semantic recognition on the voice input sent by the voice processing device to generate and issue an operation command for the target device operation corresponding to the recognized semantics. The voice processing device includes a voice collection device for collecting the user's voice signal transmitted through the air as the second voice input, and the voice processing device uploads the second voice input to the server. The server generates and issues the operation command and / or the second operation command based on a comparison between the voice input and the second voice input. When the voice input and the second voice input are almost the same, the collection function of the bone conduction device is turned off, and the voice collection device is used to collect the voice input. The generating and issuing of the operation command and / or the second operation command based on the comparison of the voice input and the second voice input includes: generating current environment information based on a comparison between the voice input and the second voice input, and generating the operation command and / or the second operation command based on the current environment information; Based on the comparison between the voice input and the second voice input, determining a multi-person interaction scenario, and generating the operation command and / or the second operation command based on the multi-person interaction scenario; and Based on a comparison between the voice input and the second voice input, filtering out irrelevant information from the second voice input, and generating the second operation command based on the processed second voice input; The generating of the current environment information is specifically as follows: In response to the second voice input acquired by air conduction including the voice input waveform acquired by bone conduction and other waveforms, it is determined that the current environment information is noisy; The determination of the multi-person interaction scenario is specifically as follows: In response to the second voice input acquired by air conduction including the voice input waveform acquired by bone conduction and the input waveform from other users, it is determined that multiple people are participating in the voice interaction; The processing of filtering out irrelevant information from the second voice input is specifically as follows: Determine and filter out irrelevant information in the second voice input through comparison; The server is further configured to perform semantic recognition on the second voice input to generate and issue a second operation command for the target device operation corresponding to the recognized semantics.
2. The system of claim 1, wherein: The bone conduction device includes a wake-up module, which sends the collected voice input to the voice processing device after the wake-up module recognizes the wake-up word from the voice input; And / or the voice processing device includes a wake-up module, and after the wake-up module recognizes the wake-up word from the voice input, the received voice input is uploaded to the server.
3. The system of claim 1, wherein: The speech processing device is used for: A voice control operation is initiated based on the wake-up word recognized from the voice input.
4. The system of claim 1, wherein: The bone conduction device is used for: Based on the scene information, turn on or off the voice input collection function, and / or The speech processing device is used for: Enable or disable the second voice input collection function based on the scene information.
5. The system of claim 4, wherein: The scene information includes at least one of the following: Scene information determined based on signals collected by sensors on the bone conduction device; Scene information determined based on signals collected by sensors on the speech processing device; Scenario information determined based on an associated function on the speech processing device; as well as Scene information determined based on a comparison of the voice input and the second voice input.
6. The system of claim 1, wherein: The voice input collection function of the bone conduction device is activated based on at least one of the following: The operation of wearing the bone conduction device; and Specific actions for the bone conduction device.
7. The system of claim 1, wherein: The speech processing device is used to perform semantic recognition on at least part of the speech input and generate an operation command for a target device operation corresponding to the recognized semantics.
8. The system of claim 1, wherein: The bone conduction device comprises: A bone conduction sensor for acquiring a user's voice input via bone conduction; The bone conduction speaker is used to transmit the content received from the voice processing device and / or the target device into the user's ear canal via bone conduction.
9. The system of claim 8, wherein: The content output by the bone conduction speaker includes at least one of the following: a statement of the execution order; The result of executing the command; and Interaction content with users.
10. The system of claim 1, wherein: The target device directly receives the operation command issued by the server and performs the operation corresponding to the operation command; and / or The voice processing device receives the operation command sent by the server, and sends the operation command to the target device by itself or via the bone conduction device.
11. The system of claim 1, wherein: The target device includes at least one of the following: The speech processing device itself; A networked smart home appliance that receives and executes the operation command issued by the server; and A traditional household appliance that obtains operation commands via the voice processing device.
12. The system of claim 1, wherein: The server includes: A local server in close communication with the voice suite, the local server being configured to perform semantic recognition on at least a portion of the voice input, generate an operation command for a target device operation corresponding to the recognized semantics, and issue the operation command.
13. A voice kit comprising: a bone conduction device that collects voice input based on bone conduction and sends the collected voice input to the voice processing device; A speech processing device includes a communication unit communicatively connected to the bone conduction device, wherein the speech processing device receives speech data collected by the bone conduction device via the communication unit to realize semantic recognition of the speech input and generate an operation command for a target device operation corresponding to the recognized semantics. The speech processing device includes a speech collection device for collecting a user's speech signal transmitted through the air as a second speech input, and the speech processing device performs the target device operation based on a comparison between the speech input and the second speech input. When the speech input and the second speech input are almost identical, the collection function of the bone conduction device is turned off, and the speech collection device is used to collect the speech input. The implementing the target device operation based on the comparison between the voice input and the second voice input includes: generating an operation command and / or a second operation command based on a comparison of the voice input and the second voice input; The step of generating an operation command and / or a second operation command based on the comparison between the voice input and the second voice input includes: generating current environment information based on a comparison between the voice input and the second voice input, and generating the operation command and / or the second operation command based on the current environment information; Based on the comparison between the voice input and the second voice input, determining a multi-person interaction scenario, and generating the operation command and / or the second operation command based on the multi-person interaction scenario; and Based on a comparison between the voice input and the second voice input, filtering out irrelevant information from the second voice input, and generating the second operation command based on the processed second voice input; The generating of the current environment information is specifically as follows: In response to the second voice input acquired by air conduction including the voice input waveform acquired by bone conduction and other waveforms, it is determined that the current environment information is noisy; The determination of the multi-person interaction scenario is specifically as follows: In response to the second voice input acquired by air conduction including the voice input waveform acquired by bone conduction and the input waveform from other users, it is determined that multiple people are participating in the voice interaction; The processing of filtering out irrelevant information from the second voice input is specifically as follows: Determine and filter out irrelevant information in the second voice input through comparison; The server is used to perform semantic recognition on the second voice input to generate and issue a second operation command for the target device operation corresponding to the recognized semantics.
14. The kit of claim 13, wherein: The bone conduction device includes a wake-up module, which sends the collected voice input to the voice processing device after the wake-up module recognizes the wake-up word from the voice input; and / or The voice processing device includes a wake-up module, which uploads the received voice input to a server after the wake-up module recognizes a wake-up word from the voice input.
15. The kit of claim 13, wherein: The bone conduction device and the voice processing device each include a low-power short-range communication module for close-range communication with each other, and the communication module includes at least one of the following: A Bluetooth communication module that communicates with the voice processing device based on Bluetooth technology; An infrared communication module that communicates with the voice processing device based on infrared technology; as well as A Zigbee communication module that communicates with the voice processing device based on Zigbee technology.
16. The kit of claim 13, wherein: The speech processing device further includes: A networking unit is used to upload the voice data from the user received from the bone conduction device to a server, wherein the server performs semantic recognition on the voice input to generate and issue an operation command for the target device operation corresponding to the recognized semantics.
17. The kit of claim 16, wherein: The networking unit is further configured to: Receive the interactive content sent by the server, and The communication unit is further configured to: The interactive content is sent to the bone conduction device for voice output.
18. The kit of claim 17, wherein: The bone conduction device comprises: A bone conduction sensor for acquiring a user's voice input via bone conduction; The bone conduction speaker is used to transmit the content received from the speech processing device and / or the target device into the user's ear canal via bone conduction.
19. The kit of claim 13, wherein: Based on the scene information, it is determined whether to turn on or off the voice input collection function of the bone conduction device and the voice processing device.
20. The kit of claim 13, wherein: The bone conduction device and the speech processing device are a set connected via wired or wireless connection; or The bone conduction device and the speech processing device are a set arranged in the same physical device; or The bone conduction device and the speech processing device are a detachable kit.
21. A bone conduction device, comprising: A bone conduction sensor for collecting user voice input via bone conduction; a bone conduction speaker, configured to transmit content received from the speech processing device and / or the target device into the user's ear canal via bone conduction; as well as a communication module configured to transmit the collected voice input to a voice processing device, so that the voice processing device can perform semantic recognition of the voice input and generate an operation command for a target device operation corresponding to the recognized semantics. The voice processing device includes a voice acquisition device configured to acquire an airborne voice signal of the user as a second voice input, and to perform the target device operation based on a comparison between the voice input and the second voice input. When the voice input and the second voice input are substantially identical, the bone conduction device's acquisition function is disabled, and the voice acquisition device is used to acquire the voice input. The implementing the target device operation based on the comparison between the voice input and the second voice input includes: Based on the comparison between the voice input and the second voice input, generating and issuing an operation command and / or a second operation command; The generating and issuing of an operation command and / or a second operation command based on the comparison of the voice input and the second voice input includes: generating current environment information based on a comparison between the voice input and the second voice input, and generating the operation command and / or the second operation command based on the current environment information; Based on the comparison between the voice input and the second voice input, determining a multi-person interaction scenario, and generating the operation command and / or the second operation command based on the multi-person interaction scenario; and Based on a comparison between the voice input and the second voice input, filtering out irrelevant information from the second voice input, and generating the second operation command based on the processed second voice input; The generating of the current environment information is specifically as follows: In response to the second voice input acquired by air conduction including the voice input waveform acquired by bone conduction and other waveforms, it is determined that the current environment information is noisy; The determination of the multi-person interaction scenario is specifically as follows: In response to the second voice input acquired by air conduction including the voice input waveform acquired by bone conduction and the input waveform from other users, it is determined that multiple people are participating in the voice interaction; The processing of filtering out irrelevant information from the second voice input is specifically as follows: Determine and filter out irrelevant information in the second voice input through comparison; The server is used to perform semantic recognition on the second voice input to generate and issue a second operation command for the target device operation corresponding to the recognized semantics.
22. The bone conduction device of claim 21, further comprising: The wake-up module is used to recognize the wake-up word from the voice input, and The communication module is used to send the collected voice input to the voice processing device after the wake-up module recognizes the wake-up word.
23. The bone conduction device according to claim 21, wherein: The content output by the bone conduction speaker includes at least one of the following: a statement of the execution order; The result of executing the command; and Interaction content with users.
24. The bone conduction device of claim 21, further comprising: A power supply module, wherein the power supply module includes at least one of the following: Wireless charging components; Battery components; USB port.
25. The bone conduction device of claim 21, further comprising: A sensor device for collecting scene or action information, and The bone conduction device turns on or off a voice input collection function based on the collected scene or action information.
26. A speech processing device, comprising: a communication unit, configured to receive voice data collected by the bone conduction device; a networking unit, configured to upload voice data received from the user through the bone conduction device to a server, wherein the server and / or the voice processing device performs semantic recognition of the voice input to generate and issue an operation command for a target device operation corresponding to the recognized semantics; and a voice collection device for collecting the user's voice signal transmitted through the air as a second voice input, and generating and issuing the operation command and / or the second operation command based on a comparison between the voice input and the second voice input, and when the voice input and the second voice input are almost identical, disabling the collection function of the bone conduction device and using the voice collection device to collect the voice input; The generating and issuing of the operation command and / or the second operation command based on the comparison of the voice input and the second voice input includes: generating current environment information based on a comparison between the voice input and the second voice input, and generating the operation command and / or the second operation command based on the current environment information; Based on the comparison between the voice input and the second voice input, determining a multi-person interaction scenario, and generating the operation command and / or the second operation command based on the multi-person interaction scenario; and Based on a comparison between the voice input and the second voice input, filtering out irrelevant information from the second voice input, and generating the second operation command based on the processed second voice input; The generating of the current environment information is specifically as follows: In response to the second voice input acquired by air conduction including the voice input waveform acquired by bone conduction and other waveforms, it is determined that the current environment information is noisy; The determination of the multi-person interaction scenario is specifically as follows: In response to the second voice input acquired by air conduction including the voice input waveform acquired by bone conduction and the input waveform from other users, it is determined that multiple people are participating in the voice interaction; The processing of filtering out irrelevant information from the second voice input is specifically as follows: Determine and filter out irrelevant information in the second voice input through comparison; The server is further configured to perform semantic recognition on the second voice input to generate and issue a second operation command for the target device operation corresponding to the recognized semantics.
27. The speech processing device according to claim 26, wherein: The voice processing device starts a voice control operation based on a wake-up word recognized from the voice input.
28. A voice control method, comprising: A bone conduction device collects voice input; The bone conduction device sends the voice input to a voice processing device; The speech processing device realizes semantic recognition of the speech input and generates corresponding target device operation commands. Wherein, the speech processing device includes a speech acquisition device, and the method further includes: The speech processing device collects a second speech input; The voice processing device generates the target device operation command and / or the second operation command based on a comparison between the voice input and the second voice input, and when the voice input and the second voice input are almost identical, turns off the collection function of the bone conduction device and uses the voice collection device to collect the voice input; The generating of the target device operation command and / or the second operation command based on the comparison of the voice input and the second voice input includes: generating current environment information based on a comparison between the voice input and the second voice input, and generating the operation command and / or the second operation command based on the current environment information; Based on the comparison between the voice input and the second voice input, determining a multi-person interaction scenario, and generating the operation command and / or the second operation command based on the multi-person interaction scenario; and Based on a comparison between the voice input and the second voice input, filtering out irrelevant information from the second voice input, and generating the second operation command based on the processed second voice input; The generating of the current environment information is specifically as follows: In response to the second voice input acquired by air conduction including the voice input waveform acquired by bone conduction and other waveforms, it is determined that the current environment information is noisy; The determination of the multi-person interaction scenario is specifically as follows: In response to the second voice input acquired by air conduction including the voice input waveform acquired by bone conduction and the input waveform from other users, it is determined that multiple people are participating in the voice interaction; The processing of filtering out irrelevant information from the second voice input is specifically as follows: Determine and filter out irrelevant information in the second voice input through comparison; The server is used to perform semantic recognition on the second voice input to generate and issue a second operation command for the target device operation corresponding to the recognized semantics.
29. The method of claim 28, wherein: The speech processing device realizes semantic recognition of the speech input and generates a corresponding target device operation command, including: The voice processing device uploads the voice input to the server; The server recognizes the semantics of the voice input to obtain an operation command for the target device operation corresponding to the recognized semantics.
30. The method of claim 28, wherein The voice processing device realizes semantic recognition of the second voice input and generates a second operation command corresponding to the target device.
Citation Information
Patent Citations
Audio data acquisition method, device, storage medium and terminal
CN109240639A