Voice acquisition and recognition system
By establishing communication connections between multiple electronic devices and utilizing multiple voice acquisition and processing modules, the problem of poor voice recognition performance of electronic devices in noisy environments or at long distances is solved, resulting in clearer voice recognition and a better user experience.
Patent Information
- Application Number
- CN202423060426.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Utility models(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2034-12-11
AI Technical Summary
Existing electronic devices are easily affected by environmental noise and distance when recognizing user voice commands, resulting in poor voice recognition performance and affecting user experience.
By establishing communication connections between multiple electronic devices and utilizing multiple voice acquisition and processing modules, multi-channel sound information can be acquired and aggregated. The voice processing module of the target electronic device is then used for processing to improve sound pickup quality and voice recognition performance.
In noisy environments or at long distances, collaborative data collection and processing by multiple devices improves the clarity and accuracy of speech recognition, thus enhancing the user experience.
Smart Images

Figure CN223679814U_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent speech recognition, and particularly relates to a speech collection and recognition system. BACKGROUND
[0002] With the development of science and technology, electronic devices have gradually become a part of people's daily life. Electronic devices include electronic products and smart appliances, such as mobile phones, earphones, smart glasses, smart sound boxes, sweeping robots, smart air conditioners, smart refrigerators, etc. These devices are all equipped with microphone sensors to receive voice instructions from users and perform corresponding operations according to the voice instructions from users.
[0003] Currently, electronic products and smart appliances use their own sound wind sensors to collect user voices when identifying voice instructions from users. However, if the user is far away from the target electronic device or the environment where the electronic device is located is relatively noisy, the sound sensor on the electronic device will have poor voice recognition effect, which affects the user experience. CONTENT OF THE UTILITY MODEL
[0004] The present application aims to provide a speech collection and recognition system, which aims to solve the technical problem that current electronic devices are easily affected by the environment and have poor voice recognition effect.
[0005] To achieve the above-mentioned purpose, the present application provides the following scheme:
[0006] A speech collection and recognition system is applied to multiple electronic devices, the multiple electronic devices are communicatively connected, and the speech collection and recognition system comprises multiple speech collection modules and at least one speech processing module.
[0007] The multiple electronic devices comprise a first electronic device and a second electronic device, the first electronic device is a target electronic device, and the second electronic device is any of the electronic devices in the multiple electronic devices except the first electronic device.
[0008] At least one of the second electronic devices is configured to collect sound information of a sound source and send the collected sound information to the first electronic device.
[0009] The speech collection and recognition system provided by the present application has the following beneficial effects:
[0010] In the embodiment, the voice collection and recognition system collects sound information of the same sound source by using the voice collection module of one or more second electronic devices, realizes the input of multi-path sound information, and transmits the collected sound information to the first electronic device, so that the multi-path sound information can be collected on the target electronic device, and the voice processing module of the target electronic device processes the multi-path sound information to obtain clear sound, thereby improving the sound pickup quality and improving the voice recognition effect. BRIEF DESCRIPTION OF DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of the drawings shown.
[0012] Figure 1 Fig. 1 is a first schematic block diagram of a voice collection and recognition system provided by the embodiment of the present application;
[0013] Figure 2 Fig. 2 is a second schematic block diagram of a voice collection and recognition system provided by the embodiment of the present application;
[0014] Figure 3 Fig. 3 is a third schematic block diagram of a voice collection and recognition system provided by the embodiment of the present application;
[0015] Figure 4 Fig. 4 is a fourth schematic block diagram of a voice collection and recognition system provided by the embodiment of the present application.
[0016] Explanation of reference numerals:
[0017] 1, voice collection and recognition system; 2, electronic device; 21, first electronic device; 22, second electronic device. DETAILED DESCRIPTION
[0018] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0019] It should be noted that all directional indications, such as upper, lower, left, right, front, back, and the like, are merely used for convenience of description and are not intended to limit the application to a particular orientation.
[0020] It should also be noted that when an element is referred to as being "fixed" or "set" on another element, it can be directly on the other element or can be indirectly on the other element with intervening elements. When an element is referred to as being "connected" or "coupled" to another element, it can be directly connected or coupled to the other element or can be indirectly connected or coupled to the other element through intervening elements.
[0021] In addition, the descriptions involving "first", "second", and the like in the present application are only for the purpose of description and cannot be understood as indicating or implying the relative importance of the technical features indicated or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first", "second" can explicitly or implicitly include at least one of the features. In addition, the technical solutions of various embodiments can be combined with each other, but it must be based on the realization of a person skilled in the art, when the combination of technical solutions appears contradictory or cannot be realized, it should be considered that the combination of technical solutions does not exist, nor within the scope of protection claimed in the present application.
[0022] As shown in Figure 1 The embodiments of the present application provide a voice acquisition and recognition system 1 applied to a plurality of electronic devices 2, which are communicatively connected, for example, connected to the same network to perform data transmission through the same network. Each electronic device 2 can receive user instructions and execute corresponding tasks or commands according to the instructions. The plurality of electronic devices 2 includes a first electronic device 21 and a second electronic device 22. The first electronic device 21 is a target electronic device, which can be understood as an electronic device 2 that the user wants to control, such as a smart air conditioner, a smart sound, or an electronic device 2 used in a call scenario, such as a mobile phone, a telephone watch, etc. The second electronic device 22 is any electronic device 2 in the plurality of electronic devices 2 except the first electronic device 21.
[0023] Specifically, the voice acquisition and recognition system 1 includes a plurality of voice acquisition modules and at least one voice processing module. The at least one second electronic device 22 is configured to acquire sound information of a sound source and send the acquired sound information to the first electronic device 21. The sound source can be a human voice.
[0024] In the embodiment, the voice collection and recognition system 1 collects sound information of the same sound source by using the voice collection module of one or more second electronic devices 22, realizes the input of multi-channel sound information, and transmits the collected sound information to the first electronic device 21 by the second electronic device 22 since the multiple electronic devices 2 are connected to the same network, so that the multi-channel sound information can be collected on the target electronic device, processed by the voice processing module of the target electronic device, and clear sound is obtained, thereby improving the sound pickup quality and enhancing the voice recognition effect.
[0025] As shown in Figure 2 In some embodiments, the voice collection and recognition system 1 further comprises a network communication device, which is used to receive the sound information collected by the second electronic device 22 and send the collected sound information to the first electronic device 21 for processing.
[0026] In combination Figure 2 In some embodiments, at least one voice processing module is built in one or more of the first electronic device 21, the second electronic device 22, and the network communication device, so that one or more of the first electronic device 21, the second electronic device 22, and the network communication device can process the sound information.
[0027] As shown in Figure 3 In some embodiments, the first electronic device 21 and the second electronic device 22 each have a voice collection module and a voice processing module, the voice collection module can transmit a signal to the voice processing module, at least two voice collection modules are used to collect sound signals of a sound source, and the voice processing module of the first electronic device 21 is used to perform noise reduction and / or amplification processing on the sound signals to obtain clear sound, thereby improving the sound pickup quality of the electronic device 2 and enhancing the voice recognition effect of the electronic device 2.
[0028] Specifically, as shown in Figure 3As shown, the first electronic device 21 is a device that the user needs to interact with voice, and the voice interaction can include conversation, control command, etc. When the first electronic device 21 and the second electronic device 22 are in a state capable of collecting the sound information of the sound source, that is, the distance between the first electronic device 21 and the second electronic device 22 and the sound source is appropriate, in this case, the first electronic device 21 and the second electronic device 22 can simultaneously collect the sound information of the sound source, and the second electronic device 22 sends the collected sound information to the first electronic device 21. Compared with collecting sound information only by the voice collection module of the first electronic device 21 itself, the embodiment adopts multiple voice collection modules to collect the sound information of the sound source, which can make the first electronic device 21 obtain a stronger signal, so that the first electronic device 21 can obtain clearer sound. When the first electronic device 21 is far away from the sound source or the environment where the first electronic device 21 is located is relatively noisy, resulting in that the first electronic device 21 cannot collect the sound information of the sound source, but the second electronic device 22 is in a state capable of collecting the sound information of the sound source, the second electronic device 22 can collect the sound information of the sound source and send it to the first electronic device 21, and then the voice processing module of the first electronic device 21 performs noise reduction and amplification processing, so that the first electronic device 21 can also obtain clear sound, and then the target electronic device can improve the voice recognition effect, and also can realize long-distance sound reception.
[0029] In the embodiments of the present application, the sound information is at least one of voice instruction information, voice call information, and voice interaction information, in combination with Figure 3 In some embodiments, the sound source is a user's voice with control instructions, in which case the sound information includes instruction information, voice interaction information, etc. In other embodiments, the sound source is a user's voice during a voice call, for example, the user's voice when using a mobile phone or a smart watch to make a call. In this case, the sound information also includes voice call information and voice conversation information.
[0030] In combination with Figure 3 In some embodiments, the plurality of electronic devices 2 includes at least two of a mobile phone, a tablet, a watch, a headset, smart glasses, a smart speaker, a sweeping robot, a smart air conditioner, a smart refrigerator, a smart television, a smart washing machine, a smart lamp, a smart kitchen appliance, and a smart water heater, to realize the collection and identification of multiple sound information.
[0031] As Figure 3 and Figure 4 As shown, in some embodiments, each voice processing module includes a converter and a processor. The converter of the first electronic device 21 is configured to convert the sound information into a first digital signal; the converter of the second electronic device 22 is configured to convert the sound information into a second digital signal; and the processor of the first electronic device 21 is configured to process the first digital signal and / or the second digital signal.
[0032] In this embodiment, the voice collection and recognition system 1 collects sound information of the same sound source by using multiple voice collection modules, and converts the sound information collected by the multiple voice collection modules into multiple digital signals by using multiple converters respectively corresponding to the multiple voice collection modules. Since the multiple electronic devices 2 are connected to the same network, the multiple digital signals (the first digital signal and the second digital signal) can be collected on the target electronic device, and the processor of the target electronic device can process the multiple digital signals, so that clear sound is obtained, the sound pickup quality is improved, and the voice recognition effect is improved.
[0033] In combination Figure 4 In some embodiments, the converter can be an existing analog-to-digital converter (ADC) for converting sound information into a digital signal. The analog-to-digital converter can also perform noise reduction processing on the sound information, such as removing noise in the signal by filtering and other processing means, thereby achieving the noise reduction effect of the sound information.
[0034] In combination Figure 4 In some embodiments, the processor can be an existing digital signal processor (DSP) for processing the first digital signal and / or the second digital signal to obtain clear sound. Specifically, the digital signal processor can control the gain of the sound information by using digital signal processing technology, thereby amplifying the sound.
[0035] As Figure 3 and Figure 4 shown, in some embodiments, the voice collection module includes a voiceprint recognizer and a sound sensor in a standby wake-up state. The voiceprint recognizer is used to identify the identity of the collected sound information according to the preset voiceprint information, thereby achieving identity verification and improving the security of the system. The sound sensor can be a microphone sensor for collecting sound information of the sound source.
[0036] In combination Figure 3 and Figure 4 In a specific application scenario, the first electronic device 21 and the second electronic device 22 are in communication connection. Specifically, the first electronic device 21 and the second electronic device 22 are connected to the same network, and the first electronic device 21 and the second electronic device 22 are located at different positions in the same area, for example, distributed at different positions of a house. Among the multiple electronic devices 2, the first electronic device 21 (i.e., the target electronic device) is an air conditioner, and the other electronic devices 2 except the first electronic device 21 can be at least two of smart glasses, a smart speaker, a sweeping robot, and a smart refrigerator. The sound source is a sound with a control instruction issued by a user, and the control instruction can be opening the air conditioner, closing the air conditioner, opening the cooling mode, opening the heating mode, etc., which can be selected as needed.
[0037] Moreover, the first electronic device 21 is far away from the sound source, and the first electronic device 21 cannot collect sound information of the sound source or the collected sound information is relatively weak, but at least two of the other electronic devices 2 in the plurality of electronic devices 2 except the first electronic device 21 are distributed near the sound source. In this case, the second electronic device 22 can collect sound information of the sound, and send the sound information to the first electronic device 21 after processing. Thus, the processor of the first electronic device 21 can perform noise reduction and amplification processing on the obtained digital signals to obtain clear sound with control instructions, so as to perform operations such as opening the air conditioner, closing the air conditioner, opening the cooling mode, or opening the heating mode according to the control instructions, so as to call the voice sensor of the electronic device 2 near the user to perform voice input when remotely controlling the target electronic device 2, and improve the user experience. It should be understood that the first electronic device 21 can also be an electronic device 2 such as a smart speaker, a sweeping robot, a smart refrigerator, or smart glasses. For example, when the first electronic device 21 is a smart speaker, the first electronic device 21 can obtain pure sound with control instructions from the obtained digital signals. In this case, the control instructions can be turning on the sound, increasing the volume, decreasing the volume, playing the previous song, playing the next song, and the like.
[0038] In combination with Figure 3 And Figure 4 In another specific application scenario, the first electronic device 21 and the second electronic device 22 are in communication connection. Specifically, the first electronic device 21 and the second electronic device 22 are connected on the same network, and the first electronic device 21 and the second electronic device 22 are located at different positions in the same area, for example, distributed at different positions of a house. Among the plurality of electronic devices 2, the first electronic device 21 (i.e., the target electronic device) is a mobile phone or a watch that can perform voice communication, and the other electronic devices 2 except the first electronic device 21 can be at least two of smart glasses, a smart speaker, a sweeping robot, and a smart refrigerator. In this case, the first electronic device 21 can collect the user's communication sound through its own sound sensor, and convert the sound information into a first digital signal through its own transducer, and send the first digital signal to its own processor. At the same time, the second electronic device 22 can collect the user's communication sound through its own sound sensor, and convert the sound information into a second digital signal through its own transducer, and send the second digital signal to the processor of the first electronic device 21. The processor of the first electronic device 21 processes the first digital signal and the second digital signal, so that the first electronic device 21 can obtain clear communication sound in a relatively noisy environment, and send the communication sound to the communication counterpart of the user in the form of a signal, so that the communication counterpart can hear the clear communication sound of the user, thereby improving the communication quality and improving the user experience.
[0039] It should be understood that in the above application scenarios, the plurality of electronic devices 2 can be arranged in any manner, and the embodiments are not specifically limited.
[0040] As shown in Figure 3 and Figure 4 In some embodiments, the plurality of electronic devices 2 are connected by a wireless communication module, which is any one of a Bluetooth module, a WIFI module, a ZigBee module, and a Thread module, to realize signal exchange between the plurality of electronic devices 2. In specific applications, the second electronic device 22 can send the second digital signal to the first electronic device 21 through the wireless communication module, or send it to a network communication terminal for voice signal conversion and processing, and then send it to the first electronic device 21. In a specific embodiment, in the first electronic device 21, the sound sensor collects sound information of the sound source and sends the sound information to the analog-to-digital converter, which converts the sound information into the first digital signal, and then sends the first digital signal to the digital signal processor; at the same time, in the second electronic device 22, the sound sensor collects sound information of the sound source and sends the sound information to the analog-to-digital converter, which converts the sound information into the second digital signal, and the second electronic device 22 sends the second digital signal to the digital signal processor of the first electronic device 21 through the wireless communication module, and the digital signal processor processes the first digital signal and the second digital signal to obtain clear instructions and then performs corresponding operations.
[0041] In combination Figure 3 In some embodiments, the distance between the first electronic device 21 and the sound source is greater than the distance between the second electronic device 22 and the sound source, or the voice collection module of the first electronic device 21 fails, resulting in that the first electronic device 21 cannot collect sound information of the sound source, which is voice instruction information. In this case, the second electronic device 22 collects the sound information of the sound source and processes the collected sound information into the second digital signal and sends it to the first electronic device 21, which is processed by the voice processing module of the first electronic device 21, so that the first electronic device 21 can also obtain clear sound, and then the first electronic device 21 executes corresponding operations according to the control instructions in the sound, realizing remote control of the first electronic device 21.
[0042] The above only describes the preferred embodiments of the present application, and does not limit the patent scope of the present application, and any equivalent structural transformation made by using the contents of the present application specification and drawings, or direct / indirect application in other related technical fields is included in the patent protection scope of the present application.
Claims
1. A voice acquisition and recognition system, applied to multiple electronic devices, wherein the multiple electronic devices are communicatively connected, characterized in that, The voice acquisition and recognition system includes multiple voice acquisition modules and at least one voice processing module; The plurality of electronic devices includes a first electronic device and a second electronic device, wherein the first electronic device is the target electronic device, and the second electronic device is any one of the plurality of electronic devices other than the first electronic device; At least one of the second electronic devices is used to collect sound information from a sound source and send the collected sound information to the first electronic device.
2. The voice acquisition and recognition system according to claim 1, characterized in that, The voice acquisition and recognition system also includes a network communication device, which is used to receive the sound information acquired by the second electronic device and send the acquired sound information to the first electronic device.
3. The voice acquisition and recognition system according to claim 2, characterized in that, The first electronic device, the second electronic device, and one or more of the network communication devices have the at least one voice processing module built into them.
4. The voice acquisition and recognition system according to claim 1, characterized in that, Both the first electronic device and the second electronic device have the aforementioned voice acquisition module and voice processing module. At least two of the voice acquisition modules are used to acquire sound information from a sound source, and the voice processing module of the first electronic device is used to perform noise reduction and / or amplification processing on the sound information.
5. The voice acquisition and recognition system according to claim 1, characterized in that, Each of the aforementioned voice processing modules includes a converter and a processor, wherein the converter of the first electronic device is used to convert the sound information into a first digital signal; The converter of the second electronic device is used to convert the sound information into a second digital signal, and the processor of the first electronic device is used to process the first digital signal and / or the second digital signal. The converter is an analog-to-digital converter. And / or, The processor is a digital signal processor.
6. The voice acquisition and recognition system according to any one of claims 1-5, characterized in that, The voice acquisition module includes a voiceprint recognizer and a sound sensor in a wake-up state. The voiceprint recognizer is used to identify the identity of the acquired voice information based on preset voiceprint information, and the sound sensor is used to acquire the sound information of the sound source.
7. The voice acquisition and recognition system according to claim 1, characterized in that, The multiple electronic devices are connected to each other via a wireless communication module, which is any one of Bluetooth, WIFI, ZigBee, and Thread modules.
8. The voice acquisition and recognition system according to claim 1, characterized in that, The plurality of electronic devices includes at least two of the following: mobile phones, tablets, watches, earphones, smart glasses, smart speakers, robot vacuum cleaners, smart air conditioners, smart refrigerators, smart TVs, smart washing machines, smart lamps, smart kitchen appliances, and smart water heaters.
9. The voice acquisition and recognition system according to claim 1, characterized in that, The sound information is one or more of the following: voice command information, voice call information, and voice interaction information.
10. The voice acquisition and recognition system according to claim 1, characterized in that, The distance between the first electronic device and the sound source is greater than the distance between the second electronic device and the sound source.