Noise processing method, device, and system

CA3160740CActive Publication Date: 2026-09-2210353744 CANADA LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CA3160740
Authority / Receiving Office
CA · CA
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-11-13
Filing Date
2020-07-30
Publication Date
2026-09-22
Estimated Expiration
2040-07-30
Patent Text Reader

Abstract

Embodiments of the present application disclose a noise processing method, and corresponding device and system, of which the method comprises: detecting collected audio information; filtering, when voice information is detected, the voice information according to prestored audio information of a target user; judging whether there is voice information after the filtering; and recognizing, if there is, the filtered voice information and making corresponding feedback according to a recognition result. The obtained voice of a target person is taken as priori information in the present application, accordingly, when a non-target person sends out an instruction, the instruction of the non-target person can be suppressed according to the priori information, and when there are other human voice interferences and environmental noises while a target person sends out an instruction, the human voice interferences and environmental noises at the same location, at a near location and at a remote location can be suppressed.
Need to check novelty before this filing date? Find Prior Art

Description

NOISE PROCESSING METHOD, DEVICE, AND SYSTEM BACKGROUND OF THE INVENTION Technical Field

[0001] The present invention relates to the field of acoustics, and more particularly to a noise processing method, and corresponding device and system. Description of Related Art

[0002] With the development of artificial intelligence, more intellectualization will be reflected in more and more such living environments as on-board environment, home environment, classroom environment and meeting-room environment, etc. Of the variegated intelligent equipments applied in these environments, the intelligent voice interaction equipment plays an important role. The intelligent voice interaction equipment achieves voice interaction between humans and equipments, whereby the equipments are enabled to perform certain operations and controls in place of humans and according to the ideas of humans, the human hands are liberated as far as possible, so the intelligent voice interaction equipment is an indispensable intelligent equipment in the future.

[0003] Since an actual living environment is usually very complicated, besides the sound of the target person, there are also many noises and interfering sounds. These noises and interfering sounds are not what we desire, their presence severely interferes with the interaction between humans and voice equipments, and deteriorates the interactive experience. In order to avoid the interference of these noises and interfering sounds, a microphone array is usually employed to perform beam forming or blind source separation, to strengthen the sound of a particular direction, to suppress the sounds of other directions, or to separate the sounds of particular target persons.

[0004] However, the traditional beam forming or blind source separation cannot effectively suppress interference or effectively separate the target sound under all environments. When an interfering sound is likewise a human sound, and is very close to, at the same location as, or very remote from the location of the target sound, the effect of the aforementioned method will be reduced abruptly. SUMMARY OF THE INVENTION

[0005] In order to solve problems pending in the state of the art, the present invention proposes a noise processing method, and corresponding device and system. The method can not only address interference from environmental noises, but can also deal with interference from human voices at the same location, at a near location, or at a remote location, whereby the interactive experience between humans and equipments is enhanced.

[0006] Specific technical solutions provided by the embodiments of the present invention are as follows.

[0007] According to the first aspect, the present invention provides a noise processing method that comprises:

[0008] detecting collected audio information;

[0009] filtering, when voice information is detected, the voice information according to prestored audio information of a target user;

[0010] judging whether there is voice information after the filtering; and

[0011] recognizing, if there is, the filtered voice information and making corresponding feedback according to a recognition result.

[0012] Preferably, the step of filtering the voice information according to prestored audio information of a target user specifically includes:

[0013] constructing an acoustic model, wherein the acoustic model is a Gaussian mixture model whose variable is the voice information and initial values of whose parameters are a covariance matrix obtained after calculating the audio information of the target user;

[0014] rectifying the parameters of the acoustic model according to an EM algorithm;

[0015] judging whether number of iterations of the EM algorithm reaches a preset value;

[0016] obtaining an output result of the acoustic model when the preset value is reached; and

[0017] filtering the voice information according to the output result.

[0018] Preferably, when voice information is detected, the method further comprises:

[0019] performing echo cancellation on the voice information.

[0020] Preferably, the method further comprises:

[0021] sending an operation instruction to the target user according to a received request sent from the target user;

[0022] receiving audio information sent from the target user according to the operation instruction; and

[0023] storing the audio information sent from the target user according to the operation instruction.

[0024] Preferably, an algorithm that detects the collected audio information includes any of a pitch detection algorithm, a double-threshold method, and a posteriori SNR frequency domain iterative algorithm.

[0025] According to the second aspect, the present invention provides a noise processing device that comprises:

[0026] a detecting module, for detecting collected audio information;

[0027] an analyzing module, for filtering, when voice information is detected, the voice information according to prestored audio information of a target user;

[0028] a judging module, for judging whether there is voice information after the filtering; and

[0029] a recognizing module, for recognizing, if there is, the filtered voice information and making corresponding feedback according to a recognition result.

[0030] Preferably, the analyzing module specifically includes:

[0031] a constructing module, for constructing an acoustic model, wherein the acoustic model is a Gaussian mixture model whose variable is the voice information and initial values of whose parameters are a covariance matrix obtained after calculating the audio information of the target user;

[0032] a rectifying module, for rectifying the parameters of the acoustic model according to an EM algorithm; and

[0033] a processing module, for judging whether number of iterations of the EM algorithm reaches a preset value, obtaining an output result of the acoustic model when the preset value is reached, and filtering the voice information according to the output result.

[0034] Preferably, the analyzing module further includes:

[0035] an echo cancelling module, for performing echo cancellation on voice information when the voice information is detected.

[0036] Preferably, the device further comprises a storing module, for:

[0037] sending an operation instruction to the target user according to a received request sent from the target user;

[0038] receiving audio information sent from the target user according to the operation instruction; and

[0039] storing the audio information sent from the target user according to the operation instruction.

[0040] According to the third aspect, the present invention provides a computer system that comprises:

[0041] one or more processor(s); and

[0042] a memory, associated with the one or more processor(s), for storing a program instruction that executes the following operations when read and executed by the one or more processor(s):

[0043] detecting collected audio information;

[0044] filtering, when voice information is detected, the voice information according to prestored audio information of a target user;

[0045] judging whether there is voice information after the filtering; and

[0046] recognizing, if there is, the filtered voice information and making corresponding feedback according to a recognition result.

[0047] The embodiments of the present invention achieve the following advantageous effects.

[0048] The present invention firstly obtains voice of a target person to serve as priori information, accordingly, when a non-target person sends out an instruction, the instruction of the non- target person can be suppressed according to the priori information, and when there are other human voice interferences and environmental noises while a target person sends out an instruction, the human voice interferences and environmental noises at the same location, at a near location and at a remote location can be suppressed according to the priori information, so as to obtain an instruction free from other human voices and environmental noises, whereby sound clarity of the target person is enhanced, and interactive experience is enhanced. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] To more clearly describe the technical solutions in the embodiments of the present invention, drawings required to be used in the description of the embodiments will be briefly introduced below. Apparently, the drawings introduced below are merely directed to some embodiments of the present invention, while it is possible for persons ordinarily skilled in the art to acquire other drawings based on these drawings without spending creative effort in the process.

[0050] Fig. 1 is a view illustrating an application environment for a noise processing method provided by the embodiments of the present application;

[0051] Fig. 2 is a flowchart illustrating a noise processing method provided by Embodiment 1 of the present application;

[0052] Fig. 3 is a view schematically illustrating the structure of a noise processing device provided by Embodiment 2 of the present application;

[0053] Fig. 4 is a view schematically illustrating positions of a noise processing device and experimenting users provided by Embodiment 2 of the present application; and

[0054] Fig. 5 is a view illustrating the framework of a computer system provided by Embodiment 3 of the present application. DETAILED DESCRIPTION OF THE INVENTION

[0055] To make more lucid and clear the objectives, technical solutions and advantages of the present invention, the technical solutions in the embodiments of the present invention will be more clearly and comprehensively described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the embodiments as described are merely partial, rather than the entire, embodiments of the present invention. All other embodiments obtainable by persons ordinarily skilled in the art on the basis of the embodiments in the present invention without spending creative effort shall all fall within the protection scope of the present invention.

[0056] The noise processing method provided by the present application is applicable to the application environment as shown in Fig. 1, in which servicing end 12 communicates with database 11 and terminal 13 through network. Terminal 13 can be, but is not limited to be, any of various personal computers, notebook computers, smart mobile phones, panel computers and portable wearable devices, and servicing end 12 can be embodied as an independent servicing end or a servicing end cluster consisting of a plurality of servicing ends.

[0057] Embodiment 1

[0058] As shown in Fig. 2, the present application provides a noise processing method that specifically comprises the following steps.

[0059] S21 - detecting collected audio information.

[0060] The detection algorithm can include any of a pitch detection algorithm, a double- threshold method, and a posteriori SNR frequency domain iterative algorithm.

[0061] In addition, the detection algorithm can as well be any other algorithm that can realize voice breakpoint detection, and selection of the algorithm is not restricted in this solution.

[0062] S22 - filtering, when voice information is detected, the voice information according to prestored audio information of a target user.

[0063] When voice information is detected, the method further comprises the following step:

[0064] performing echo cancellation on the voice information.

[0065] The echo in this solution means acoustic echo, when echo cancellation is performed, this can be realized by such an acoustic echo cancellation method frequently employed in this field of specialty as the echo suppression algorithm or the acoustic echo cancellation algorithm, to which no restriction is made in the present invention.

[0066] The detected voice information contains environmental noises and / or human sound interfering noises.

[0067] The step of filtering the voice information according to prestored audio information of a target user specifically includes the following steps:

[0068] 1. constructing an acoustic model,

[0069] wherein the acoustic model is a Gaussian mixture model whose variable is the voice information and initial values of whose parameters are a covariance matrix obtained after calculating the audio information of the target user;

[0070] the Gaussian mixture model (GMM) can be expressed by the following formula: [Image disponible dans le document PDF, Image available in the PDF document]

[0072] where x is the voice information, <semantics>N(x|μk,∑k)<annotation encoding="application / x-tex">N(x|\mu_k, \sum k)< / annotation>< / semantics> is the kth component in the model, <semantics>πk<annotation encoding="application / x-tex">\pi_k< / annotation>< / semantics> is a mixture coefficient, namely the weight of each component, <semantics>πk<annotation encoding="application / x-tex">\pi_k< / annotation>< / semantics>, <semantics>μk<annotation encoding="application / x-tex">\mu_k< / annotation>< / semantics>, <semantics>∑k<annotation encoding="application / x-tex">\sum k< / annotation>< / semantics> are parameters of the Gaussian mixture model, and their initial values are a covariance matrix obtained after calculating the audio information of the target user;

[0073] 2. rectifying the parameters of the acoustic model according to an EM algorithm,

[0074] wherein the EM algorithm is an expectation maximum algorithm;

[0075] the aforementioned step 2 specifically includes the following two sub-steps:

[0076] a. calculating a posteriori probability according to the initial values of the current parameters; and

[0077] b. rectifying the parameters according to the posteriori probability;

[0078] 3. judging whether number of iterations of the EM algorithm reaches a preset value;

[0079] in this solution, the number of iterations is set according to empirical values, when the number of times by which the EM algorithm is executed reaches a preset value (the number of executions of the aforementioned steps a, b), this means that iteration ends at this time;

[0080] 4. obtaining an output result of the acoustic model when the preset value is reached;

[0081] the output result is precisely the posteriori probability calculated and obtained according to the parameters at the last iteration;

[0082] 5. filtering the voice information according to the output result.

[0083] It is thusly made possible to effectively suppress environmental noises, and human sounds at the same location, human sounds at a near location, and human sounds at a remote location.

[0084] S23 – judging whether there is voice information after the filtering.

[0085] When there is no voice information, this indicates that the detected voice information is voice uttered by a non-target user; when there is voice information, this indicates that the detected voice information contains voice uttered by the target user.

[0086] S24 – recognizing, if there is, the filtered voice information and making corresponding feedback according to a recognition result.

[0087] Specifically, the filtered voice information is transformed to text content and recognized by means of the word-segmentation technology, etc., to judge the user's intention, and to make corresponding feedback, and an evaluating indicator is further output at the same time for appraising the exactness of the voice recognizing process.

[0088] The evaluating indicator can be sentence error rate (SER), sentence correct rate (S. Corr) and word / character error rate (WER / CER), etc.

[0089] Besides, obtainment of prestored audio information of the target user includes the following steps:

[0090] 1. sending an operation instruction to the target user according to a received request sent from the target user;

[0091] this method is applicable to an intelligent voice interaction equipment, so the request sent from the target user can be a request to reset the equipment; in accordance with the request sent from the target user, an operation instruction is sent to the target user, for instance:

[0092] voice reminder: "please well adjust sitting posture", "please utter little Biu little Biu", "please tilt your head to the left for about 10 cm before uttering little Biu little Biu", "please tilt your head to the right for about 10 cm before uttering little Biu little Biu", "please tilt your body forwards for about 10 cm before uttering little Biu little Biu", and so on;

[0093] 2. receiving audio information sent from the target user according to the operation instruction;

[0094] the target user sends corresponding audio information according to the operation instruction, for instance, on receiving an operation instruction "please well adjust sitting" posture", the target user replies "sitting posture already adjusted"; on receiving an operation instruction "please utter little Biu little Biu", the target user replies "little Biu" little Biu"; on receiving an operation instruction "please tilt your head to the left for about 10 cm before uttering little Biu little Biu", the target user continues to perform the corresponding action in accordance with the instruction and makes a reply;

[0095] as should be noted, when there are plural operation instructions, these operation instructions are sent according to a set time interval,

[0096] for example, sending an operation instruction at every 2 seconds;

[0097] 3. storing the audio information sent from the target user according to the operation instruction;

[0098] for instance, such voices as replied by the target user, "sitting posture already adjusted", "little Biu little Biu", are stored.

[0099] The present invention firstly obtains voice of a target person to serve as priori information, accordingly, when a non-target person sends out an instruction, the instruction of the non- target person can be suppressed according to the priori information, and when there are other human voice interferences and environmental noises while a target person sends out an instruction, the human voice interferences and environmental noises at the same location, at a near location and at a remote location can be suppressed according to the priori information, so as to obtain an instruction free from other human voices and environmental noises, whereby sound clarity of the target person is enhanced, and interactive experience is enhanced.

[0100] Embodiment 2

[0101] As shown in Fig. 3, the present application provides a noise processing device that specifically comprises:

[0102] a detecting module 31, for detecting collected audio information;

[0103] an analyzing module 32, for filtering, when voice information is detected, the voice information according to prestored audio information of a target user;

[0104] a judging module 33, for judging whether there is voice information after the filtering; and

[0105] a recognizing module 34, for recognizing, if there is, the filtered voice information and making corresponding feedback according to a recognition result.

[0106] Preferably, the analyzing module 32 specifically includes:

[0107] a constructing module 321, for constructing an acoustic model, wherein the acoustic model is a Gaussian mixture model whose variable is the voice information and initial values of whose parameters are a covariance matrix obtained after calculating the audio information of the target user;

[0108] a rectifying module 322, for rectifying the parameters of the acoustic model according to an EM algorithm; and

[0109] a processing module 323, for judging whether number of iterations of the EM algorithm reaches a preset value, obtaining an output result of the acoustic model when the preset value is reached, and filtering the voice information according to the output result.

[0110] Preferably, the analyzing module 32 further includes:

[0111] an echo cancelling module 324, for performing echo cancellation on voice information when the voice information is detected.

[0112] The device further comprises a storing module 35, for:

[0113] sending an operation instruction to the target user according to a received request sent from the target user;

[0114] receiving audio information sent from the target user according to the operation instruction; and

[0115] storing the audio information sent from the target user according to the operation instruction.

[0116] Preferably, the algorithm that detects the collected audio information includes any of a pitch detection algorithm, a double-threshold method, and a posteriori SNR frequency domain iterative algorithm.

[0117] When the aforementioned noise processing device is an intelligent interaction equipment, the intelligent interaction equipment comprises a voice interacting system and a voice recognizing system, of which the voice interacting system includes the aforementioned detecting module 31, analyzing module 32, judging module 33 and storing module 35, and the voice recognizing system includes the aforementioned recognizing module 34.

[0118] An interaction experiment is carried out by means of the aforementioned intelligent inaction equipment, and users are arranged according to preset positions.

[0119] Refer to Fig. 4, in which are included 5 users that are respectively user No. 1, user No. 2, user No. 3, user No. 4 and user No. 5.

[0120] The experimenting process is as follows:

[0121] 1. user No. 1 and user No. 2 simultaneously speak, and user No. 1 is the target user;

[0122] 2. user No. 1 and user No. 3 simultaneously speak, and user No. 1 is the target user;

[0123] 3. user No. 1 and user No. 4 simultaneously speak, and user No. 1 is the target user;

[0124] 4. user No. 1 and user No. 5 simultaneously speak, and user No. 1 is the target user;

[0125] 5. user No. 1, user No. 2 and user No. 3 simultaneously speak, and user No. 1 is the target user;

[0126] 6. user No. 1, user No. 3 and user No. 4 simultaneously speak, and user No. 1 is the target user;

[0127] 7. user No. 1, user No. 4 and user No. 5 simultaneously speak, and user No. 1 is the target user;

[0128] 8. all users simultaneously speak, and user No. 1 is the target user;

[0129] The recognizing module 34 in the voice recognizing system is employed for recognizing the filtered voice information and making corresponding feedback according to the recognition result, and additionally employed for outputting an evaluating indicator for appraising the exactness of the voice recognizing process.

[0130] In the present solution, the evaluating indicator is WER (word error rate).

[0131] Experiment results derived from the above experiment are as shown in the following Table 1.

[0132] Table 1

[0133] [Image disponible dans le document PDF, Image available in the PDF document]

[0134] The prior-art noise-reducing method is likewise a Gaussian mixture model, its variable is the voice information, and the initial value of the parameter is a preset value, rather than the covariance matrix obtained after calculating the audio information of the target user as in the present solution; in addition, this Gaussian mixture model employs the EM algorithm for parameter rectification, and the optimal parameter is obtained by an adaptive algorithm during rectification.

[0135] As can be deduced according to the above experiment results, since the present application utilizes the audio information of the target user as the priori information, the effect of subsequent voice recognition can be enhanced, and interactive experience is hence enhanced.

[0137] As shown in Fig. 5, Embodiment 3 of the present application provides a computer system that comprises:

[0138] one or more processor(s); and

[0139] a memory, associated with the one or more processor(s), for storing a program instruction that executes the following operations when read and executed by the one or more processor(s):

[0140] detecting collected audio information;

[0141] filtering, when voice information is detected, the voice information according to prestored audio information of a target user;

[0142] judging whether there is voice information after the filtering; and

[0143] recognizing, if there is, the filtered voice information and making corresponding feedback according to a recognition result.

[0144] Fig. 5 exemplarily illustrates the framework of a computer system that can specifically include a processor 52, a video display adapter 54, a magnetic disk driver 56, an input / output interface 58, a network interface 510, and a memory 512. The processor 52, the video display adapter 54, the magnetic disk driver 56, the input / output interface 58, the network interface 510, and the memory 512 can be communicably connected with one another via a communication bus 514.

[0145] The processor 52 can be embodied as a general CPU (Central Processing Unit), a microprocessor, an ASIC (Application Specific Integrated Circuit), or one or more integrated circuit(s) for executing relevant program(s) to realize the technical solutions provided by the present application.

[0146] The memory 512 can be embodied in such a form as an ROM (Read Only Memory), an RAM (Random Access Memory), a static storage device, or a dynamic storage device. The memory 512 can store an operating system 516 for controlling the running of a computer system 50, and a basic input / output system (BIOS) 518 for controlling lower- level operations of the computer system. In addition, the memory 512 can also store a web browser 520, a data storage administration system 522, etc. To sum it up, when the technical solutions provided by the present application are to be realized via software or firmware, the relevant program codes are stored in the memory 512, and invoked and executed by the processor 52.

[0147] The input / output interface 58 is employed to connect with an input / output module to realize input and output of information. The input / output module can be equipped in the device as a component part (not shown in the drawings), and can also be externally connected with the device to provide corresponding functions. The input means can include a keyboard, a mouse, a touch screen, a microphone, and various sensors etc., and the output means can include a display screen, a loudspeaker, a vibrator, an indicator light etc.

[0148] The network interface 510 is employed to connect to a communication module (not shown in the drawings) to realize intercommunication between the current device and other devices. The communication module can realize communication in a wired mode (via USB, network cable, for example) or in a wireless mode (via mobile network, WIFI, Bluetooth, etc.).

[0149] The communication bus 514 includes a passageway transmitting information between various component parts of the device (such as the processor 52, the video display adapter 54, the magnetic disk driver 56, the input / output interface 58, the network interface 510, and the memory 512).

[0150] Additionally, the computer system may further obtain information of specific collection conditions from a virtual resource object collection condition information database for judgment on conditions, and so on.

[0151] As should be noted, although merely the processor 52, the video display adapter 54, the magnetic disk driver 56, the input / output interface 58, the network interface 510, the memory 512, and the communication bus 514 are illustrated for the aforementioned device, the device may further include other component parts prerequisite for realizing normal running during specific implementation. In addition, as can be understood by persons skilled in the art, the aforementioned device may as well only include component parts necessary for realizing the solutions of the present application, without including the entire component parts as illustrated.

[0152] As can be known through the description to the aforementioned embodiments, it is clearly learnt by person skilled in the art that the present application can be realized through software plus a general hardware platform. Based on such understanding, the technical solutions of the present application, or the contributions made thereby over the state of the art, can be essentially embodied in the form of a software product, and such a computer software product can be stored in a storage medium, such as an ROM / RAM, a magnetic disk, an optical disk etc., and includes plural instructions enabling a computer equipment (such as a personal computer, a server, or a network device etc.) to execute the methods described in various embodiments or some sections of the embodiments of the present application.

[0153] Although preferred embodiments in the embodiments of the present invention have been described so far, it is possible for persons skilled in the art to make additional change and modification to these embodiments once they have learnt of the basic inventive concept. Accordingly, the present invention is meant to subsume the preferred embodiments and all changes and modifications that fall within the scope in the embodiments of the present invention. In addition, the noise processing device and computer system provided by the above embodiments pertain to the same conception as the noise processing method – see the method embodiment for their specific realization processes, while no repetition is made in this context.

[0154] Apparently, it is possible for persons skilled in the art to make various alterations and variations to the present invention without departing from the principle and scope of the present invention. Accordingly, if such alterations and variations to the present invention fall within the principle and scope of the present invention and the range of equivalent technology, the present invention is also meant to subsume these alterations and variations.

Claims

2. The device of claim 1 further comprising a collecting module communicatively connected to the processor, configured to collect the collected audio information.

3. The device of any one of claims 1 to 2 further comprising: a storage media communicatively connected to the processor configured to store the prestored audio information; and a pre-collecting module communicatively connected to one or more of the storage media and the processor, configured to obtain the prestored audio information and provide the prestored audio information to the storage media.

4. The device of any one of claims 1 to 3 wherein the detection is one or more of a pitch detection, a double threshold method detection, and a posteriori SNR frequency domain iterative detection.

5. The device of any one of claims 1 to 4 wherein the judging module further comprises an echo canceling module configured to one or more of echo cancellation and echo suppression on the collected audio information.

6. The device of any one of claims 1 to 5 wherein the analyzing module is further comprises: a constructing module for constructing an acoustic model; a rectifying module for rectifying the parameters of the acoustic model according to an EM algorithm; a processing module for judging whether a number of iterations of the EM algorithm reaches a preset value; an obtaining module for obtaining an output result of the acoustic model when the preset value is reached; and a filtering module for filtering the voice information according to the output result.

7. The device of claim 6 wherein the acoustic model is a Gaussian mixture model.

8. The device of claim 7 wherein a variable of the Gaussian mixture model is the collected audio information.

9. The device of any one of claims 7 to 8 wherein one or more parameters of initial values of the Gaussian mixture model are a covariance matrix obtained based on calculating an audio information of the target user.

10. The device of any one of claims 1 to 9 wherein the EM algorithm is an expectation maximum algorithm.

11. The device of claim any one of claims 6 to 10 wherein rectifying the parameters of the acoustic model further comprises: calculating a posteriori probability according to one or more initial values of the parameters; and rectifying the parameters according to the posteriori probability.

12. The device of claim any one of claims 6 to 11 wherein the number of iterations are set according to empirical values.

13. The device of any one of claims 6 to 12 wherein preset value is reached if the output result is the calculated posteriori probability at the last iteration.

14. The device of any one of claims 1 to 13 wherein the recognizing module is further configured to transform the filtered audio information to text content.

15. The device of claim 14 wherein the recognizing module if further configured to recognize the filtered audio information by word-segmentation technology.

16. The device of any one of claims 1 to 15 wherein the recognizing module is further configured to output an evaluating indicator for appraising an exactness of a voice recognizing process.

17. The device of claim 16 wherein the evaluating indicator is one or more of a sentence error rate, sentence correct rate, word error rate and character error rate.

18. The device of any one of claims 13 to 17 wherein the pre-collecting module is further configured to send a plurality of operation instructions to the target user based on a received request from the target user and to receive the prestored audio information based on the target user response to the operation instructions.

19. The device of claim 18 wherein the pre-collecting module is further configured to send the operation instructions based on a set time interval.

20. The device of claim 19 wherein the set time interval is 2 seconds.

21. The device of any one of claims 18 to 20 wherein the received request is a request to reset the equipment.

22. A system for making feedback based on a collected audio information the system comprising: a processor for processing the collected audio information further comprising: an obtaining module configured to obtain a collected audio information and a prestored audio information of a target user; a detecting module configured to detect the collected audio information wherein the collected audio information is detected by a voice breakpoint detection; an analyzing module configured to filter the collected audio information based on the prestored audio information wherein audio of the collected audio information remaining after filtering is output as filtered audio information; a judging module configured to judge whether there is filtered audio information wherein filtered audio information remaining of the collected audio information indicates that the collected audio information contains audio uttered by the target user; and a recognizing module configured to recognize the filtered audio information and make corresponding feedback based on the recognition result.

23. The system of claim 22 further comprising a collecting module communicatively connected to the processor, configured to collect the collected audio information.

24. The system of any one of claims 22 to 23 further comprising: a storage media communicatively connected to the processor configured to store the prestored audio information; and a pre-collecting module communicatively connected to one or more of the storage media and the processor, configured to obtain the prestored audio information and provide the prestored audio information to the storage media.

25. The system of any one of claims 22 to 24 wherein the detection is one or more of a pitch detection, a double threshold method detection, and a posteriori SNR frequency domain iterative detection.

26. The system of any one of claims 22 to 25 wherein the judging module further comprises an echo canceling module configured to one or more of echo cancellation and echo suppression on the collected audio information.

27. The system of any one of claims 22 to 26 wherein the analyzing module is further comprises: a constructing module for constructing an acoustic model; a rectifying module for rectifying the parameters of the acoustic model according to an EM algorithm; a processing module for judging whether a number of iterations of the EM algorithm reaches a preset value; an obtaining module for obtaining an output result of the acoustic model when the preset value is reached; and a filtering module for filtering the voice information according to the output result.

28. The system of claim 27 wherein the acoustic model is a Gaussian mixture model.

29. The system of claim 28 wherein a variable of the Gaussian mixture model is the collected audio information.

30. The system of any one of claims 28 to 29 wherein one or more parameters of initial values of the Gaussian mixture model are a covariance matrix obtained based on calculating an audio information of the target user.

31. The system of any one of claims 22 to 30 wherein the EM algorithm is an expectation maximum algorithm.

32. The system of claim any one of claims 27 to 31 wherein rectifying the parameters of the acoustic model further comprises: calculating a posteriori probability according to one or more initial values of the parameters; and rectifying the parameters according to the posteriori probability.

33. The system of claim any one of claims 27 to 32 wherein the number of iterations are set according to empirical values.

34. The system of any one of claims 27 to 33 wherein preset value is reached if the output result is the calculated posteriori probability at the last iteration.

35. The system of any one of claims 22 to 34 wherein the recognizing module is further configured to transform the filtered audio information to text content.

36. The system of claim 35 wherein the recognizing module if further configured to recognize the filtered audio information by word-segmentation technology.

37. The system of any one of claims 22 to 36 wherein the recognizing module is further configured to output an evaluating indicator for appraising an exactness of a voice recognizing process.

38. The system of claim 37 wherein the evaluating indicator is one or more of a sentence error rate, sentence correct rate, word error rate and character error rate.

39. The system of any one of claims 24 to 38 wherein the pre-collecting module is further configured to send a plurality of operation instructions to the target user based on a received request from the target user and to receive the prestored audio information based on the target user response to the operation instructions.

40. The system of claim 39 wherein the pre-collecting module is further configured to send the operation instructions based on a set time interval.

41. The system of claim 40 wherein the set time interval is 2 seconds.

42. The system of any one of claims 39 to 41 wherein the received request is a request to reset the equipment.

43. A computer implemented method for making feedback based on a collected audio information, the method comprising: obtaining the collected audio information and a prestored audio information of a target user; detecting the collected audio information by a voice breakpoint detection; filtering the collected audio information based on the prestored audio information wherein audio of the collected audio information remaining after filtering is output as filtered audio information; judging whether there is filtered audio information wherein filtered audio information remaining of the collected audio information indicates that the collected audio information contains audio uttered by the target user; recognizing the filtered audio information; and making corresponding feedback based on the recognition result.

44. The method of claim 43 further comprising collecting the collected audio information.

45. The method of any one of claims 43 to 44 further comprising: obtaining the prestored audio information; providing the prestored audio information to a storage media; and storing the prestored audio information.

46. The method of any one of claims 43 to 45 wherein the detection is one or more of a pitch detection, a double threshold method detection, and a posteriori SNR frequency domain iterative detection.

47. The method of any one of claims 43 to 46 wherein the judging further comprises one or more of echo cancellation and echo suppression on the collected audio information.

48. The method of any one of claims 43 to 47 wherein the filtering further comprises: constructing an acoustic model; rectifying the parameters of the acoustic model according to an EM algorithm; judging whether a number of iterations of the EM algorithm reaches a preset value; obtaining an output result of the acoustic model when the preset value is reached; and filtering the voice information according to the output result.

49. The method of claim 48 wherein the acoustic model is a Gaussian mixture model.

50. The method of claim 49 wherein a variable of the Gaussian mixture model is the collected audio information.

51. The method of any one of claims 49 to 50 wherein one or more parameters of initial values of the Gaussian mixture model are a covariance matrix obtained based on calculating an audio information of the target user.

52. The method of any one of claims 43 to 51 wherein the EM algorithm is an expectation maximum algorithm.

53. The method of claim any one of claims 48 to 52 wherein rectifying the parameters of the acoustic model further comprises: calculating a posteriori probability according to one or more initial values of the parameters; and rectifying the parameters according to the posteriori probability.

54. The method of claim any one of claims 48 to 53 wherein the number of iterations are set according to empirical values.

55. The method of any one of claims 48 to 54 wherein preset value is reached if the output result is the calculated posteriori probability at the last iteration.

56. The method of any one of claims 43 to 55 wherein the recognizing further comprises transforming the filtered audio information to text content.

57. The method of claim 56 wherein the recognizing further comprises recognizing the filtered audio information by word-segmentation technology.

58. The method of any one of claims 43 to 57 wherein the recognizing further comprises outputting an evaluating indicator for appraising an exactness of a voice recognizing process.

59. The method of claim 58 wherein the evaluating indicator is one or more of a sentence error rate, sentence correct rate, word error rate and character error rate.

60. The method of any one of claims 45 to 59 wherein the pre-collecting of the prestored audio information further comprises: sending a plurality of operation instructions to the target user based on a received request from the target user; and receiving the prestored audio information based the target user response to the operation instructions.

61. The method of claim 60 wherein the operation instructions are sent based on a set time interval.

62. The method of claim 61 wherein the set time interval is 2 seconds.

63. The method of any one of claims 60 to 62 wherein the received request is a request to reset the equipment.