Voice consultation device, voice consultation method, and storage medium
By combining a microphone module, acoustic echo cancellation circuit, and camera, along with beamforming algorithms and recognition mode switching, the problem of inaccurate voice recognition in noisy environments has been solved, improving the accuracy of voice recognition and the human-computer interaction effect.
Patent Information
- Application Number
- CN202310064900.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-17
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2043-01-17
AI Technical Summary
Voice consultation devices may not accurately recognize voice data in noisy environments, affecting their effectiveness.
By employing a combination of microphone module, acoustic echo cancellation circuit, camera and processor, noise is suppressed and audio data clarity is improved through sealing layer, beamforming algorithm and recognition mode switching.
It improves the accuracy of speech recognition and response, and enhances the human-computer interaction effect.
Smart Images

Figure CN116132878B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to a voice consultation device, a voice consultation method and a storage medium, and belongs to the technical field of computers. BACKGROUND
[0002] At present, with the development of artificial intelligence, self-service terminal devices are widely used. The self-service terminal device is generally composed of a man-machine interface, and a user can operate it independently according to the prompt of the device without the participation of staff, thereby realizing business handling.
[0003] A typical self-service terminal device is also configured with a voice consultation function, and a voice consultation device is obtained. At this time, a microphone is installed on the voice consultation device, the microphone collects voice data issued by a user, identifies the semantics corresponding to the voice data, and performs an operation corresponding to the semantics.
[0004] However, the application environment of the voice consultation device is usually noisy, for example, in a government office, there may be broadcast sound, people talking sound and the like, at this time, the voice consultation device may be inaccurate in identifying voice data, and the use effect of the voice consultation device may be affected. SUMMARY
[0005] The application provides a voice consultation device, a voice consultation method and a storage medium, which can solve the problems of inaccurate voice recognition and poor use effect of the traditional voice consultation device. The application provides the following technical solutions:
[0006] In a first aspect, a voice consultation device is provided, which comprises:
[0007] A microphone module comprising a PCB, a plurality of microphones installed on the PCB, and a module panel located above each microphone; a sealing layer corresponding to each microphone is arranged between the PCB and the module panel, and the sealing layer is adapted to seal the microphone; the plurality of microphones form a first collection channel and a second collection channel; the first collection channel is adapted to collect recording data, and the second collection channel is adapted to collect back collection data;
[0008] An acoustic echo cancellation (AEC) circuit connected to the microphone module, which is adapted to use the recording data and the back collection data to perform echo cancellation of a loudspeaker, and obtain to-be-recognized audio data;
[0009] A camera adapted to collect image data in front of the voice consultation device;
[0010] A processor connected to the AEC circuit and the camera, wherein the processor is configured to:
[0011] Obtain the image data collected by the camera and the to-be-recognized audio data.
[0012] determine a human-machine distance and a human-machine angle between the face and the voice consultation device when the image data indicates that the face exists;
[0013] perform noise suppression on the to-be-recognized audio data based on the human-machine angle using a beamforming algorithm to obtain processed audio data;
[0014] determine an identification mode based on the human-machine distance, the identification mode including a near-field identification mode and a far-field identification mode;
[0015] perform audio identification on the processed audio data based on the identification mode to obtain an identification result;
[0016] execute a business handling action corresponding to the business handling instruction based on the identification result being the business handling instruction, and return a business handling result;
[0017] when the identification result is the consultation task, determine a query result corresponding to the consultation task based on a knowledge base of a preset field, the preset field matching voice consultation services provided by the voice consultation device.
[0018] Optionally, the processor is further configured to:
[0019] compare the identification result with instruction parameters of a preconfigured business handling instruction;
[0020] when the identification result has a matching instruction parameter, determine that the identification result is the business handling instruction corresponding to the instruction parameter;
[0021] when the identification result does not have a matching instruction parameter, determine that the identification result is the consultation task.
[0022] Optionally, when the identification result is the business handling instruction, the business handling action corresponding to the business handling instruction is executed, and a business handling result is returned, including:
[0023] the business handling instruction is sent to a business handling page of the voice consultation device in a command manner to trigger the business handling page to handle a business and obtain the business handling result;
[0024] or,
[0025] the identification result is sent to a preset service, and the preset service sends the identification result to a business handling page in a command manner to trigger the business handling page to handle a business and obtain the business handling result.
[0026] Optionally, in the case that the identification result is a consultation task, the query result corresponding to the consultation task is determined based on a knowledge base of a preset field, and the method comprises the following steps of:
[0027] The query result corresponding to the consultation task is determined in the knowledge base based on a heuristic dialogue manner.
[0028] Optionally, the noise of the to-be-identified audio data is suppressed based on the human-machine angle using a beamforming algorithm to obtain processed audio data, and the method comprises the following steps of:
[0029] The voice signal outside the human-machine angle is suppressed and the multi-channel voice signal inside the human-machine angle is integrated into a single-channel audio signal using the beamforming algorithm to obtain the processed audio data.
[0030] Optionally, the thickness of the sealing layer corresponding to each microphone is equal; and in a direction parallel to the PCB, the distance between the edge of the sealing layer and the edge of the microphone is greater than a preset distance.
[0031] Optionally, the AEC circuit comprises a right-channel cancellation circuit and a left-channel cancellation circuit, and each channel cancellation circuit comprises an audio positive input end, an audio negative input end, an audio positive output end and an audio negative output end; the audio positive input end is connected to the audio positive output end through a first resistor, and the first resistor and the audio positive output end are grounded through a second resistor; the audio negative input end is connected to the audio negative output end through a third resistor, and the third resistor and the audio negative output end are grounded through a fourth resistor.
[0032] Optionally, the processor is further configured to:
[0033] generate a virtual customer service image based on AIGC technology;
[0034] fuse the business handling result or the query result with the virtual customer service image to obtain a fused customer service image;
[0035] output the fused customer service image.
[0036] In a second aspect, a voice consultation method is provided, and the method comprises the following steps of:
[0037] obtaining image data collected by a camera in a voice consultation device and to-be-identified audio data output by an AEC circuit in the voice consultation device;
[0038] in the case that the image data indicates that a human face exists, determining a human-machine distance and a human-machine angle between the human face and the voice consultation device;
[0039] perform noise suppression on the to-be-identified audio data based on the human-machine angle using a beamforming algorithm, to obtain processed audio data;
[0040] determine an identification mode based on the human-machine distance, the identification mode including a near-field identification mode and a far-field identification mode;
[0041] perform audio identification on the processed audio data based on the identification mode, to obtain an identification result;
[0042] in a case where the identification result is a service handling instruction, perform a service handling action corresponding to the service handling instruction, and return a service handling result;
[0043] in a case where the identification result is a consultation task, determine a query result corresponding to the consultation task based on a knowledge base of a preset field, the preset field matching a voice consultation service provided by the voice consultation device.
[0044] In a third aspect, a computer-readable storage medium is provided, the storage medium storing a program, the program being executed by a processor to implement the voice consultation method of the first aspect.
[0045] The beneficial effects of the present application include at least the following: the voice of the voice consultation device itself is separated by the AEC circuit and is not sent to the voice recognition; the clarity of the audio collected by the voice consultation device can be improved from the source; at the same time, the angle at which the user faces the microphone is analyzed by the beamforming algorithm, so that the voice signals outside the angle are suppressed, and the multi-channel human voice signals inside the angle are integrated into single-channel audio first, which is sent to the voice recognition, which can further improve the clarity of the audio data and the accuracy of the voice recognition. In addition, the human-machine distance between the user and the voice consultation device is determined by the camera, and the identification mode corresponding to the distance is automatically switched to, which can improve the accuracy of the voice recognition and the response accuracy of the voice consultation device.
[0046] In addition, the combination of heuristic dialogue and knowledge base can make the voice consultation device more intelligent and improve the human-computer interaction effect.
[0047] The above description is only a summary of the technical solutions of the present application. In order to more clearly understand the technical means of the present application and can be implemented according to the content of the description, the following will be described in detail with the preferred embodiments of the present application and with the help of the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS
[0048] Figure 1 is a structural schematic diagram of a voice consultation device provided by an embodiment of the present application;
[0049] Figure 2 is a structural schematic diagram of a microphone module provided by an embodiment of the present application;
[0050] Figure 3 is a schematic diagram of a collection channel of a microphone module provided by an embodiment of the present application;
[0051] Figure 4 is a structural schematic diagram of an AEC circuit provided by an embodiment of the present application;
[0052] Figure 5 is a flowchart of a voice consultation method provided by an embodiment of the present application;
[0053] Figure 6 is a flowchart of a voice consultation method provided by an embodiment of the present application applied to the medical insurance field. DETAILED DESCRIPTION
[0054] The specific embodiments of the present application will be further described in detail below with reference to the accompanying drawings and embodiments. The following embodiments are used to illustrate the present application, but not to limit the scope of the present application.
[0055] Figure 1 is a structural schematic diagram of a voice consultation device provided by an embodiment of the present application. The voice consultation device is used to automatically provide services in a preset field for a user. For example, services in the medical insurance field, services in the government field, or services in the financial field, etc. The present embodiment does not limit the preset field to which the voice consultation device is applicable. The voice consultation device includes a microphone module 110, an acoustic echo cancellation (AEC) circuit 120, a camera 130, a processor 140, and a memory 150.
[0056] The microphone module 110 is used to collect audio data in the current scene of the voice consultation device. In order to suppress noise in the audio data, with reference to Figure 2 In the present embodiment, the microphone module 110 includes a PCB board 21, a plurality of microphones 22 mounted on the PCB board, and a module panel 23 located above each microphone. A sealing layer 24 corresponding to each microphone is arranged between the PCB board and the module panel, and the sealing layer is adapted to seal the microphone.
[0057] The sealing layer is made of silica gel material, which can improve the air tightness and dust resistance of the microphone. At the same time, the sealing layer can also avoid the problem of resonance of the microphone and inconsistency of the channels between the microphones.
[0058] According to Figure 2It can be known that the thicknesses of the sealing layers corresponding to the respective microphones are equal; and in the direction parallel to the PCB, the distance between the edge of the sealing layer and the edge of the microphone is greater than a preset distance. The microphone can be a Microelectro Mechanical Systems (MEMS) microphone, and correspondingly, the preset distance can be 3 millimeters (mm).
[0059] The plurality of microphones form a first collection channel and a second collection channel; the first collection channel is adapted to collect recording data, and the second collection channel is adapted to collect back collection data. For example, referring to Figure 3 , the microphone module 110 includes 8 collection channels, of which channels 1-6 are first collection channels to collect recording data, and channels 7 and 8 are second collection channels to collect back collection data.
[0060] Illustratively, the microphone module 110 uses a standard USB Audio to collect audio data, and the sampling rate of each channel is 16 KHZ, and the sampling bit is 16 bits.
[0061] The AEC circuit 120 is connected to the microphone module to obtain the recording data and the back collection data collected by the microphone module 110, and is adapted to use the recording data and the back collection data to perform echo cancellation on the loudspeaker to obtain the to-be-recognized audio data.
[0062] In this embodiment, the AEC circuit is adapted to separate the sound of the loudspeaker of the voice consultation device itself from the audio data collected by the microphone module, and does not perform audio processing, so that part of the noise can be eliminated from the source to improve the audio processing effect.
[0063] Referring to the AEC circuit shown in Figure 4 , the AEC circuit includes a right channel cancellation circuit 41 and a left channel cancellation circuit 42, and each channel cancellation circuit includes an audio positive input end SPKL+ and SPKR+, an audio negative input end SPKL- and SPKR-, an audio positive output end LINE_L_P and LINE_R_P, and an audio negative output end LINE_L_N and LINE_R_N; the audio positive input end is connected to the audio positive output end through a first resistor R1, and the first resistor and the audio positive output end are grounded (analog ground AGND) through a second resistor R2; the audio negative input end is connected to the audio negative output end through a third resistor R3, and the third resistor and the audio negative output end are grounded AGND through a fourth resistor R4.
[0064] V LINE_L_P = V SPKL+ (R2 / (R1+R2));
[0065] V LINE_L_N = V SPKL- (R4 / (R3+R4));
[0066] V LINE_R_P = V SPKR+ (R2 / (R1+R2));
[0067] V LINE_R_N = V SPKR- (R4 / (R3+R4))。
[0068] The camera 130 is adapted to collect image data in front of the voice consultation device. The camera 130 can implement collection of image data; or, the camera 130 can also collect image data when a living body is detected to be close. For the latter implementation, the voice consultation device can further be provided with an infrared sensor, which can detect whether there is a living body within a certain range in front of the voice consultation device. In other embodiments, the camera 130 can also collect image data in a case where the voice consultation device receives an interactive operation through a human-computer interaction interface, and the present embodiment does not limit the timing of image data collection by the camera 130.
[0069] The processor 140 connected with the AEC circuit and the camera is used for processing the output of the camera 130 and the AEC circuit 120. The processor 140 is further connected with the memory 150.
[0070] The processor 140 can include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 140 can be implemented in at least one of a hardware form of a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), and a PLA (Programmable Logic Array). The processor 140 can also include a main processor and a coprocessor. The main processor is a processor for processing data in a wake-up state, also known as a CPU (Central Processing Unit). The coprocessor is a low-power processor for processing data in a standby state. In some embodiments, the processor 140 can be integrated with a GPU (Graphics Processing Unit) that is responsible for rendering and drawing the content required to be displayed by the display screen. In some embodiments, the processor 140 can further include an AI (Artificial Intelligence) processor that is used for processing computing operations related to machine learning.
[0071] The memory 150 can include one or more computer-readable storage media that can be non-transitory. The memory 150 can also include high-speed random access memory and nonvolatile, computer-readable storage media such as one or more disk storage devices, flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 150 is used to store at least one instruction for being executed by the processor 140 to implement the following steps:
[0072] Obtaining image data collected by a camera and audio data to be recognized;
[0073] In a case where the image data indicates that a human face exists, determining a human-machine distance and a human-machine angle between the human face and the voice consultation device;
[0074] Using a beamforming algorithm to perform noise suppression on the audio data to be recognized based on the human-machine angle, to obtain processed audio data;
[0075] Determining a recognition mode based on the human-machine distance, the recognition mode including a near-field recognition mode and a far-field recognition mode;
[0076] Performing audio recognition on the processed audio data based on the recognition mode, to obtain a recognition result;
[0077] In a case where the recognition result is a service handling instruction, performing a service handling action corresponding to the service handling instruction, and returning a service handling result;
[0078] In a case where the recognition result is a consultation task, determining a query result corresponding to the consultation task based on a knowledge base of a preset field, the preset field matching a voice consultation service provided by the voice consultation device.
[0079] In some embodiments, the voice consultation device can also optionally include a peripheral device interface and at least one peripheral device. The processor 140, the memory 150, and the peripheral device interface can be connected through a bus or a signal line. Each peripheral device can be connected to the peripheral device interface through a bus, a signal line, or a circuit board. Illustratively, the peripheral devices include but are not limited to radio frequency circuitry, a touch display screen, audio circuitry, a loudspeaker, and a power supply, etc.
[0080] Of course, the voice consultation device can also include fewer or more components, which are not limited in the present embodiment.
[0081] Using a beamforming algorithm to perform noise suppression on the audio data to be recognized based on the human-machine angle, to obtain processed audio data, including: using the beamforming algorithm to suppress voice signals outside the human-machine angle, and integrating multi-channel voice signals within the human-machine angle into a single-channel audio signal, to obtain the processed audio data.
[0082] The recognition mode is determined based on the human-machine distance, including: determining whether the human-machine distance is less than a preset distance; if the human-machine distance is less than the preset distance, determining that the recognition mode is a near-field recognition mode; and if the human-machine distance is greater than or equal to the preset distance, determining that the recognition mode is a far-field recognition mode.
[0083] The preset distance is pre-stored in the voice consultation device, and can be 1 meter or other numerical values, and the embodiment does not limit the value of the preset distance.
[0084] The processed audio data is subjected to audio recognition based on the recognition mode to obtain a recognition result, including: in the case that the recognition mode is the near-field recognition mode, using an automatic speech recognition (ASR) algorithm corresponding to the near-field recognition mode to perform audio recognition on the processed audio data, performing natural language processing (NLP) on the recognized data, and obtaining the recognition result; and in the case that the recognition mode is the far-field recognition mode, using an ASR algorithm corresponding to the far-field recognition mode to perform audio recognition on the processed audio data, performing NLP on the recognized data, and obtaining the recognition result. The recognition result is returned in a JSON manner.
[0085] The ASR algorithm corresponding to the near-field recognition mode and the ASR algorithm corresponding to the far-field recognition mode are trained using different voice sample data, and the trained algorithm parameters are different. Specifically, the ASR algorithm corresponding to the near-field recognition mode is trained using language sample data collected within the preset distance, and the ASR algorithm corresponding to the far-field recognition mode is trained using language sample data collected outside the preset distance.
[0086] In the embodiment, through automatic switching of the near-field and far-field recognition modes, not only can the switching be achieved without feeling, but also the accuracy of voice recognition can be improved, and the use effect of the voice consultation device can be improved.
[0087] The recognition result obtained by the processor can have two needs, one for business handling and the other for business inquiry. Based on this, the processor also needs to determine the demand corresponding to the recognition result. Specifically, the processor compares the recognition result with instruction parameters of a pre-configured business handling instruction; in the case that the recognition result has a matching instruction parameter, it is determined that the recognition result is the business handling instruction corresponding to the instruction parameter; and in the case that the recognition result does not have a matching instruction parameter, it is determined that the recognition result is a consultation task.
[0088] In one example, in the case where the recognition result is a service handling instruction, a service handling action corresponding to the service handling instruction is performed, and a service handling result is returned, including: issuing the service handling instruction to the service handling page of the voice consultation device in the form of a command command, to trigger the service handling page to handle the service, and obtain the service handling result.
[0089] Taking the service handling page as an example, the system page of the medical insurance bureau, the processor issues the command command of the command to the system page of the medical insurance bureau service handling through the protocol of the Chrome Inspector, to handle the service. At this time, the whole process can be completed without contact to complete the medical insurance center off-site record, birth reimbursement and other service handling. At the same time, the client directly interfaces with the service handling system can reduce the response delay.
[0090] In another example, in the case where the recognition result is a service handling instruction, a service handling action corresponding to the service handling instruction is performed, and a service handling result is returned, including: sending the recognition result to a preset service, issuing the service handling page to the service handling page in the form of a command command through the preset service, to trigger the service handling page to handle the service, and obtain the service handling result.
[0091] Chrome inspector is a protocol for web page and interface service interaction, which aims to issue the result of NLP semantic understanding to the page in the form of command, so that the page knows which specific business scenario to jump to handle the service. Based on this, the voice consultation device can also implement a preset service 160 on the server to receive the NLP recognition result, and then receive the recognition result through cloud-to-cloud, that is, service request service HTTP or HTTPS or WSS protocol, so that the preset service obtains the recognition result to complete the medical insurance service handling and page navigation task. The cloud-to-cloud method (or service-to-service method) can ensure the effect of the request link, which is more stable than the client directly interfacing.
[0092] In the case where the recognition result is a consultation task, the knowledge base of the preset field is used to determine the query result corresponding to the consultation task, including: determining the query result corresponding to the consultation task in the knowledge base based on the heuristic dialogue method.
[0093] Taking the preset field as the medical insurance field as an example, the knowledge base is the knowledge base of the medical insurance field. The processor generalizes the consultation task to obtain the consultation intent; determines the query result corresponding to the consultation intent and the topic to which the consultation intent belongs in the knowledge base; and recommends at least one heuristic question related to the topic to the user for selection.
[0094] Optionally, the processor is further configured to generate the virtual customer service image based on an artificial intelligence technology (AIGenerated Content, AIGC) technology, fuse the service handling result or the query result with the virtual customer service image to obtain a fused customer service image, and output the fused customer service image.
[0095] In the method, the fusing of the service handling result or the query result with the virtual customer service image comprises matching voice data corresponding to the service handling result or the query result with a mouth shape of the virtual customer service image, and the outputting of the fused customer service image comprises displaying the virtual customer service image and playing the voice data, the mouth shape of the virtual customer service image being synchronized with the voice data.
[0096] In summary, the voice consultation device provided in the embodiment separates the sound of the voice consultation device itself from the loudspeaker and does not send the sound to the voice recognition, so that the clarity of the audio collected by the voice consultation device can be improved from the source. Meanwhile, the angle at which the user faces the microphone is analyzed by using the beamforming algorithm, so that the voice signals outside the angle are suppressed, and the multi-channel voice signals within the angle are integrated into single-channel audio and sent to the voice recognition, so that the clarity of the audio data can be further improved, and the accuracy of the voice recognition is improved, thereby improving the response accuracy of the voice consultation device.
[0097] In addition, the combination of the heuristic dialogue and the knowledge base can make the voice consultation device more intelligent and improve the human-computer interaction effect.
[0098] Figure 5 is a flowchart of a voice consultation method provided in an embodiment of the present application. The embodiment takes the voice consultation method as an example for description in the voice consultation device shown in Figure 1 The method comprises at least the following steps:
[0099] Step 501: acquiring image data collected by a camera in the voice consultation device and audio data to be recognized output by an AEC circuit in the voice consultation device;
[0100] Step 502: in a case where the image data indicates that there is a face, determining a human-computer distance and a human-computer angle between the face and the voice consultation device;
[0101] Step 503: using a beamforming algorithm to perform noise suppression on the audio data to be recognized based on the human-computer angle to obtain processed audio data;
[0102] Step 504: determining a recognition mode based on the human-computer distance, the recognition mode comprising a near-field recognition mode and a far-field recognition mode;
[0103] At step 505, audio recognition is performed on the processed audio data based on the identified mode to obtain a recognition result, and step 506 or 507 is executed;
[0104] At step 506, if the recognition result is a service handling instruction, a service handling action corresponding to the service handling instruction is executed, and a service handling result is returned.
[0105] At step 507, if the recognition result is a consultation task, a query result corresponding to the consultation task is determined based on a knowledge base of a preset field, and the preset field matches the voice consultation service provided by the voice consultation device.
[0106] For related descriptions of this embodiment, refer to Figure 1 The embodiments shown in the drawings will not be described here.
[0107] Taking the medical insurance field as an example, refer to Figure 6 The method includes: after detecting a face through a camera, calculating the human-machine distance between the voice consultation device and the user; determining whether it is a near-field recognition mode based on the human-machine distance; if yes, using the near-field recognition mode for voice recognition; if not, using the far-field recognition mode for voice recognition. After NLP processing of the voice recognition result, if the recognition result is a command, the command is issued to a medical insurance service handling page, and a service handling result is fed back to the user; if the recognition result is a policy consultation demand, a policy consultation result corresponding to the consultation demand is queried in a medical insurance knowledge base, and fed back to the user.
[0108] In summary, the voice consultation method provided in this embodiment separates the sound of the voice consultation device's own loudspeaker through the AEC circuit, and does not send it to the voice recognition; the clarity of the audio collected by the voice consultation device can be improved from the source; at the same time, the angle of the user facing the microphone interaction is analyzed through the beamforming algorithm, so as to suppress the voice signal outside the angle, and the multi-channel human voice signal inside the angle is first integrated into a single-channel audio and sent to the voice recognition, which can further improve the clarity of the audio data and the accuracy of the voice recognition. In addition, the human-machine distance between the user and the voice consultation device is determined through the camera, and the recognition mode corresponding to the distance is automatically switched, which can improve the accuracy of the voice recognition and the response accuracy of the voice consultation device.
[0109] In addition, the combination of heuristic dialogue and knowledge base can make the voice consultation device more intelligent and improve the human-computer interaction effect.
[0110] Optionally, the present application also provides a computer readable storage medium, the computer readable storage medium stores a program, the program is loaded and executed by a processor to realize the voice consultation method of the above method embodiment.
[0111] Optionally, the present application also provides a computer product, the computer product includes a computer readable storage medium, the computer readable storage medium stores a program, the program is loaded and executed by a processor to realize the voice consultation method of the above method embodiment.
[0112] The technical features of the above embodiments can be combined in any manner. To make the description concise, all possible combinations of the technical features in the above embodiments are not described, but as long as the combinations of the technical features do not exist contradictions, they should be considered as the scope of the present application.
[0113] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the scope of the present application. Therefore, the scope of the patent of the present application should be subject to the appended claims.
Claims
1. A voice consultation device, characterized by, The voice consultation device comprises: A microphone module comprising a PCB, a plurality of microphones mounted on the PCB, and a module panel located above each microphone; a sealing layer corresponding to each microphone is arranged between the PCB and the module panel, and the sealing layer is adapted to seal the microphone; the plurality of microphones form a first acquisition channel and a second acquisition channel; the first acquisition channel is adapted to acquire recording data, and the second acquisition channel is adapted to acquire back acquisition data; An acoustic echo cancellation (AEC) circuit connected to the microphone module, adapted to use the recording data and the back acquisition data to perform echo cancellation on a loudspeaker to obtain to-be-recognized audio data; A camera adapted to acquire image data in front of the voice consultation device; A processor connected to the AEC circuit and the camera, wherein the processor is configured to: acquire the image data acquired by the camera and the to-be-recognized audio data; in a case where the image data indicates that a human face exists, determine a human-machine distance and a human-machine angle between the human face and the voice consultation device; use a beamforming algorithm to perform noise suppression on the to-be-recognized audio data based on the human-machine angle to obtain processed audio data; the use of the beamforming algorithm to perform noise suppression on the to-be-recognized audio data based on the human-machine angle to obtain the processed audio data comprises: using the beamforming algorithm to suppress voice signals outside the human-machine angle and integrate multi-channel voice signals within the human-machine angle into single-channel audio signals to obtain the processed audio data; determine an identification mode based on the human-machine distance, wherein the identification mode comprises a near-field identification mode and a far-field identification mode; perform audio identification on the processed audio data based on the identification mode to obtain an identification result; the performance of audio identification on the processed audio data based on the identification mode to obtain the identification result comprises: in a case where the identification mode is the near-field identification mode, using an ASR algorithm corresponding to the near-field identification mode to perform audio identification on the processed audio data, performing natural language processing on the identified data to obtain the identification result; in a case where the identification mode is the far-field identification mode, using an ASR algorithm corresponding to the far-field identification mode to perform audio identification on the processed audio data, performing NLP processing on the identified data to obtain the identification result; wherein the ASR algorithm corresponding to the near-field identification mode is trained using language sample data collected within a preset distance, and the ASR algorithm corresponding to the far-field identification mode is trained using language sample data collected outside the preset distance; in a case where the identification result is a business handling instruction, performing a business handling action corresponding to the business handling instruction and returning a business handling result; in a case where the identification result is a consultation task, determining a query result corresponding to the consultation task based on a knowledge base of a preset field, wherein the preset field matches a voice consultation service provided by the voice consultation device.
2. The voice consulting device according to claim 1, characterized by The processor is further configured to: compare the identification result with an instruction parameter of a preconfigured business handling instruction. In a case where the recognition result matches an instruction parameter, the recognition result is determined as a service handling instruction corresponding to the instruction parameter; In a case where the recognition result does not match an instruction parameter, the recognition result is determined as the consultation task.
3. The voice consulting device according to claim 1, characterized by, In a case where the recognition result is a service handling instruction, a service handling action corresponding to the service handling instruction is performed, and a service handling result is returned, including: The service handling instruction is sent to a service handling page of the voice consultation device in a command mode to trigger the service handling page to handle a service, and the service handling result is obtained; Or, The recognition result is sent to a preset service, and the service handling page is sent in a command mode through the preset service to trigger the service handling page to handle a service, and the service handling result is obtained.
4. The voice advisory device of claim 1, wherein, In a case where the recognition result is a consultation task, a query result corresponding to the consultation task is determined based on a knowledge base of a preset field, including: The query result corresponding to the consultation task is determined in the knowledge base in a heuristic dialogue mode.
5. The voice advisory device of claim 1, wherein, The thickness of the sealing layer corresponding to each microphone is equal; in a direction parallel to the PCB, the distance between the edge of the sealing layer and the edge of the microphone is greater than a preset distance.
6. The voice advisory device of claim 1, wherein, The AEC circuit includes a right channel cancellation circuit and a left channel cancellation circuit, and each channel cancellation circuit includes an audio positive input end, an audio negative input end, an audio positive output end and an audio negative output end; the audio positive input end is connected to the audio positive output end through a first resistor, and the first resistor and the audio positive output end are grounded through a second resistor; the audio negative input end is connected to the audio negative output end through a third resistor, and the third resistor and the audio negative output end are grounded through a fourth resistor.
7. The voice advisory device according to any one of claims 1 to 6, characterized in that, The processor is further configured to: generate a virtual customer service image based on AIGC technology; fuse the service handling result or the query result with the virtual customer service image to obtain a fused customer service image; output the fused customer service image.
8. A voice consultation method characterized by, The method comprises: obtaining image data collected by a camera in a voice consultation device and audio data to be recognized output by an AEC circuit in the voice consultation device; in a case where the image data indicates that there is a face, determining a human-machine distance and a human-machine angle between the face and the voice consultation device; using a beamforming algorithm to suppress noise in the audio data to be recognized based on the human-machine angle to obtain processed audio data; using the beamforming algorithm to suppress voice signals outside the human-machine angle and integrating multi-channel voice signals within the human-machine angle into a single-channel audio signal to obtain the processed audio data; determining an identification mode based on the human-machine distance, the identification mode including a near-field identification mode and a far-field identification mode; perform audio recognition on the processed audio data based on the recognition mode, to obtain a recognition result; perform audio recognition on the processed audio data based on the recognition mode, to obtain a recognition result, including: in a case where the recognition mode is a near-field recognition mode, performing audio recognition on the processed audio data using an ASR algorithm corresponding to the near-field recognition mode, performing natural language processing on the recognized data, and obtaining the recognition result; in a case where the recognition mode is a far-field recognition mode, performing audio recognition on the processed audio data using an ASR algorithm corresponding to the far-field recognition mode, performing NLP processing on the recognized data, and obtaining the recognition result; wherein the ASR algorithm corresponding to the near-field recognition mode is trained using language sample data collected within a preset distance, and the ASR algorithm corresponding to the far-field recognition mode is trained using language sample data collected outside the preset distance; in a case where the recognition result is a service handling instruction, performing a service handling action corresponding to the service handling instruction, and returning a service handling result; in a case where the recognition result is a consultation task, determining a query result corresponding to the consultation task based on a knowledge base of a preset field, the preset field matching a voice consultation service provided by the voice consultation device.
9. A computer-readable storage medium, characterized in that, The storage medium has a program stored therein, and the program, when executed by a processor, is configured to implement the voice consultation method of claim 8.
Citation Information
Patent Citations
Artificial intelligence digital signage
CN112908221A
Intelligent robot , power consumption service system
CN207630055U