Voice conversation device based on deep learning
By adopting deep learning-based speech recognition and noise reduction technology in voice conversation devices, the problem that existing devices cannot accurately recognize and obtain clean speech is solved, and high-accuracy automatic response and acquisition of clean speech signals are achieved.
Patent Information
- Application Number
- CN202510146064.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing voice conversation devices cannot accurately recognize voice, resulting in reduced accuracy of automatic responses and difficulty in obtaining clean noise-reduced voice.
The voice dialogue device based on deep learning is adopted, including a receiving module, an average computing module, a voice noise reduction module, a voice recognition module, a voice synthesis module, a repetitive generation module and a storage module. The target speech is decoded through the deep learning model, extract multi-resolution cochlear map features for noise reduction processing, obtain decoded text and extract text intentions, realize online response and speech conversion.
It significantly improves the accuracy of speech recognition, realizes automatic response, and obtains a clean voice signal through more accurate noise reduction processing.
Smart Images

Figure CN119993144A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of voice dialogue technology, and in particular to a voice dialogue device based on deep learning. Background Art
[0002] With the development of intelligent voice, intelligent devices equipped with intelligent voice have gradually entered the daily life of users, providing intelligent voice services to users in all aspects of life. These intelligent devices equipped with intelligent voice perform speech recognition on the user's voice. In order to reduce the time consumption of the dialogue system, semantic recognition is usually performed in advance. When the intermediate results of speech recognition are obtained, the complete expression of the user is predicted, and the reply information is generated in advance based on the predicted complete expression, so that the reply information can be output immediately when the reply conditions are met, for example, after the user has finished speaking a paragraph.
[0003] However, existing voice dialogue devices cannot accurately recognize speech, and thus cannot provide text responses and convert speech through answering strategies, which reduces the accuracy of automatic answers; at the same time, clean speech after noise reduction is obtained. Therefore, we propose a voice dialogue device based on deep learning to solve the above-mentioned problems. Summary of the invention
[0004] The purpose of the present invention is to provide a voice dialogue device based on deep learning to solve the problems raised in the above background technology.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A voice dialogue device based on deep learning, comprising a receiving module, an average calculation module, a voice noise reduction module, a voice recognition module, a voice synthesis module, a repetitive generation module, and a storage module; the receiving module is connected to the average calculation module, the average calculation module is connected to the voice noise reduction module, the voice noise reduction module is connected to the voice recognition module, the voice recognition module is connected to the repetitive generation module, and the repetitive generation module is connected to the storage module;
[0007] A receiving module, used for the voice dialogue device to receive voice;
[0008] an average calculation module, for calculating an average of speech sentence lengths based on past speech of the user stored in the storage module, each speech sentence length indicating the length of the speech of the user;
[0009] The speech noise reduction module is used to perform noise reduction on the received speech and obtain a clean speech file after noise reduction;
[0010] The speech recognition module uses a deep learning model to decode the target speech and obtain the decoded text;
[0011] a speech synthesis module for combining a dependent word that has established a dependency relationship with a noun included in the acquired user's speech with the noun to generate a plurality of answer sentence candidates according to the text intention;
[0012] a repetitive generation module for selecting a response sentence candidate from the plurality of response sentence candidates generated by the candidate generation module in association with the average value of the speech sentence length calculated by the average calculation module, and using the selected response sentence candidate as is or processing the selected response sentence candidate to generate a parrot-like response sentence;
[0013] The storage module is used to store the user's past voice.
[0014] As a further solution of the present invention: the speech noise reduction module includes a file receiving unit, a feature extraction unit, a first calculation unit, a second calculation unit, and a noise reduction processing unit. The file receiving unit is connected to the feature extraction unit, the feature extraction unit is connected to the first calculation unit, the first calculation unit is connected to the second calculation unit, and the second calculation unit is connected to the noise reduction processing unit.
[0015] As a further solution of the present invention: the file receiving unit is used to receive the noisy speech file to be processed; the feature extraction unit is used to extract the multi-resolution cochlear map features in the noisy speech file; the first calculation unit is used to obtain the prior signal-to-noise ratio estimation value of the noisy speech file based on the multi-resolution cochlear map features and the network model; the second calculation unit is used to obtain the gain function based on the prior signal-to-noise ratio estimation value; the noise reduction processing unit is used to perform noise reduction processing on the noisy speech file based on the gain function to obtain a clean speech file after noise reduction.
[0016] As a further solution of the present invention: the method for using the speech denoising module includes receiving a noisy speech file to be processed; extracting multi-resolution cochlear map features in the noisy speech file; obtaining a priori signal-to-noise ratio estimation value of the noisy speech file based on the multi-resolution cochlear map features and the network model; obtaining a gain function based on the priori signal-to-noise ratio estimation value; and performing denoising on the noisy speech file based on the gain function to obtain a clean speech file after denoising.
[0017] As a further solution of the present invention: the method of obtaining a priori signal-to-noise ratio estimation value of a noisy speech file based on multi-resolution cochlear map features and a network model includes: obtaining a mapping of the priori signal-to-noise ratio based on multi-resolution cochlear map features; loading the mapping of the priori signal-to-noise ratio into the network model to obtain a priori signal-to-noise ratio estimation value.
[0018] As a further solution of the present invention: the mapping of the prior signal-to-noise ratio based on the multi-resolution cochlear map features includes: obtaining the instantaneous prior signal-to-noise ratio according to the multi-resolution cochlear map features; mapping the instantaneous prior signal-to-noise ratio to obtain a mapping of the prior signal-to-noise ratio, and the mapping method adopted is the cumulative distribution function.
[0019] As a further solution of the present invention: the speech recognition module also includes an intention recognition unit and an idiom understanding unit, the intention recognition unit is connected to the idiom understanding unit, the intention recognition unit decodes the target speech through a deep learning model, and then extracts text intent from the decoded text; the idiom understanding unit is used to detect idioms in the original text of the decoded text in the intention recognition unit to obtain idioms and idiom explanations of the idioms.
[0020] As a further solution of the present invention: the idiom understanding unit includes a detection subunit, an expansion subunit, an analysis subunit, a determination subunit, and an association subunit, the detection subunit is connected to the expansion subunit, the expansion subunit is connected to the analysis subunit, the analysis subunit is connected to the determination subunit, and the determination subunit is connected to the association subunit.
[0021] As a further solution of the present invention: the detection subunit is configured to obtain an original text containing dialogue content, perform idiom detection on the original text, and obtain idioms and idiom explanations of the idioms; the expansion subunit is configured to expand each word in the idiom to obtain a number of candidate sentences, perform similarity matching between the idiom explanations and the candidate sentences, and determine the idiom sentence corresponding to the idiom based on the similarity matching results; the analysis subunit is configured to analyze the idiom sentence to obtain the subject of the idiom sentence, and use the subject of the idiom sentence as the idiom role corresponding to the robot; the determination subunit is configured to use the sentence containing the idiom in the original text as the target sentence, and use the subject or object in the target sentence as the dialogue role of the robot; the association subunit is configured to associate the idiom role of the robot with the dialogue role, and generate dialogue information based on the role association result.
[0022] As a further solution of the present invention: the speech recognition module uses a deep learning model to decode the target speech, and the steps of obtaining the decoded text include: extracting high-dimensional speech features of the target speech; using an acoustic model to convert the high-dimensional speech features into acoustic model scores; using a deep learning model to maintain the acoustic model score sequence. Decoding the sequence so that the weighted sum of the acoustic model score and the language model score is maximized, thereby obtaining the decoded text.
[0023] Compared with the prior art, the present invention has the following beneficial effects:
[0024] 1. The deep learning-based voice dialogue device can decode the target voice by using the deep learning model to obtain the decoded text; extract the text intent from the decoded text; use the answering strategy to answer online and give the corresponding text answer according to the text intent; convert the text answer into a voice signal. Compared with the prior art, the present invention introduces neural network technology into the voice recognition problem, greatly improves the recognition accuracy, and uses the answering strategy to answer the text and convert the voice to realize automatic answering;
[0025] 2. The deep learning-based voice dialogue device can obtain more feature information by setting up a voice noise reduction module and extracting multi-resolution cochlear map features, so that a more accurate prior signal-to-noise ratio estimate can be obtained based on more feature information, and then a cleaner clean voice after noise reduction can be obtained. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 It is a structural schematic diagram of the voice dialogue device based on deep learning of the present invention.
[0027] Figure 2 It is a structural schematic diagram of the idiom understanding unit in the present invention. DETAILED DESCRIPTION
[0028] In one embodiment, Figure 1-Figure 2 As shown, a voice dialogue device based on deep learning includes a receiving module, an average calculation module, a voice noise reduction module, a voice recognition module, a voice synthesis module, a repetitive generation module, and a storage module; the receiving module is connected to the average calculation module, the average calculation module is connected to the voice noise reduction module, the voice noise reduction module is connected to the voice recognition module, the voice recognition module is connected to the repetitive generation module, and the repetitive generation module is connected to the storage module;
[0029] A receiving module, used for the voice dialogue device to receive voice;
[0030] an average calculation module, for calculating an average of speech sentence lengths based on past speech of the user stored in the storage module, each speech sentence length indicating the length of the speech of the user;
[0031] The speech noise reduction module is used to perform noise reduction on the received speech and obtain a clean speech file after noise reduction;
[0032] The speech recognition module uses a deep learning model to decode the target speech and obtain the decoded text;
[0033] a speech synthesis module for combining a dependent word that has established a dependency relationship with a noun included in the acquired user's speech with the noun to generate a plurality of answer sentence candidates according to the text intention;
[0034] a repetitive generation module for selecting a response sentence candidate from the plurality of response sentence candidates generated by the candidate generation module in association with the average value of the speech sentence length calculated by the average calculation module, and using the selected response sentence candidate as is or processing the selected response sentence candidate to generate a parrot-like response sentence;
[0035] A storage module, used to store the user's past voice;
[0036] The speech noise reduction module includes a file receiving unit, a feature extraction unit, a first calculation unit, a second calculation unit, and a noise reduction processing unit. The file receiving unit is connected to the feature extraction unit, the feature extraction unit is connected to the first calculation unit, the first calculation unit is connected to the second calculation unit, and the second calculation unit is connected to the noise reduction processing unit.
[0037] A file receiving unit, used for receiving a noisy speech file to be processed; a feature extraction unit, used for extracting multi-resolution cochlear spectrogram features in the noisy speech file; a first calculation unit, used for obtaining a priori signal-to-noise ratio estimation value of the noisy speech file based on the multi-resolution cochlear spectrogram features and the network model; a second calculation unit, used for obtaining a gain function based on the priori signal-to-noise ratio estimation value; a noise reduction processing unit, used for performing noise reduction processing on the noisy speech file based on the gain function to obtain a clean speech file after noise reduction;
[0038] The method for using the speech denoising module includes receiving a noisy speech file to be processed; extracting multi-resolution cochlear spectral features from the noisy speech file; obtaining a priori signal-to-noise ratio estimation value of the noisy speech file based on the multi-resolution cochlear spectral features and the network model; obtaining a gain function based on the priori signal-to-noise ratio estimation value; performing denoising on the noisy speech file based on the gain function to obtain a clean speech file after denoising;
[0039] Obtaining a priori signal-to-noise ratio estimation value of a noisy speech file based on multi-resolution cochlear spectral features and a network model includes: obtaining a mapping of the priori signal-to-noise ratio based on the multi-resolution cochlear spectral features; loading the mapping of the priori signal-to-noise ratio into the network model to obtain a priori signal-to-noise ratio estimation value;
[0040] The mapping of the prior signal-to-noise ratio based on the multi-resolution cochlear map features includes: obtaining the instantaneous prior signal-to-noise ratio according to the multi-resolution cochlear map features; mapping the instantaneous prior signal-to-noise ratio to obtain the mapping of the prior signal-to-noise ratio, and the mapping method adopted is the cumulative distribution function;
[0041] The speech recognition module also includes an intention recognition unit and an idiom understanding unit. The intention recognition unit is connected to the idiom understanding unit. The intention recognition unit decodes the target speech through a deep learning model and then extracts the text intention from the decoded text. The idiom understanding unit is used to detect the original text memory idioms of the decoded text in the intention recognition unit to obtain the idioms and the idiom explanations of the idioms.
[0042] The idiom understanding unit includes a detection subunit, an expansion subunit, an analysis subunit, a determination subunit, and a correlation subunit. The detection subunit is connected to the expansion subunit, the expansion subunit is connected to the analysis subunit, the analysis subunit is connected to the determination subunit, and the determination subunit is connected to the correlation subunit.
[0043] The detection subunit is configured to obtain the original text containing the dialogue content, perform idiom detection on the original text, and obtain the idiom and the idiom explanation of the idiom; the expansion subunit is configured to expand each word in the idiom to obtain a number of candidate sentences, perform similarity matching between the idiom explanation and the candidate sentences, and determine the idiom sentence corresponding to the idiom according to the similarity matching result; the analysis subunit is configured to analyze the idiom sentence to obtain the subject of the idiom sentence, and use the subject of the idiom sentence as the idiom role corresponding to the robot; the determination subunit is configured to use the sentence containing the idiom in the original text as the target sentence, and use the subject or object in the target sentence as the dialogue role of the robot; the association subunit is configured to associate the idiom role of the robot with the dialogue role, and generate dialogue information according to the role association result;
[0044] By obtaining the original text containing the content of the dialogue, the original text is detected for idioms to obtain the idioms and the idiom explanations of the idioms; each word in the idiom is expanded to obtain several candidate sentences, the idiom explanation is matched with each candidate sentence for similarity, and the idiom sentence corresponding to the idiom is determined based on the similarity matching result; the idiom sentence is analyzed to obtain the subject of the idiom sentence, and the subject of the idiom sentence is used as the idiom role corresponding to the robot; the sentence containing the idiom in the original text is used as the target sentence, and the subject or object in the target sentence is used as the dialogue role of the robot; the idiom role of the robot is associated with the dialogue role, and the dialogue information is generated based on the role association result. This application can realize human-computer dialogue in the scene containing idiom sentences, improve the application scope of voice dialogue, and enhance the intelligence level of voice dialogue.
[0045] The speech recognition module uses a deep learning model to decode the target speech, and the steps of obtaining the decoded text include: extracting high-dimensional speech features from the target speech; using an acoustic model to convert the high-dimensional speech features into acoustic model scores; using a deep learning model to perform maintenance ratio decoding on the acoustic model score sequence so that the weighted sum of the acoustic model score and the language model score is maximized, thereby obtaining the decoded text;
[0046] The present invention can decode the target speech by using a deep learning model to obtain a decoded text; extract the text intent from the decoded text; use an answering strategy to answer online and give a corresponding text answer based on the text intent; convert the text answer into a speech signal. Compared with the prior art, the present invention introduces neural network technology into the speech recognition problem, greatly improving the recognition accuracy, and uses an answering strategy to answer text and convert speech to achieve automatic answering; at the same time, by setting a speech denoising module, multi-resolution cochlear map features can be extracted to obtain more feature information, so that a more accurate prior signal-to-noise ratio estimate can be obtained based on more feature information, and then a cleaner clean speech after denoising can be obtained.
[0047] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A voice dialogue device based on deep learning, characterized in that: It includes a receiving module, an average calculation module, a speech noise reduction module, a speech recognition module, a speech synthesis module, a repetitive generation module, and a storage module; the receiving module is connected to the average calculation module, the average calculation module is connected to the speech noise reduction module, the speech noise reduction module is connected to the speech recognition module, the speech recognition module is connected to the repetitive generation module, and the repetitive generation module is connected to the storage module; A receiving module, used for the voice dialogue device to receive voice; an average calculation module, for calculating an average of speech sentence lengths based on past speech of the user stored in the storage module, each speech sentence length indicating the length of the speech of the user; The speech noise reduction module is used to perform noise reduction on the received speech and obtain a clean speech file after noise reduction; The speech recognition module uses a deep learning model to decode the target speech and obtain the decoded text; a speech synthesis module for combining a dependent word that has established a dependency relationship with a noun included in the acquired user's speech with the noun to generate a plurality of answer sentence candidates according to the text intention; a repetitive generation module for selecting a response sentence candidate from the plurality of response sentence candidates generated by the candidate generation module in association with the average value of the speech sentence length calculated by the average calculation module, and using the selected response sentence candidate as is or processing the selected response sentence candidate to generate a parrot-like response sentence; The storage module is used to store the user's past voice.
2. A voice dialogue device based on deep learning according to claim 1, characterized in that: The speech noise reduction module includes a file receiving unit, a feature extraction unit, a first calculation unit, a second calculation unit, and a noise reduction processing unit. The file receiving unit is connected to the feature extraction unit, the feature extraction unit is connected to the first calculation unit, the first calculation unit is connected to the second calculation unit, and the second calculation unit is connected to the noise reduction processing unit.
3. A voice dialogue device based on deep learning according to claim 2, characterized in that: The file receiving unit is used to receive a noisy speech file to be processed; the feature extraction unit is used to extract multi-resolution cochlear spectrogram features in the noisy speech file; the first calculation unit is used to obtain a priori signal-to-noise ratio estimation value of the noisy speech file based on the multi-resolution cochlear spectrogram features and the network model; The second calculation unit is used to obtain a gain function based on a priori signal-to-noise ratio estimation value; the noise reduction processing unit is used to perform noise reduction processing on the noisy speech file based on the gain function to obtain a clean speech file after noise reduction.
4. The deep learning-based voice dialogue device according to claim 3, characterized in that: The method for using the speech denoising module includes receiving a noisy speech file to be processed; extracting multi-resolution cochlear map features in the noisy speech file; obtaining a priori signal-to-noise ratio estimation value of the noisy speech file based on the multi-resolution cochlear map features and a network model; obtaining a gain function based on the priori signal-to-noise ratio estimation value; and performing denoising on the noisy speech file based on the gain function to obtain a clean speech file after denoising.
5. The deep learning-based voice dialogue device according to claim 4, characterized in that: The method of obtaining a priori signal-to-noise ratio estimation value of a noisy speech file based on multi-resolution cochlear spectral features and a network model includes: obtaining a mapping of the priori signal-to-noise ratio based on the multi-resolution cochlear spectral features; and loading the mapping of the priori signal-to-noise ratio into the network model to obtain a priori signal-to-noise ratio estimation value.
6. The deep learning-based voice dialogue device according to claim 5, characterized in that: The mapping of the prior signal-to-noise ratio based on the multi-resolution cochlear map features includes: obtaining the instantaneous prior signal-to-noise ratio according to the multi-resolution cochlear map features; mapping the instantaneous prior signal-to-noise ratio to obtain the mapping of the prior signal-to-noise ratio, and the mapping method adopted is the cumulative distribution function.
7. The deep learning-based voice dialogue device according to claim 1, characterized in that: The speech recognition module also includes an intention recognition unit and an idiom understanding unit. The intention recognition unit is connected to the idiom understanding unit. The intention recognition unit decodes the target speech through a deep learning model and then extracts text intent from the decoded text. The idiom understanding unit is used to detect idioms in the original text of the decoded text in the intention recognition unit to obtain idioms and idiom explanations of the idioms.
8. The deep learning-based voice dialogue device according to claim 7, characterized in that: The idiom understanding unit includes a detection subunit, an expansion subunit, an analysis subunit, a determination subunit, and an association subunit. The detection subunit is connected to the expansion subunit, the expansion subunit is connected to the analysis subunit, the analysis subunit is connected to the determination subunit, and the determination subunit is connected to the association subunit.
9. The deep learning-based voice dialogue device according to claim 8, characterized in that: The detection subunit is configured to obtain an original text containing the dialogue content, perform idiom detection on the original text, and obtain the idiom and the idiom explanation of the idiom; the expansion subunit is configured to expand each word in the idiom to obtain a plurality of candidate sentences, perform similarity matching between the idiom explanation and the candidate sentences, and determine the idiom sentence corresponding to the idiom according to the similarity matching result; The analyzing subunit is configured to analyze the idiom sentence to obtain the subject of the idiom sentence, and use the subject of the idiom sentence as the idiom role corresponding to the robot; the determining subunit is configured to use the sentence containing the idiom in the original text as the target sentence, and use the subject or object in the target sentence as the dialogue role of the robot; The association subunit is configured to associate the robot's idiom role with the dialogue role and generate dialogue information according to the role association result.
10. The deep learning-based voice dialogue device according to claim 1, characterized in that: The speech recognition module uses a deep learning model to decode the target speech, and the steps of obtaining the decoded text include: extracting high-dimensional speech features from the target speech; using an acoustic model to convert the high-dimensional speech features into acoustic model scores; using the deep learning model to perform maintenance ratio decoding on the acoustic model score sequence so that the weighted sum of the acoustic model score and the language model score is maximized, thereby obtaining the decoded text.