Response method and device based on deep learning, equipment and medium
By using deep learning technology to convert speech into text and analyze intent in real time, the problems of insufficient context changes and intent recognition in response technology are solved, and the safe and efficient implementation of intelligent voice services in highly sensitive areas is achieved.
Patent Information
- Application Number
- CN202510713235.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-12
AI Technical Summary
Existing response technologies have difficulty processing contextual changes in natural speech in real time and are unable to effectively identify user intent, resulting in insufficient service continuity and accuracy, which may have serious consequences, especially in highly sensitive areas such as healthcare and finance.
Using a deep learning-based method, the speech recognition model converts voice information into text, combines the natural language model to parse potential intentions and contextual features, and uses the intent recognition model to dynamically analyze changes in user intentions to generate professional and compliant response suggestions.
It improves the response speed and interaction quality of financial services, reduces business risks, enhances the communication efficiency and situational adaptability of medical services, and improves user experience and security.
Smart Images

Figure CN120632026A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a deep learning-based response method, device, equipment, and medium. Background Art
[0002] With the continuous advancement of digital and intelligent healthcare services, users are placing higher demands on service response speed and semantic understanding capabilities in remote medical consultations, health consultations, and home-based elderly care scenarios. For example, applications such as remote medical consultations, medication guidance, and emergency response require the ability to capture changes in the patient's or elderly user's language in real time to accurately understand their health status, needs, and emotional fluctuations. Similarly, in the financial services sector, banks, securities firms, and insurance companies are increasingly relying on intelligent voice systems to handle high-frequency interactions such as customer service calls, remote financial consulting, and financial product recommendations. These financial services often involve high-value transactions, complex product explanations, and regulatory compliance requirements, placing higher demands on system response speed, semantic understanding accuracy, and context retention. Traditional systems, in particular, often provide irrelevant answers, delayed responses, or even communication breakdowns when customer intent changes suddenly, semantics jump, or ambiguity arises, severely impacting customer experience and business conversion.
[0003] Current voice response technologies generally face the following challenges: first, it is difficult to process contextual changes in natural speech in real time; second, the ability to dynamically recognize user intent is insufficient; and third, it is unable to effectively perform semantic correction and dialogue connection, which in turn affects service continuity and accuracy. In highly sensitive and high-response fields such as healthcare, elderly care, and finance, service failures can have serious consequences, placing more stringent requirements on the stability, accuracy, compliance, and intelligent responsiveness of intelligent voice systems. Therefore, an answering method with real-time speech processing capabilities, a dynamic intent recognition mechanism, and contextual correction functions is urgently needed to improve the quality of intelligent services in real-time financial call scenarios. Summary of the Invention
[0004] The present invention provides a deep learning-based response method, device, computer equipment and medium to solve the technical problems of poor prediction timeliness and low prediction accuracy of cash flow data in related technologies.
[0005] In a first aspect, a deep learning-based response method is provided, comprising:
[0006] Obtain the user's current voice information and historical conversation information, and analyze the current voice information through a speech recognition model to obtain corresponding text information;
[0007] Parsing the text information using a natural language model to obtain a parsing result corresponding to the text information, wherein the parsing result includes potential intent information and context feature information;
[0008] Analyzing the context feature information and the historical conversation information through an intent recognition model to obtain current intent information and a response prediction result;
[0009] If the current intention information is consistent with the potential intention information, a response suggestion is output based on the response prediction result.
[0010] In a second aspect, a deep learning-based response device is provided, comprising:
[0011] An acquisition module is used to acquire the user's current voice information and historical conversation information, and analyze the current voice information through a speech recognition model to obtain corresponding text information;
[0012] A parsing module, configured to parse the text information using a natural language model to obtain a parsing result corresponding to the text information, wherein the parsing result includes potential intent information and context feature information;
[0013] An analysis module is used to analyze the context feature information and the historical conversation information through an intent recognition model to obtain current intent information and a response prediction result;
[0014] A response module is used to output a response suggestion based on the response prediction result if the current intention information is consistent with the potential intention information.
[0015] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned deep learning-based response method when executing the computer program.
[0016] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the above-mentioned deep learning-based answering method. In the solution implemented by the above-mentioned deep learning-based answering method, device, computer equipment, and storage medium, the user's current voice information and historical conversation information can be obtained, and the current voice information can be analyzed by a speech recognition model to obtain corresponding text information. Furthermore, the text information can be parsed by a natural language model to obtain a corresponding parsing result of the text information, wherein the parsing result includes potential intent information and contextual feature information. Thus, the contextual feature information and historical conversation information can be analyzed by the intent recognition model to obtain current intent information and answer prediction results. Finally, if the current intent information is consistent with the potential intent information, an answer suggestion is output based on the answer prediction result. In the present invention, the user's speech is transcribed in real time by a speech recognition model, and the potential intent and contextual information in the speech are analyzed in combination with natural language processing technology, so that the user's needs can be accurately understood. This method is suitable for semantically complex and terminology-intensive conversation environments in financial and medical scenarios. Furthermore, by combining historical conversation records, the intention recognition model is used to dynamically judge the changes in the user's current intention and adjust the response strategy in real time to ensure the continuity and accuracy of the response. Figure 1 When a customer is prompted, professional and compliant response suggestions are generated based on the response prediction results. This not only significantly improves the response speed and interaction quality in financial services, reducing business risks and compliance issues caused by misunderstandings, but also enhances communication efficiency and situational adaptability in medical services, improving the service experience and safety of patients and elderly users, and enabling the safe and efficient implementation of intelligent voice services in highly sensitive industries. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0018] Figure 1 1 is a schematic diagram of an application environment of a deep learning-based response method according to an embodiment of the present invention;
[0019] Figure 2 This is a flowchart of a deep learning-based response method according to an embodiment of the present invention;
[0020] Figure 3 yes Figure 1 A schematic flow chart of a specific implementation of step S10;
[0021] Figure 4 1 is a schematic structural diagram of a response device based on deep learning in one embodiment of the present invention;
[0022] Figure 5 is a structural diagram of a computer device in one embodiment of the present invention;
[0023] Figure 6 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0024] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0025] The deep learning-based response method provided by the embodiment of the present invention can be applied in Figure 1 In an application environment, the client communicates with the server through a network. The server can obtain the user's current voice information and historical conversation information through the client, and analyze the current voice information through a speech recognition model to obtain corresponding text information; parse the text information through a natural language model to obtain a parsing result corresponding to the text information, wherein the parsing result includes potential intention information and context feature information; analyze the context feature information and the historical conversation information through an intention recognition model to obtain current intention information and a response prediction result; if the current intention information is consistent with the potential intention information, then output a response suggestion based on the response prediction result, and finally feed the response suggestion back to the client. In the present invention, the user's voice is transcribed in real time by a speech recognition model, and the potential intention and context information in the speech are parsed in combination with natural language processing technology, so as to accurately understand the user's needs. This method is suitable for conversation environments with complex semantics and dense professional terms in financial and medical scenarios. Furthermore, combined with historical conversation records, the intention recognition model is used to dynamically judge the changes in the user's current intention, and the response strategy is adjusted in real time to ensure the continuity and accuracy of the response. When the current intention is consistent with the potential intention identified by the system, the user's voice is transcribed in real time by a speech recognition model, and the potential intention and context information in the speech are analyzed ... the user's needs can be accurately understood. Figure 1When a response is received, professional and compliant reply suggestions are generated based on the answer prediction results. It not only significantly improves the response speed and interaction quality in financial services, reduces business risks and compliance issues caused by misunderstandings, but also enhances communication efficiency and situational adaptability in medical services, improves the service experience and safety of patients or elderly users, and realizes the safe and efficient implementation of intelligent voice services in highly sensitive industries. Among them, the client can be but is not limited to various personal computers, laptops, smart phones, tablets and portable wearable devices. The server can be implemented with an independent server or a server cluster composed of multiple servers. The present invention is described in detail below through specific embodiments.
[0026] See also Figure 2 As shown, Figure 2 A flowchart of a deep learning-based response method provided in an embodiment of the present invention includes the following steps:
[0027] S10: Acquire the user's current voice information and historical conversation information, and analyze the current voice information through a voice recognition model to obtain corresponding text information.
[0028] For example, in step S10, the user's current voice information can be first obtained, that is, the voice content input by the user in real time through a terminal device (such as a phone, smartphone, or voice customer service system or telemedicine platform). At the same time, the user's historical conversation information in the system is retrieved, including past call records, text interaction content, or processed business requests, or historical consultation records, health consultation content, medication recommendations, and monitoring data related to the user in medical scenarios. By fusing current voice input with historical interaction data, more complete contextual support is provided for subsequent semantic understanding and intent recognition, which helps to provide more accurate, continuous, and professionally qualified intelligent response services in complex application environments such as finance and healthcare.
[0029] It should be noted that voice information currently typically exists in the form of an audio data stream, and therefore this voice data is analyzed and processed by a speech recognition model. Speech recognition models can be based on deep neural networks (such as end-to-end speech recognition networks or long short-term memory networks) and are capable of processing multilingual, accented, and unstructured speech expressions. The speech recognition model first extracts features from the speech signal, such as Mel-Frequency Cepstral Coefficients (MFCCs) or spectrogram data. It then uses a language model to decode the continuous semantics of the speech, accurately transcribing the unstructured speech data into structured text. This text information not only preserves the original content of the user's expression but also provides the foundational input for subsequent natural language understanding and intent recognition. Using a high-precision speech recognition model can significantly improve the system's ability to understand user input and the accuracy of its responses. Incorporating historical conversation information can also enhance contextual understanding, further improving the coherence and personalization of intelligent responses.
[0030] Among them, such as Figure 3 As shown, in step S10, that is, analyzing the current voice information by the voice recognition model to obtain corresponding text information, the following steps are included:
[0031] S11: performing a preprocessing operation on the current voice information to obtain preprocessed voice information.
[0032] The preprocessing operation includes a noise filtering operation and a speech rate normalization processing operation.
[0033] S12: Performing vector feature extraction on the preprocessed speech information to obtain a corresponding speech feature vector.
[0034] S13: Analyze the speech feature vector using the speech recognition model to obtain the text information.
[0035] In steps S11-S13, the current voice information input by the user may first be pre-processed to improve the accuracy and robustness of subsequent voice recognition. The pre-processing operation mainly includes two steps: noise filtering and speech rate normalization.
[0036] Noise filtering removes background noise, telecommunications network interference, or noise generated by terminal devices from user speech. Filter algorithms and adaptive noise reduction algorithms (such as spectral subtraction and Wiener filtering) are used to clean the original speech signal, making the target speech clearer. Speech rate normalization uses dynamic time warping (DTW) or frame alignment to unify speech signal differences caused by speech speed and intonation fluctuations among different users into a speech rate range that the model can stably process, thereby reducing recognition bias caused by differences in speaking rhythm. The preprocessed speech data is more consistent and resolvable, laying a solid foundation for subsequent feature extraction and recognition.
[0037] Furthermore, speech feature vectors can be extracted from the preprocessed speech data. This process is typically based on frame-level processing techniques, which divide the continuous speech signal into frames and extract feature information from each frame, such as Mel-Frequency Cepstral Coefficients (MFCCs), Mel-spectrum, and Perceptual Linear Prediction Parameters (PLPs), to form a high-dimensional vector reflecting the speech content. These speech feature vectors retain the core information of the speech while filtering out individual differences in the speaker. The extracted speech feature vectors can then be input into a trained speech recognition model, which can be based on a deep neural network architecture such as a recurrent neural network (RNN), a long short-term memory network (LSTM), or an end-to-end speech recognition network (such as a Transformer or CTC). The model analyzes and decodes the speech sequence, ultimately outputting corresponding structured text information. This text information fully preserves the user's semantic expression, providing a reliable data foundation for subsequent natural language understanding, intent recognition, and intelligent response.
[0038] S20: Parse the text information using a natural language model to obtain a parsing result corresponding to the text information.
[0039] The analysis results include potential intent information and context feature information.
[0040] For example, the text information obtained by speech recognition can be subjected to in-depth semantic analysis through a natural language model, with the goal of understanding the true intention behind the user's speech and its contextual position in the current conversation.
[0041] It should be noted that natural language models are typically built based on deep learning structures, such as Bidirectional Encoder Representation (BERT), Transformer networks, or specially trained domain language models. These models have the ability to model complex sentence structures, word meaning ambiguity, and contextual dependencies. During the parsing process, the text is first subjected to basic processing such as word segmentation, part-of-speech tagging, and syntactic structure analysis. Then, semantic embedding technology is used to convert it into a high-dimensional semantic vector to capture the key expressions contained therein. Through the contextual attention mechanism, the natural language model can not only identify the topic and tone of the current sentence, but also combine historical conversation information to determine whether the user is repeating, supplementing, or negating previous content, thereby outputting accurate contextual feature information.
[0042] At the same time, the natural language model can also identify the user's underlying operational intent, or latent intent information, based on semantic clues and keyword structure. The latent intent information and contextual features in the parsed results serve as the core basis for the system's subsequent intent recognition and response generation. This ensures that the system can understand not only the "surface" of the words but also the "user's intended meaning," enabling more intelligent and context-sensitive automated or collaborative responses.
[0043] In some embodiments, parsing the text information through a natural language model to obtain a parsing result corresponding to the text information includes: extracting key information from the text information, wherein the key information includes keyword information and syntactic structure; identifying the keyword information through the natural language model to obtain a parsing result corresponding to the text information.
[0044] For example, to achieve a deep understanding of user text messages, not only can natural language models be used to parse the entire sentence, but key information can also be pre-extracted to enhance the relevance and accuracy of semantic processing. Specifically, the text message is first structured and analyzed to extract keyword information and syntactic structure. Keyword information typically includes verbs, noun phrases, and domain terms that represent the user's core intent, such as "reimbursement," "registration," "balance inquiry," and "risk assessment," as well as "registration," "reimbursement," "medication advice," and "symptom consultation" in the medical field. Syntactic structure refers to the grammatical relationships between words, such as subject-verb-object structures, modification relationships, and clause dependencies. This structural information helps understand the overall organization of semantics and identify semantic focus and contextual dependencies. In medical contexts, it helps distinguish between the user's main complaint, medical history, and consultation purpose, thereby improving understanding of complex health issues and providing more accurate and professional intelligent responses in both financial and medical scenarios.
[0045] Furthermore, after the extraction is completed, the keyword information can be used as the input focus to be sent to the natural language model for further semantic recognition. The natural language model uses context modeling capabilities, combined with the position and usage of keywords in the sentence, to judge its semantic category, contextual intention, tone state, etc., and thus output a complete parsing result. The parsing result not only includes an abstract expression of the user's current speech content, such as the intention of "complaint", "consultation", and "appointment", but also includes the logical connection with the previous conversation, such as whether it is a supplementary explanation or a change of intention. Through this "extract first, then parse" processing mechanism, the understanding ability in multi-round conversations, long text expressions, and fuzzy request scenarios has been significantly improved, ensuring the provision of stable and accurate semantic basic support in application scenarios with strong professionalism and changing intentions such as the financial field.
[0046] S30: Analyze the context feature information and the historical conversation information through the intention recognition model to obtain current intention information and response prediction results.
[0047] In some embodiments, the analyzing the context feature information and the historical conversation information through the intention recognition model to obtain current intention information and a response prediction result includes: analyzing the context feature information and the historical conversation information through the intention recognition model to obtain the current intention information; performing semantic matching on the current intention information and the past intention information to obtain a matching result; and determining the response prediction result based on the matching result.
[0048] For example, in order to more accurately understand the user's actual needs in the current session, the intent recognition model can be used to jointly analyze context feature information and historical conversation information to output current intent information and response prediction results.
[0049] It should be noted that the intent recognition model is built based on a deep semantic understanding algorithm, which comprehensively considers both explicit semantic features (such as keywords, emotional expressions, and intonation) and implicit contextual information (such as conversation history, logical connections, and user behavior trajectories) in the current text. The intent recognition model first encodes the contextual feature information in the current round of conversation and historical conversation data. It then combines time series modeling techniques (such as recurrent neural networks or attention mechanisms) to analyze the semantic evolution trajectory, thereby outputting the current user's core intent information, such as "apply for a claim," "book a follow-up appointment," "check risk level," etc.
[0050] After obtaining the current intent information, it can be further semantically matched with the user's intent information in the historical session to determine whether the current expression continues, repeats, changes or transfers the original intent. The matching process can be performed by using semantic similarity calculation, intent label comparison, etc. Figure 1If the matching result shows a new intention or a significant deviation, it will automatically switch to the new task processing path and output the new predicted response content.
[0051] This mechanism, which integrates current judgments with historical semantic matching, greatly enhances the consistency and accuracy of responses in multi-round conversations, task jumps, and complex business scenarios, improving user experience and service efficiency.
[0052] S40: If the current intention information is consistent with the potential intention information, a response suggestion is output according to the response prediction result.
[0053] In some embodiments, if the current intention information is consistent with the potential intention information, a response suggestion is output based on the response prediction result, including: if the current intention information is consistent with the potential intention information, a corresponding corpus segment is retrieved based on the response prediction result; the corpus segment is adjusted according to the context feature information to obtain an adjusted natural language text; and the adjusted natural language text is determined as the response suggestion.
[0054] For example, in order to ensure that the generated response content is consistent with the user's current intention and has a natural and fluent expression effect, when it is determined that the current intent information is consistent with the potential intent information, the preset corpus fragments will be further retrieved based on the response prediction results. These corpus fragments come from professional knowledge bases or historical high-quality conversation records, and usually correspond one-to-one to different intent categories, such as "account inquiry", "drug use", "claims process", etc. Therefore, after locating the most matching corpus fragment based on the response prediction results, it will not be output directly, but will be personalized based on the current context feature information (such as user identity, tone, recent conversation content), including replacing the name, supplementing time and place information, adjusting the tone and voice, etc., so as to generate more natural and actual context-appropriate response content. Finally, the adjusted natural language text is determined as the response suggestion output by the system to the user, ensuring that while accurately matching the intent, higher conversation fluency and user satisfaction are achieved.
[0055] In some embodiments, the method also includes: calculating the difference vector between the current intention information and the potential intention information, and determining whether the difference vector exceeds a preset threshold; if it exceeds the preset threshold, determining that the current intention information is inconsistent with the potential intention information; if the current intention information is inconsistent with the potential intention information, determining a secondary response suggestion based on the current intention information.
[0056] For example, in order to improve the sensitivity and response capabilities to changes in user intentions, after obtaining the current intention information and potential intention information, the two are first mapped into high-dimensional semantic vectors through vectorized representation technology, and the difference vector between them is calculated. The difference vector reflects the extent of change in user expression in the semantic space. Subsequently, the value of the difference vector is compared with the preset threshold. If the difference exceeds the threshold, it means that the user's intention has changed significantly, and it is determined that the current intention is inconsistent with the originally predicted potential intention. In this case, the original response prediction result is no longer used. Instead, the response content is regenerated based on the current intent information, and targeted secondary response suggestions are output to ensure that it can flexibly respond to the shift in user needs, improve interaction coherence and service accuracy, and is especially suitable for scenarios with multiple rounds of complex questions and answers and rapid switching of intentions.
[0057] In some embodiments, the method further includes: obtaining historical voice information and dividing the historical voice information into a training set and a validation set; inputting the training set into a deep learning model, performing model fitting using a supervised learning method, and obtaining the speech recognition model.
[0058] For example, to obtain a trained speech recognition model, the historical speech information is first divided into a training set and a validation set. The data contained in the training set is used to learn and fit the model parameters, while the validation set is used to monitor model performance during training and prevent overfitting. The training set is then input into the deep learning model, and the model is trained using supervised learning. This involves using known labels as the target and minimizing the error between the predicted output and the true value. After multiple rounds of training and validation, a capable speech recognition model is obtained.
[0059] It can be seen that in the above scheme, the user's voice can be transcribed in real time through the speech recognition model, and the potential intention and context information in the speech can be analyzed by combining natural language processing technology to accurately understand the user's needs. This method is suitable for conversation environments with complex semantics and dense professional terms in financial and medical scenarios. Furthermore, combined with historical conversation records, the intention recognition model is used to dynamically judge the changes in the user's current intention, and the response strategy is adjusted in real time to ensure the continuity and accuracy of the response. When the current intention is consistent with the potential intention recognized by the system, Figure 1 When a customer is prompted, professional and compliant response suggestions are generated based on the response prediction results. This not only significantly improves the response speed and interaction quality in financial services, reducing business risks and compliance issues caused by misunderstandings, but also enhances communication efficiency and situational adaptability in medical services, improving the service experience and safety of patients and elderly users, and enabling the safe and efficient implementation of intelligent voice services in highly sensitive industries.
[0060] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0061] In one embodiment, a response device based on deep learning is provided, and the response device based on deep learning corresponds one-to-one with the response method based on deep learning in the above embodiment. Figure 4 As shown, the deep learning-based response device includes an acquisition module 101, a parsing module 102, an analysis module 103, and a response module 104. The functional modules are described in detail as follows:
[0062] Acquisition module 101, for acquiring the user's current voice information and historical conversation information, and analyzing the current voice information through a speech recognition model to obtain corresponding text information;
[0063] The parsing module 102 is configured to parse the text information using a natural language model to obtain a parsing result corresponding to the text information, wherein the parsing result includes potential intent information and context feature information;
[0064] An analysis module 103 is configured to analyze the context feature information and the historical conversation information using an intent recognition model to obtain current intent information and a response prediction result;
[0065] The response module 104 is configured to output a response suggestion based on the response prediction result if the current intention information is consistent with the potential intention information.
[0066] In one embodiment, the acquisition module 101 is specifically configured to:
[0067] Performing a preprocessing operation on the current voice information to obtain preprocessed voice information; wherein the preprocessing operation includes a noise filtering operation and a speech rate normalization operation;
[0068] Performing vector feature extraction on the preprocessed speech information to obtain a corresponding speech feature vector;
[0069] The speech feature vector is analyzed by the speech recognition model to obtain the text information.
[0070] In one embodiment, the parsing module 102 is specifically configured to:
[0071] Extracting key information from the text information, wherein the key information includes keyword information and syntactic structure;
[0072] The keyword information is identified through the natural language model to obtain a parsing result corresponding to the text information.
[0073] In one embodiment, the analysis module 103 is further configured to:
[0074] Analyzing the context feature information and the historical conversation information through the intention recognition model to obtain the current intention information;
[0075] Performing semantic matching on the current intention information and the past intention information to obtain a matching result;
[0076] The response prediction result is determined according to the matching result.
[0077] In one embodiment, the response module 104 is specifically configured to:
[0078] If the current intention information is consistent with the potential intention information, then searching for the corresponding corpus segment according to the response prediction result;
[0079] Adjusting the corpus segment according to the context feature information to obtain an adjusted natural language text;
[0080] The adjusted natural language text is determined as the answer suggestion.
[0081] In one embodiment, the response module 104 is specifically configured to:
[0082] Calculating a difference vector between the current intention information and the potential intention information, and determining whether the difference vector exceeds a preset threshold;
[0083] If the preset threshold is exceeded, determining that the current intention information is inconsistent with the potential intention information;
[0084] If the current intention information is inconsistent with the potential intention information, a secondary response suggestion is determined based on the current intention information.
[0085] In one embodiment, the acquisition module 101 is specifically configured to:
[0086] Acquire historical speech information, and divide the historical speech information into a training set and a validation set;
[0087] The training set is input into a deep learning model, and the model is fitted using a supervised learning method to obtain the speech recognition model.
[0088] The present invention provides a response device based on deep learning, which can transcribe the user's speech in real time through a speech recognition model, and analyze the potential intention and context information in the speech by combining natural language processing technology, so as to accurately understand the user's needs. This method is suitable for dialogue environments with complex semantics and dense professional terms in financial and medical scenarios. Furthermore, combined with historical dialogue records, the intention recognition model is used to dynamically judge the changes in the user's current intention, and the response strategy is adjusted in real time to ensure the continuity and accuracy of the response. When the current intention is consistent with the potential intention recognized by the system, the user can Figure 1 When a customer is prompted, professional and compliant response suggestions are generated based on the response prediction results. This not only significantly improves the response speed and interaction quality in financial services, reducing business risks and compliance issues caused by misunderstandings, but also enhances communication efficiency and situational adaptability in medical services, improving the service experience and safety of patients and elderly users, and enabling the safe and efficient implementation of intelligent voice services in highly sensitive industries.
[0089] For the specific definition of the deep learning-based response device, please refer to the definition of the deep learning-based response method above and will not be repeated here. Each module in the above-mentioned deep learning-based response device can be implemented in whole or in part through software, hardware, or a combination thereof. Each of the above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each of the above modules.
[0090] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, memory, network interface and database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the server side of a deep learning-based response method.
[0091] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 6As shown. The computer device includes a processor, memory, network interface, display screen, and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the client side of a deep learning-based response method.
[0092] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:
[0093] Obtain the user's current voice information and historical conversation information, and analyze the current voice information through a speech recognition model to obtain corresponding text information;
[0094] Parsing the text information using a natural language model to obtain a parsing result corresponding to the text information, wherein the parsing result includes potential intent information and context feature information;
[0095] Analyzing the context feature information and the historical conversation information through an intent recognition model to obtain current intent information and a response prediction result;
[0096] If the current intention information is consistent with the potential intention information, a response suggestion is output based on the response prediction result.
[0097] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0098] Obtain the user's current voice information and historical conversation information, and analyze the current voice information through a speech recognition model to obtain corresponding text information;
[0099] Parsing the text information using a natural language model to obtain a parsing result corresponding to the text information, wherein the parsing result includes potential intent information and context feature information;
[0100] Analyzing the context feature information and the historical conversation information through an intent recognition model to obtain current intent information and a response prediction result;
[0101] If the current intention information is consistent with the potential intention information, a response suggestion is output based on the response prediction result.
[0102] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0103] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchl ink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0104] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0105] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A response method based on deep learning, characterized in that: The method comprises: Obtain the user's current voice information and historical conversation information, and analyze the current voice information through a speech recognition model to obtain corresponding text information; Parsing the text information using a natural language model to obtain a parsing result corresponding to the text information, wherein the parsing result includes potential intent information and context feature information; Analyzing the context feature information and the historical conversation information through an intent recognition model to obtain current intent information and a response prediction result; If the current intention information is consistent with the potential intention information, a response suggestion is output based on the response prediction result.
2. The method according to claim 1, characterized in that The current voice information is analyzed by the voice recognition model to obtain corresponding text information, including: Performing a preprocessing operation on the current voice information to obtain preprocessed voice information; wherein the preprocessing operation includes a noise filtering operation and a speech rate normalization operation; Performing vector feature extraction on the preprocessed speech information to obtain a corresponding speech feature vector; The speech feature vector is analyzed by the speech recognition model to obtain the text information.
3. The method according to claim 1, characterized in that The parsing of the text information by a natural language model to obtain a parsing result corresponding to the text information includes: Extracting key information from the text information, wherein the key information includes keyword information and syntactic structure; The keyword information is identified through the natural language model to obtain a parsing result corresponding to the text information.
4. The method according to claim 1, wherein The analysis of the context feature information and the historical conversation information by the intent recognition model to obtain current intent information and a response prediction result includes: Analyzing the context feature information and the historical conversation information through the intention recognition model to obtain the current intention information; Performing semantic matching on the current intention information and the past intention information to obtain a matching result; The response prediction result is determined according to the matching result.
5. The method according to claim 1, characterized in that If the current intention information is consistent with the potential intention information, outputting a response suggestion based on the response prediction result includes: If the current intention information is consistent with the potential intention information, then searching for the corresponding corpus segment according to the response prediction result; Adjusting the corpus segment according to the context feature information to obtain an adjusted natural language text; The adjusted natural language text is determined as the answer suggestion.
6. The method according to claim 1, characterized in that The method further comprises: Calculating a difference vector between the current intention information and the potential intention information, and determining whether the difference vector exceeds a preset threshold; If the preset threshold is exceeded, determining that the current intention information is inconsistent with the potential intention information; If the current intention information is inconsistent with the potential intention information, a secondary response suggestion is determined based on the current intention information.
7. The method according to claim 1, characterized in that The method further comprises: Acquire historical speech information, and divide the historical speech information into a training set and a validation set; The training set is input into a deep learning model, and the model is fitted using a supervised learning method to obtain the speech recognition model.
8. A response device based on deep learning, characterized in that: include: An acquisition module is used to acquire the user's current voice information and historical conversation information, and analyze the current voice information through a speech recognition model to obtain corresponding text information; A parsing module, configured to parse the text information using a natural language model to obtain a parsing result corresponding to the text information, wherein the parsing result includes potential intent information and context feature information; An analysis module is used to analyze the context feature information and the historical conversation information through an intent recognition model to obtain current intent information and a response prediction result; A response module is used to output a response suggestion based on the response prediction result if the current intention information is consistent with the potential intention information.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the deep learning-based response method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the deep learning-based response method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Medical dialogue search term rewriting method and device, equipment and storage medium
CN120994798A