Intelligent question answering method and device, computer device and storage medium
By using a pre-set question-answering model for speech recognition, semantic recognition, and style generation, the problem of poor semantic matching in existing intelligent question-answering technologies is solved, and continuous style response speech generation is achieved, improving accuracy and efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2024-10-30
- Publication Date
- 2026-05-01
AI Technical Summary
Existing intelligent question-answering technologies rely on semantic matching, resulting in poor answer quality, and the responses are presented in text form, lacking continuity in voice style.
The system uses a pre-defined question-and-answer model to perform speech recognition, semantic recognition, pronoun replacement, style recognition, and generation modules to generate response speech that matches the user's voice style, including speech synthesis using a speech conversion model.
It improves the style and accuracy and efficiency of reply voice and text, ensures that reply voice matches the user's style, reduces computing resources, and achieves continuity of voice content and style.
Smart Images

Figure CN119479644B_ABST
Abstract
Description
Intelligent question answering methods, devices, computer equipment and storage media Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to an intelligent question-answering method, apparatus, computer device, and storage medium. Background Technology
[0002] Currently, intelligent customer service or intelligent assistants are one application scenario of natural language processing technology. They are typically based on retrieval-based question-and-answer technology. This involves business experts pre-defining a question-and-answer knowledge base, a series of question-and-answer pairs, with each standard question corresponding to one answer. Then, semantic matching models are used to find synonyms for user questions and provide the corresponding answers. However, the effectiveness of this approach relies heavily on semantic matching and is usually presented in text form, leading to unsatisfactory results. Therefore, a technological solution is urgently needed to address these issues. Summary of the Invention
[0003] This invention provides an intelligent question-answering method, apparatus, computer device, and storage medium to solve the problem that the response process in the prior art relies on semantic matching and the response is presented in text form, resulting in poor answering effect.
[0004] An intelligent question-answering method includes:
[0005] The system acquires the user's voice and performs speech recognition on the user's voice using the speech recognition module in a preset question-and-answer model to obtain the recognized text.
[0006] The semantic recognition module in the preset question-and-answer model performs semantic recognition on the recognized text to obtain the semantic recognition result.
[0007] The semantic recognition results are used to replace pronouns in the recognized text to obtain text information;
[0008] The user's style is obtained by performing style recognition on the user's voice and text information through the style recognition module in the preset question-and-answer model.
[0009] The generation module in the preset question-and-answer model generates style text based on the user's voice, user style, and text information to obtain the response style and response text.
[0010] The reply style and reply text are synthesized using a speech conversion model to obtain a reply voice corresponding to the user's voice.
[0011] A smart question-answering device, comprising:
[0012] The speech conversion module is used to acquire the user's speech of the target user and perform speech recognition on the user's speech through the speech recognition module in the preset question-and-answer model to obtain the recognized text;
[0013] The semantic recognition module is used to perform semantic recognition on the recognized text through the semantic recognition module in the preset question-answering model, and obtain the semantic recognition result;
[0014] The pronoun replacement module is used to replace pronouns in the recognized text based on the semantic recognition result to obtain text information;
[0015] The style recognition module is used to perform style recognition on the user's voice and text information through the style recognition module in the preset question-and-answer model to obtain the user's style;
[0016] The information response module is used to generate style text from the user's voice, user style, and text information through the generation module in the preset question and answer model, so as to obtain the response style and response text.
[0017] The text conversion module is used to perform speech synthesis on the response style and the response text using a speech conversion model to obtain a response speech corresponding to the user's speech.
[0018] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described intelligent question-answering method.
[0019] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described intelligent question-answering method.
[0020] This invention provides an intelligent question-answering method, apparatus, device, and medium. The method utilizes a semantic recognition module within a pre-defined question-answering model to perform semantic recognition and pronoun replacement on the identified text, thereby achieving semantic recognition of the text and replacing pronouns in the identified text. This ensures the accuracy of speech content and style recognition, thus improving the accuracy and efficiency of the generated style and text. The style recognition module within the pre-defined question-answering model performs style recognition on user speech and text information, achieving user style identification. The generation module within the pre-defined question-answering model generates style-text based on user style and text information, achieving the generation of response style and response text, ensuring continuity in speech style, thereby generating response speech, reducing computational resources, and ensuring that the generated response speech conforms to the user's style. Attached Figure Description
[0021] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 is a schematic diagram of the application environment of the intelligent question-answering method in an embodiment of the present invention;
[0023] Figure 2 is a flowchart of an intelligent question-answering method according to an embodiment of the present invention;
[0024] Figure 3 is a flowchart of an intelligent question-answering method in another embodiment of the present invention;
[0025] Figure 4 is a schematic block diagram of an intelligent question-answering device according to an embodiment of the present invention;
[0026] Figure 5 is a schematic diagram of a computer device according to an embodiment of the present invention. Detailed Implementation
[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0028] The intelligent question-answering method provided in this invention can be applied to the application environment shown in Figure 1. Specifically, the intelligent question-answering method is applied in an intelligent question-answering device, which includes a client and a server as shown in Figure 1. The client and server communicate via a network to solve the problem that local image stylization in the prior art cannot achieve the expected results. The server can be a standalone server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The client, also known as the user terminal, refers to the program that provides categorization services to customers, corresponding to the server. The client can be installed on, but is not limited to, various computers, laptops, smartphones, tablets, and portable wearable devices.
[0029] In one embodiment, as shown in FIG2, an intelligent question-answering method is provided. Taking the application of this method to the server in FIG1 as an example, the method includes the following steps:
[0030] S10: Acquire the user's voice of the target user, and perform voice recognition on the user's voice through the voice recognition module in the preset question-and-answer model to obtain the recognized text.
[0031] In essence, user voice refers to pre-processed audio data of a user, such as the voice of a user asking a question or making a consultation. A pre-set question-and-answer model is a pre-configured model trained on a large amount of sample data. Recognized text refers to the text content corresponding to the user's voice. The speech recognition module is a model used to convert speech into text. The target user is the user who asked the question.
[0032] Specifically, the user's speech is acquired and input into a pre-defined question-and-answer model. The speech recognition module within this model performs speech recognition, which involves segmenting the speech into frames. This means dividing the speech into multiple frame units at fixed time intervals (e.g., 25 milliseconds). To avoid excessive variation between adjacent frame units, an overlap is allowed between them. Each frame unit is multiplied by a window function to ensure continuity between its left and right ends, resulting in a continuous time window. A short-time Fourier transform is performed on all windowed frame units to obtain the corresponding spectrum, i.e., the spectrum distributed across different time windows on the time axis. A Mel filter is then used to process the spectrum, yielding the Mel spectrum corresponding to the speech spectrogram, converting the linear natural spectrum into a Mel spectrum that reflects human auditory characteristics. The logarithm of the Mel spectrum is obtained to acquire its logarithmic energy. This logarithmic energy is then inversely transformed using the Discrete Cosine Transform (DCT), and the second to thirteenth coefficients after the DCT are taken as Mel frequency cepstral coefficients (MFCCs). These MFCCs are then defined as MFCC features. An acoustic model is used to perform pattern matching on the MFCC features, specifically by calculating the distance between the MFCC features and each pronunciation template to find the best-matching pronunciation template, thus obtaining the speech text. A language model is then used to perform error correction and detection on the speech text, further judging and correcting the recognition results based on grammatical structure and semantic rules, resulting in the recognized text. For example, in the financial field, a user asks a smart assistant through a microphone, "I want to buy car insurance, what are your recommendations?" The user's voice is captured, recognized, and the text content "I want to buy car insurance, what are your recommendations?" is obtained.
[0033] S20, the semantic recognition module in the preset question-and-answer model performs semantic recognition on the recognized text to obtain the semantic recognition result.
[0034] In essence, a semantic recognition module refers to a model used to identify the meaning represented by sentences in text. The semantic recognition result refers to the meaning represented by the content in the text as identified by the model.
[0035] Specifically, the semantic recognition module in the pre-set question-answering model performs semantic recognition on the identified text. That is, the identified text is input into the semantic recognition model, and the semantic recognition ability learned during training is used to perform semantic recognition on the identified text, thereby obtaining the semantics represented by the identified text, and the identified semantics is determined as the semantic recognition result.
[0036] In one specific embodiment, the identified text is vectorized, that is, the identified text is segmented into words, and the segmented words are vectorized. Positional encoding of each word is added during the vectorization process to increase recognition accuracy. Then, feature extraction is performed on the vectorized text through an extraction layer, that is, features are extracted from the vectorized text based on semantic relationships and syntactic structure. An attention mechanism is used to capture long-distance dependencies in the identified text. Finally, a fully connected layer performs semantic recognition on the extracted features and captured relationships to obtain the semantic recognition result. For example, in the financial field, in the question "I want to buy car insurance, what are some recommendations?", semantic recognition can determine that the user wants suitable car insurance recommendations.
[0037] S30, the semantic recognition result is used to replace the pronouns in the recognized text to obtain text information.
[0038] Understandably, textual information refers to the identified text after the pronouns have been replaced, which has coherent contextual information.
[0039] Specifically, the semantic recognition results are used to replace pronouns in the identified text. This involves identifying pronouns in the text using a pre-defined bag-of-words model, thus obtaining pronoun recognition results. The position of each pronoun in the text is then determined based on these results. Next, based on the context information corresponding to each pronoun and the semantic recognition results, the content that the pronoun refers to is replaced with the specific content it represents. In this way, by replacing all pronouns in the identified text based on the semantic recognition results, the text information can be obtained. For example, in a financial scenario, after recommending multiple insurance policies to a user, the user inquires about different policies. When inquiring about a third insurance policy, what are the differences between this policy and the previous ones? Here, "this policy" refers to the currently inquired insurance, while "previous" refers to the previously inquired insurance. After replacement, it is the current insurance policy; what are the differences between it and the previous insurance policy?
[0040] S40, the user's voice and text information are style-recognized by the style recognition module in the preset question-and-answer model to obtain the user's style.
[0041] Intuitively, user style refers to the style in which a user speaks. The style recognition module refers to the model used to identify speaking styles.
[0042] Specifically, the style recognition module within the pre-defined question-answering model performs style recognition on user speech and text information. Specifically, the speech style recognition unit within the style recognition module performs style recognition on user speech, extracting speech style features that reflect the user's style. Simultaneously, the text style recognition unit within the style recognition module performs style recognition on text information, extracting text style features that reflect the user's style. Then, the extracted speech and text style features are fused to obtain style features. Based on these style features, style recognition is performed on the user's speech, i.e., style prediction is performed on the style features to determine the user's style.
[0043] S50, the generation module in the preset question-and-answer model generates style text based on the user's voice, the user's style, and the text information to obtain the response style and response text.
[0044] In other words, response style refers to the style of responding to a user's voice message. Response text refers to the content of the response to a user's question or inquiry.
[0045] Specifically, the generation module in the pre-defined question-answering model performs style-based text generation on the user's speech, style, and text information. In other words, the text generation module generates text based on the user's speech and text information. This means that the user's question is identified based on their speech and text information. Then, using the generation capabilities learned during training, the text generation module provides a content-based response to the identified user question, generating a response text corresponding to the user's speech. Similarly, the style generation module performs style generation on the user's speech and style. This means that using the generation capabilities learned during training, the text generation module provides a style-based response based on the identified user's style, generating a response style corresponding to the user's style.
[0046] S60, the reply style and the reply text are synthesized using a speech conversion model to obtain a reply voice corresponding to the user's voice.
[0047] In essence, response speech refers to speech synthesized based on response text and response style. A speech-to-speech model is a model used for speech synthesis, that is, converting text into speech.
[0048] Specifically, the response text and response style are input into a speech conversion model. The model then synthesizes the response text and style into speech. First, the response text is segmented into words, phrases, or sentences. Semantic analysis is used to add punctuation to determine sentence structure and pauses. Grammatical and semantic analysis is then performed to capture deeper meanings and intonation features. Next, the processed text undergoes phoneme conversion to obtain all phonemes corresponding to the response text. Then, prosodic features, including syllable stress, rhythm, and intonation variations, are determined using the response style. All phonemes and prosodic features are input into an acoustic model, which generates acoustic feature parameters corresponding to the response text and style. Finally, waveform conversion is performed based on these acoustic feature parameters. A vocoder converts the acoustic feature parameters into corresponding audio waveforms, and a decoder converts these audio waveforms into speech, yielding the response speech.
[0049] In this embodiment of the invention, an intelligent question-answering method performs semantic recognition and pronoun replacement on the identified text through a semantic recognition module in a preset question-answering model. This achieves semantic recognition of the text and thus replaces pronouns in the identified text, ensuring the accuracy of speech content and style recognition, thereby improving the accuracy and efficiency of the generated style and text. The style recognition module in the preset question-answering model performs style recognition on the user's speech and text information, achieving user style identification. The generation module in the preset question-answering model generates style text based on the user's style and text information, achieving the generation of response style and response text, ensuring continuity in speech style, thereby achieving the generation of response speech, reducing computational resources, and ensuring that the generated response speech conforms to the user's style.
[0050] In one embodiment, as shown in FIG3, after step S60, that is, after synthesizing the response style and the response text using a speech conversion model to obtain the response speech corresponding to the user's speech, the method further includes:
[0051] S70, the response speech is converted into text to obtain parsed text, and the parsed text is encoded to obtain the parsed text encoding result.
[0052] S80, perform style analysis on the response speech to obtain the analyzed style, and perform style encoding on the analyzed style to obtain the analyzed style encoding result.
[0053] S90, after verifying that the parsed text and the parsing style are correct, add the parsed text encoding result and the parsing style encoding result to the target user's historical text information.
[0054] Understandably, parsed text refers to the text recognition obtained from the response speech. Parsed style refers to the style recognition obtained from the response speech. Parsed text encoding result refers to the text encoding result after response speech recognition. Parsed style encoding result refers to the style encoding result of the response speech.
[0055] Specifically, after obtaining the response speech corresponding to the user's voice, a preset recognition model is acquired, and the response speech is input into the preset recognition model. The preset recognition model performs speech recognition on the response speech to obtain the parsed text. Alternatively, the response speech is input into the preset recognition model, and the speech recognition module in the preset recognition model performs speech recognition on the response speech to obtain the parsed text. Then, the parsed text is encoded by the text encoding module in the preset recognition model, that is, the parsed text is segmented into words, all segmentation results are vectorized, and the positional encoding of each segmentation result is added to each word vector to obtain the encoded parsed text result. Similarly, a style recognition model is acquired, and the response speech is input into the style recognition model. The style recognition model performs style recognition on the response speech to obtain the parsed style. Alternatively, the response speech is input into the preset recognition model, and the style recognition module in the preset recognition model performs style recognition on the response speech to obtain the parsed style. Then, the parsed style is style encoded by the style encoding module in the preset recognition module, that is, the parsed style is processed using the capabilities learned during training to obtain the parsed style encoding result. Furthermore, the system acquires the response text and response style output by the preset question-answering model, and performs text style verification on the parsed text and parsing style based on the response text and response style. That is, it verifies the parsed text using the response text by calculating text similarity to determine if the two texts are completely identical; if not, the response text is identified as the parsed text. Simultaneously, it verifies the parsed style using the response style to determine if the response style and parsing style are completely identical; if not, the response style is identified as the parsed style. Finally, after verifying that the parsed text and parsing style are correct, the encoded results of the parsed text and parsing style are added to the target user's historical text information.
[0056] In this embodiment of the invention, by recognizing the reply speech, the parsed text and parsing style are obtained, thereby enabling the parsed text and parsing style to be added to the text information, thus ensuring the continuity of the text and solving the problem of speech style discontinuity.
[0057] In one embodiment, step S40, namely, performing style recognition on the user's voice and text information through the style recognition module in the preset question-and-answer model to obtain the user's style, includes:
[0058] S401, the user's speech is encoded by the pitch encoding unit in the style recognition module to obtain a pitch encoding sequence.
[0059] S402, the user's speech is encoded by the timbre encoding unit in the style recognition module to obtain a timbre encoding sequence.
[0060] S403, the user's speech is encoded by the prosody coding unit in the style recognition module to obtain a prosody coding sequence.
[0061] S404, the text information is encoded by the text encoding unit in the style recognition module to obtain a text encoding sequence.
[0062] S405, perform style recognition on the user's speech based on the pitch coding sequence, the timbre coding sequence, the prosody coding sequence and the text coding sequence to obtain the user's style.
[0063] In essence, pitch coding sequence refers to the encoding result of the pitch level of a user's speech. Timbre coding sequence refers to the encoding result of the timbre and tone quality of a user's speech. Prosody coding sequence refers to the encoding result of the prosodic features of a user's speech, such as rhythm, intonation, and speech rate. Text coding sequence refers to the result of encoding text information.
[0064] Specifically, the pitch coding unit in the style recognition module encodes the user's speech, which involves preprocessing steps such as noise reduction, filtering, and framing. Framing involves dividing the continuous speech signal into shorter frames for frame-by-frame analysis. The fundamental frequency is extracted from each frame using a fundamental frequency algorithm, and then the envelope of the fundamental frequency is extracted to obtain the pitch coding sequence. Then, the timbre coding unit in the style recognition module encodes the user's speech, extracting timbre features and encoding them to obtain the timbre coding sequence.
[0065] Next, the prosodic encoding unit in the style recognition module encodes the user's speech, extracting and encoding its prosodic features to obtain a prosodic encoding sequence. Then, the text encoding unit in the style recognition module encodes the text information, segmenting it into words and vectorizing the segmented text to obtain word vectors. The position vector of each word is then added to these word vectors to obtain the text encoding sequence. Finally, style recognition is performed on the user's speech based on the pitch, timbre, prosodic, and text encoding sequences. This involves learning style recognition capabilities during training by fusing or concatenating the features of these sequences to form a comprehensive feature vector. The style recognition capability then performs style recognition on this comprehensive feature vector to determine the user's style.
[0066] In this embodiment of the invention, the encoding of pitch coding sequence, the encoding of timbre coding sequence, and the encoding of prosody coding sequence are realized, thereby realizing the recognition of user style, ensuring the accuracy of user style recognition, and thus improving the accuracy of subsequent responses.
[0067] In one embodiment, step S50, namely, generating style text from the user's voice, user style, and text information using the generation module in the preset question-and-answer model to obtain the response style and response text, includes:
[0068] S501, the user's voice is encoded by the voice encoding module in the preset question-and-answer model to obtain a voice embedding sequence.
[0069] S502, the text information and the historical text information are encoded by the text encoding module in the preset question-and-answer model to obtain a text embedding sequence.
[0070] S503, the user style is style-encoded by the style encoding module in the preset question-and-answer model to obtain a style embedding sequence.
[0071] S504, the speech embedding sequence, the text embedding sequence and the style embedding sequence are generated by the generation module in the preset question-and-answer model to obtain the response style and response text.
[0072] In essence, speech embedding sequences refer to converting speech into a serialized data format that captures its semantic and temporal features. Text embedding sequences refer to converting words or phrases in text into corresponding embedding vectors and arranging these vectors in the order they appear in the text to form a serialized representation. Style embedding sequences refer to representing user style features as a series of ordered elements or vectors.
[0073] Specifically, after obtaining the user's style, the user's speech is encoded using the speech encoding module in the pre-defined question-answering model. This involves sampling the user's speech at a pre-defined sampling frequency to obtain a series of discrete sample values. These sample values are then quantized and encoded to obtain a speech embedding sequence. Next, the text encoding module in the pre-defined question-answering model encodes the text information and historical text information. This involves preprocessing the text information, such as cleaning, word segmentation, and stop word removal, to improve the accuracy and efficiency of subsequent encoding. Each word or character in the text is then mapped to a high-dimensional vector space, allowing word embeddings to capture the semantic relationships between words and ensuring that similar words have similar distances in the vector space. Finally, all word embeddings are encoded, and the encoded results of the historical text information are concatenated with the encoded text information to obtain a text embedding sequence. Further, the user's style is encoded using the style encoding module in the pre-defined question-answering model. This involves extracting features from the user's style to obtain style features. These style features are then encoded using an encoding network to obtain a style embedding sequence. Next, the generation module in the preset question-answering model generates the speech embedding sequence, text embedding sequence, and style embedding sequence. That is, the generation module uses the style generation ability learned during training to generate a response style corresponding to the user's style based on the speech embedding sequence and style embedding sequence. At the same time, the generation module uses the text generation ability learned during training to generate response text corresponding to the user's text information based on the speech embedding sequence and text embedding sequence, and determines the generated response style and response text as the reply style and reply text.
[0074] In this embodiment of the invention, the encoding of the speech embedding sequence, the encoding of the text embedding sequence, and the encoding of the user style are implemented, thereby realizing the generation of response style and response text, which improves the accuracy and efficiency of the generated style and text, and ensures that the generated response speech conforms to the user style.
[0075] In one embodiment, step S10, namely acquiring the target user's voice, includes:
[0076] S101, acquire the original speech, and perform frame segmentation processing on the original speech to obtain at least one frame data.
[0077] S102, perform endpoint detection on all the frame data to obtain the start point and end point corresponding to each frame data.
[0078] S103, the original speech is denoised based on the start and end points of all the frame data to obtain the user speech.
[0079] Understandably, raw speech refers to the user's voice asking a question, or the user's voice asking a question and responding to it. Framed data refers to fixed segments of speech. Start and end points refer to the beginning and end positions of each frame of data.
[0080] Specifically, the original speech is acquired and segmented. This means dividing the original speech into segments using fixed frequency bands (e.g., 10 seconds). For example, a 3-minute original speech can be divided into 18 segments. Each segment contains the same number of signal sampling points, and these segments are defined as frames. Then, the energy value of the signal in each frame is calculated. If the energy value of several consecutive frames at the beginning of each segment is lower than a preset energy value threshold (this threshold can be set as needed), and the energy value of the next several consecutive frames is greater than or equal to the preset energy value threshold, then the point where the signal energy value increases is the starting point of that segment. Similarly, if the energy value of the speech is high for several consecutive frames, and then decreases for a certain duration in subsequent frames, the point where the energy value decreases is considered the ending point of that segment. This determines the start and end points of each frame. The audio data between the start and end points of each segment of data is retained. Audio data between individual segments (between the end point of the first segment and the start point of the second segment) is deleted, and all non-audio data is deleted sequentially. All retained segments are then concatenated according to the segmentation order to obtain the user's audio. For example, in 18 audio segments, segments 5, 9, and 14 are noise or blank audio; these are deleted sequentially, and the remaining audio data is concatenated to obtain the user's audio. For instance, in an intelligent customer service scenario, after acquiring the user's original audio, the intelligent customer service performs segmentation processing on the audio data, determines the start and end points of each frame, and removes noise or blank audio from the user's original audio to obtain the user's audio.
[0081] In this embodiment of the invention, the start point and / or end point of each segment of frame data are determined by calculating the energy value of the signal in each segment and comparing the energy value of the segment with a preset energy value threshold. Based on the start point and / or end point of each segment of frame data, speech data is deleted, thereby extracting the user's speech, reducing speech data redundancy, and improving the accuracy of subsequent style recognition.
[0082] In one embodiment, after step S102, that is, after performing endpoint detection on all the framed data to obtain the start point and end point corresponding to each framed data, the method further includes:
[0083] S104, determine the time interval between the start point and the end point in two consecutive frames of data, and compare the time interval with a preset time interval threshold to obtain a time comparison result.
[0084] S105, when the time comparison result indicates that the interval time is greater than the preset interval threshold, it is determined that the user's voice has ended.
[0085] Understandably, the interval time refers to the time between the end point and the start point in different frames of data.
[0086] Specifically, after obtaining the start and end points corresponding to each frame of data, the time between the start and end points of two consecutive different frames of data is statistically analyzed to obtain the interval time. Thus, the interval time between the start and end points of every two consecutive different frames of data is determined. Then, a preset interval threshold is obtained, and each interval time is compared with the preset interval threshold to obtain a time comparison result corresponding to each interval time. When the time comparison result indicates that the interval time is less than or equal to the preset interval threshold, the original speech is denoised to obtain the user's speech. When the time comparison result indicates that the interval time is greater than the preset interval threshold, the user's speech is considered to have ended. In this way, the meaning represented by all time comparison results is determined, thereby determining the position where the user's speech ends. For example, in the financial field, when inquiring about a certain type of insurance, the silence time may be relatively long due to considering or recalling the insurance name. In this case, the previous speech is retained, and subsequent speech is recorded.
[0087] In this embodiment of the invention, by determining the time interval between the start and end points of two consecutive frames of data, the time between the start and end points is statistically analyzed. By comparing the time interval with a preset time interval threshold, the time comparison result is obtained, thereby determining the end of the speech.
[0088] In one embodiment, before step S10, that is, before performing speech recognition on the user's speech using a preset question-and-answer model to obtain text information corresponding to the user's speech, the following steps are included:
[0089] S106, Obtain a sample dataset, the sample dataset including at least one sample training data.
[0090] Understandably, the sample training data can be audio information corresponding to each user, or multiple audio information from a single user. The sample training data can be collected from different databases, or it can be pre-prepared data sent from the client to the database, and then a sample training dataset can be constructed based on all the acquired sample training data.
[0091] S107, the training data of all the samples are identified and predicted by a preset training model to obtain the prediction result corresponding to each of the training data of the samples.
[0092] Understandably, the prediction results are used to characterize the response style and response information obtained by identifying and predicting from the sample training data.
[0093] Specifically, all sample training data are input into a preset training model. The preset training model performs recognition and prediction on all sample training data. First, the speech recognition module in the preset training model performs speech recognition on the sample training data to obtain sample recognition text corresponding to each sample training data. Then, semantic recognition is performed on the sample recognition text, and pronouns in the sample recognition text are replaced based on the semantic recognition results to obtain sample text information. The style recognition module in the preset training model performs style recognition on each sample training data to obtain the sample recognition style. Next, the generation module in the preset training model generates style text from the sample training data, sample recognition style, and sample text information, that is, it generates the style corresponding to the sample recognition style and the text corresponding to the sample text information. The generated style and text are determined as the prediction results. In this way, the prediction results corresponding to each of the sample training data can be obtained.
[0094] S108, perform loss processing on all the prediction results using a preset loss function to obtain the prediction loss value.
[0095] Understandably, the prediction loss is generated during the prediction process using the sample training data.
[0096] Specifically, after obtaining the predicted labels, a preset loss function is acquired, and loss is calculated on all prediction results using the preset loss function to obtain the loss value corresponding to the training data of each sample. Then, the total loss value generated by the preset training model during the prediction process is calculated based on all loss values. That is, all loss values can be summed directly, or all loss values can be weighted to obtain the prediction loss value corresponding to the preset training model.
[0097] S109, when the predicted loss value reaches the preset convergence condition, the preset training model after convergence is determined as the preset question answering model.
[0098] Understandably, the convergence condition can be either the predicted loss value being less than a set threshold, or the predicted loss value being very small after 50,000 calculations and no longer decreasing, at which point training can stop.
[0099] Specifically, after obtaining the predicted loss value, if the predicted loss value does not reach the preset convergence condition, the initial parameters of the preset training model are adjusted based on the predicted loss value. All sample training data are then re-input into the preset training model with adjusted initial parameters, and iterative training is performed to obtain the predicted loss value corresponding to the preset training model with adjusted initial parameters. Then, if the predicted loss value does not reach the preset convergence condition, the initial parameters of the preset training model are adjusted again based on the predicted loss value, so that the predicted loss value of the preset training model with adjusted initial parameters reaches the preset convergence condition. In this way, the accuracy of the preset training model increases, and the obtained quality inspection results continuously approach the correct results, until the predicted loss value of the preset training model reaches the preset convergence condition. At this point, the converged preset training model is determined as the preset question-answering model.
[0100] In this embodiment of the invention, a preset training model is iteratively trained using a large amount of sample training data, and the overall loss value of the preset training model is calculated by comparing the loss function, thereby determining the predicted loss value of the preset training model. The initial parameters of the preset training model are adjusted based on the predicted loss value until the model converges, thus achieving the training of the preset question-answering model and ensuring that the preset question-answering model achieves a high expected performance.
[0101] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0102] In one embodiment, an intelligent question-answering device is provided, which corresponds one-to-one with the intelligent question-answering methods described in the above embodiments. As shown in Figure 4, the intelligent question-answering device includes a speech conversion module 10, a semantic recognition module 20, a pronoun replacement module 30, a style recognition module 40, an information reply module 50, and a text conversion module 60. Detailed descriptions of each functional module are as follows:
[0103] The speech conversion module 10 is used to acquire the user's speech of the target user and perform speech recognition on the user's speech through the speech recognition module in the preset question-and-answer model to obtain the recognized text;
[0104] The semantic recognition module 20 is used to perform semantic recognition on the recognized text through the semantic recognition module in the preset question-and-answer model to obtain the semantic recognition result;
[0105] The pronoun replacement module 30 is used to replace pronouns in the recognized text based on the semantic recognition result to obtain text information;
[0106] Style recognition module 40 is used to perform style recognition on the user's voice and text information through the style recognition module in the preset question-and-answer model to obtain the user's style;
[0107] The information response module 50 is used to generate style text from the user's voice, the user's style, and the text information through the generation module in the preset question and answer model, so as to obtain the response style and response text.
[0108] The text conversion module 60 is used to perform speech synthesis on the response style and the response text through a speech conversion model to obtain a response speech corresponding to the user's speech.
[0109] In one embodiment, the device further includes:
[0110] The text parsing module is used to convert the response speech into text to obtain parsed text, and to encode the parsed text to obtain the parsed text encoding result;
[0111] The style parsing module is used to perform style parsing on the response speech to obtain the parsed style, and to perform style encoding on the parsed style to obtain the parsed style encoding result;
[0112] The information adding module is used to add the parsed text encoding result and the parsing style encoding result to the target user's historical text information after verifying that the parsed text and the parsing style are correct.
[0113] In one embodiment, the information response module 50 includes:
[0114] A speech encoding unit is used to encode the user's speech using the speech encoding module in the preset question-and-answer model to obtain a speech embedding sequence.
[0115] A text encoding unit is used to encode the text information and the historical text information using the text encoding module in the preset question-and-answer model to obtain a text embedding sequence.
[0116] A style encoding unit is used to encode the user's style through the style encoding module in the preset question-and-answer model to obtain a style embedding sequence.
[0117] The information processing unit is used to generate the speech embedding sequence, the text embedding sequence and the style embedding sequence through the generation module in the preset question-and-answer model to obtain the response style and response text.
[0118] In one embodiment, the style recognition module 40 includes:
[0119] The pitch coding unit is used to encode the user's speech through the pitch coding unit in the style recognition module to obtain a pitch coding sequence;
[0120] A timbre encoding unit is used to encode the user's speech through the timbre encoding unit in the style recognition module to obtain a timbre encoding sequence;
[0121] The prosody coding unit is used to encode the user's speech through the prosody coding unit in the style recognition module to obtain a prosody coding sequence;
[0122] A text encoding unit is used to encode the text information through the text encoding unit in the style recognition module to obtain a text encoding sequence;
[0123] A style recognition unit is used to perform style recognition on the user's speech based on the pitch encoding sequence, the timbre encoding sequence, the prosody encoding sequence, and the text encoding sequence to obtain the user's style.
[0124] In one embodiment, the speech conversion module 10 includes:
[0125] A frame-segmentation processing unit is used to acquire the original speech and perform frame-segmentation processing on the original speech to obtain at least one frame data.
[0126] An endpoint detection unit is used to perform endpoint detection on all the frame data to obtain the start point and end point corresponding to each frame data.
[0127] The noise reduction processing unit is used to perform noise reduction processing on the original speech based on the start point and end point of all the frame data to obtain the user speech.
[0128] In one embodiment, the speech conversion module 10 further includes:
[0129] A threshold comparison unit is used to determine the time interval between the start point and the end point in two consecutive frames of data, and compare the time interval with a preset time threshold to obtain a time comparison result.
[0130] The voice ending unit is used to determine the end of the user's voice when the time comparison result indicates that the interval time is greater than the preset interval threshold.
[0131] In one embodiment, the device further includes:
[0132] A sample acquisition unit is used to acquire a sample dataset, the sample dataset including at least one sample training data.
[0133] The identification and prediction unit is used to identify and predict all the sample training data through a preset training model, and obtain the prediction result corresponding to each of the sample training data.
[0134] The loss prediction unit is used to perform loss processing on all the prediction results through a preset loss function to obtain the predicted loss value;
[0135] The model convergence unit is used to determine the converged preset training model as the preset question-answering model when the predicted loss value reaches the preset convergence condition.
[0136] For specific limitations regarding the intelligent question-answering device, please refer to the limitations of the intelligent question-answering method above, which will not be repeated here. Each module in the aforementioned intelligent question-answering device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0137] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram is shown in Figure 5. The computer device includes a processor, memory, a network interface, and a database connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores data used in the intelligent question-answering method described in the above embodiment. The network interface of the computer device communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements an intelligent question-answering method.
[0138] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described intelligent question-answering method.
[0139] In one embodiment, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described intelligent question-answering method.
[0140] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0141] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0142] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. An intelligent question-answering method, characterized in that, include: The system acquires the user's voice and performs speech recognition on the user's voice using the speech recognition module in a preset question-and-answer model to obtain the recognized text. The semantic recognition module in the preset question-answering model performs semantic recognition on the identified text to obtain semantic recognition results. Specifically, the identified text is segmented into words, and each segmented word is vectorized. Positional encoding is added to each word during vectorization. Features are extracted from the vectorized text based on semantic relationships and grammatical structure. An attention mechanism is used to capture long-distance dependencies in the identified text. A fully connected layer performs semantic recognition on the extracted features and captured relationships to obtain semantic recognition results. The semantic recognition results are used to replace pronouns in the identified text to obtain text information. The style recognition module in the preset question-answering model performs style recognition on the user's speech and text information to obtain the user's style. The generation module in the preset question-answering model generates style text from the user's speech, user style, and text information to obtain the response style and response text. A speech conversion model synthesizes the response style and response text to obtain the response speech corresponding to the user's speech. Specifically, the response text is segmented into words, phrases, or sentences. Punctuation marks are added semantically to determine the sentence structure and pause positions. Synthetic and semantic analysis is performed on the response text to capture deep semantic features. The text is processed by phoneme conversion to obtain all phonemes corresponding to the response text, based on the meaning and tone features of the layer. Prosodic features are determined through the response style. All phonemes and prosodic features are input into an acoustic model, which generates acoustic feature parameters corresponding to the response text and style. A vocoder converts these acoustic feature parameters into corresponding audio waveforms, and a decoder converts the audio waveforms into speech to obtain the response speech. The step of performing style recognition on the user's speech and text information through the style recognition module in the preset question-and-answer model to obtain the user style includes: [further details about the style recognition module are needed for a complete translation]. The pitch encoding unit in the style recognition module encodes the user's speech to obtain a pitch encoding sequence; the timbre encoding unit in the style recognition module encodes the user's speech to obtain a timbre encoding sequence; the prosody encoding unit in the style recognition module encodes the user's speech to obtain a prosody encoding sequence; the text encoding unit in the style recognition module encodes the text information to obtain a text encoding sequence; and the user's style is obtained by style recognition based on the pitch encoding sequence, the timbre encoding sequence, the prosody encoding sequence, and the text encoding sequence.
2. The intelligent question-answering method as described in claim 1, characterized in that, After synthesizing the response style and the response text using a speech conversion model to obtain a response voice corresponding to the user's voice, the method further includes: converting the response voice into text to obtain parsed text, and encoding the parsed text to obtain a parsed text encoding result; parsing the response voice into a style to obtain a parsed style, and encoding the parsed style to obtain a parsed style encoding result; and after verifying that the parsed text and the parsed style are correct, adding the parsed text encoding result and the parsed style encoding result to the target user's historical text information.
3. The intelligent question-answering method as described in claim 2, characterized in that, The step of generating style text from the user's voice, user style, and text information using the generation module in the preset question-and-answer model to obtain the response style and response text includes: encoding the user's voice using the voice encoding module in the preset question-and-answer model to obtain a voice embedding sequence; encoding the text information and historical text information using the text encoding module in the preset question-and-answer model to obtain a text embedding sequence; encoding the user's style using the style encoding module in the preset question-and-answer model to obtain a style embedding sequence; and generating the voice embedding sequence, the text embedding sequence, and the style embedding sequence using the generation module in the preset question-and-answer model to obtain the response style and response text.
4. The intelligent question-answering method as described in claim 1, characterized in that, The step of obtaining the target user's voice includes: obtaining the original voice and performing frame segmentation processing on the original voice to obtain at least one frame data; performing endpoint detection on all the frame data to obtain the start point and end point corresponding to each frame data; and performing noise reduction processing on the original voice based on the start point and end point of all the frame data to obtain the user voice.
5. The intelligent question-answering method as described in claim 4, characterized in that, After performing endpoint detection on all the frame data to obtain the start point and end point corresponding to each frame data, the method further includes: determining the interval time between the start point and end point in two consecutive frame data, and comparing the interval time with a preset interval threshold to obtain a time comparison result; when the time comparison result indicates that the interval time is greater than the preset interval threshold, the user's voice is determined to have ended.
6. The intelligent question-answering method as described in claim 1, characterized in that, Before performing speech recognition on the user's speech using a preset question-answering model to obtain text information corresponding to the user's speech, the method includes: acquiring a sample dataset, the sample dataset including at least one sample training data; performing recognition and prediction on all the sample training data using a preset training model to obtain prediction results corresponding to each sample training data; performing loss processing on all the prediction results using a preset loss function to obtain prediction loss values; and determining the preset training model after convergence as the preset question-answering model when the prediction loss value reaches a preset convergence condition.
7. An intelligent question-and-answer device, characterized in that, include: The speech conversion module is used to acquire the user's speech of the target user and perform speech recognition on the user's speech through the speech recognition module in the preset question-and-answer model to obtain the recognized text; A semantic recognition module is used to perform semantic recognition on the identified text through the semantic recognition module in the preset question-answering model to obtain semantic recognition results. Specifically, the identified text is segmented into words, and the segmented words are vectorized. Positional encoding of each word is added during the vectorization process. Features are extracted from the vectorized text based on semantic relationships and grammatical structure. An attention mechanism is used to capture long-distance dependencies in the identified text. A fully connected layer is used to perform semantic recognition on the extracted features and captured relationships to obtain semantic recognition results. A pronoun replacement module is used to replace pronouns in the identified text based on the semantic recognition results to obtain text information. The style recognition module is used to perform style recognition on the user's voice and text information through the style recognition module in the preset question-and-answer model to obtain the user's style; the information reply module is used to generate style text from the user's voice, user style, and text information through the generation module in the preset question-and-answer model to obtain the reply style and reply text; the text conversion module is used to perform speech synthesis on the reply style and reply text through a speech conversion model to obtain the reply voice corresponding to the user's voice; wherein, the reply text is segmented into words, phrases, or sentences, and punctuation marks are added semantically to determine the sentence structure and pause positions, and the reply is processed... The text undergoes syntactic and semantic analysis to capture its deeper meaning and intonation. The processed response text is then phoneme-converted to obtain all phonemes corresponding to the response text. Prosodic features are determined through the response style. All phonemes and prosodic features are input into an acoustic model, which generates acoustic feature parameters corresponding to the response text and style. A vocoder converts these acoustic feature parameters into corresponding audio waveforms, and a decoder converts the audio waveforms into speech to obtain the response speech. The style recognition module includes a pitch encoding unit, used to encode the user's speech to obtain... The system comprises: a pitch encoding sequence; a timbre encoding unit, used to encode the user's speech using the timbre encoding unit in the style recognition module to obtain a timbre encoding sequence; a prosody encoding unit, used to encode the user's speech using the prosody encoding unit in the style recognition module to obtain a prosody encoding sequence; a text encoding unit, used to encode the text information using the text encoding unit in the style recognition module to obtain a text encoding sequence; and a style recognition unit, used to perform style recognition on the user's speech based on the pitch encoding sequence, the timbre encoding sequence, the prosody encoding sequence, and the text encoding sequence to obtain the user's style.
8. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the intelligent question-answering method as described in any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the intelligent question-answering method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Intelligent voice interaction implementation methods and devices, computer equipment, and storage medium
CN108711423A