Speech recognition and natural language processing integration method and system
Through the methods of multimodal data fusion and dynamic model parameter adjustment, the recognition accuracy problem of existing speech recognition technology in complex contexts is solved, and higher speech recognition accuracy and semantic understanding ability of natural language processing are achieved.
Patent Information
- Application Number
- CN202510345171.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-24
AI Technical Summary
The existing speech recognition technology is not robust enough when facing environmental noise, user pronunciation changes and contextual context, resulting in a decrease in recognition accuracy.
A method of integrating speech recognition and natural language processing is adopted. Through multimodal data fusion, combined with deep learning and reinforcement learning technology, model parameters are dynamically adjusted to adapt to different environmental noise and user pronunciation habits, and the speech recognition accuracy and semantic understanding ability of natural language processing are improved.
It significantly improves the accuracy of speech recognition and the semantic understanding ability of natural language processing, enhances the robustness and adaptability of the system, and can more accurately identify user intentions and handle complex contexts.
Smart Images

Figure CN120220652A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of language processing, and particularly relates to a method and system for integrating speech recognition and natural language processing. Background Art
[0002] As important research directions in the field of artificial intelligence, speech recognition technology and natural language processing (NLP) technology have received extensive attention and rapid development in recent years. With the popularization of smartphones, smart home devices, and various voice assistants, users' demand for voice interaction has been increasing day by day, which has promoted the further innovation and integration of related technologies. Traditional speech recognition technology mainly focuses on the process of converting speech signals into text, usually relying on the combination of acoustic models and language models. However, these methods often show insufficient robustness when facing environmental noise, user pronunciation changes, and context. Especially in real scenarios, background noise, accents or dialects, and users' speech habits will seriously affect the recognition accuracy. Summary of the Invention
[0003] The purpose of the present invention is to provide a method and system for integrating speech recognition and natural language processing to solve the deficiencies in the prior art and improve the accuracy of speech recognition and the semantic understanding ability of natural language processing.
[0004] An embodiment of the present application provides a method for integrating speech recognition and natural language processing, the method comprising:
[0005] Performing multimodal data fusion based on the user's speech signal, context text information, and environmental sensor data, wherein the multimodal data fusion adopts a multimodal alignment network based on deep learning, and through cross-modal attention technology, extracts the correlation between the acoustic features and context semantic features in the speech signal to obtain a fused multimodal feature vector;
[0006] Inputting the multimodal feature vector into a speech recognition model based on an adaptive deep neural network for speech-to-text processing, wherein the speech recognition model combines a dynamic noise suppression algorithm and context awareness technology to adjust the recognition parameters in real time to adapt to different environmental noises and user pronunciation habits, and obtains a text output;
[0007] Inputting the text output into a semantic parsing model based on a graph neural network for context semantic analysis and user intention recognition, wherein the semantic parsing model captures the implicit semantic relationships and long-distance dependencies in the text through dynamic graph structure learning and multi-hop reasoning technology, and obtains the semantic representation of the user intention and its confidence score;
[0008] According to the semantic representation and its confidence score, the parameters of the speech recognition model and the semantic parsing model are dynamically adjusted using feedback optimization technology based on reinforcement learning. Among them, the feedback optimization technology generates a model optimization strategy by combining user interaction feedback data and real-time performance monitoring, continuously improving the speech recognition accuracy and semantic understanding accuracy.
[0009] Optionally, based on the user's speech signal, context text information, and environmental sensor data, multi-modal data fusion is performed. Among them, the multi-modal data fusion adopts a multi-modal alignment network based on deep learning. Through cross-modal attention technology, the correlation between the acoustic features and context semantic features in the speech signal is extracted to obtain a fused multi-modal feature vector, including:
[0010] Based on the user's speech signal, context text information, and environmental sensor data, a data acquisition framework based on edge computing is used to obtain multi-modal data in real time. Through an adaptive data cleaning algorithm, the data is filtered for noise and unified in format to generate a preliminary standardized data set;
[0011] For the preliminary standardized data set, a feature extraction method based on deep learning is used to extract the acoustic features in the speech signal and the semantic features in the context text information respectively. Through multi-head attention technology, the feature expressions of different modalities are captured to generate a preliminary feature representation;
[0012] For the preliminary feature representation, a feature fusion method based on cross-modal attention technology is used to perform weighted fusion of the acoustic features and semantic features. Through context awareness technology, the correlation between the acoustic features and semantic features is captured to generate a preliminary fused feature representation;
[0013] For the preliminary fused feature representation, a feature integration method based on matrix factorization is used to map the fused features into a unified feature space. Through dynamic weight adjustment technology, the timeliness and consistency of the feature vector are ensured to generate the final multi-modal feature vector.
[0014] Optionally, the multi-modal feature vector is input into a speech recognition model based on an adaptive deep neural network for speech-to-text processing. Among them, the speech recognition model adjusts the recognition parameters in real time by combining a dynamic noise suppression algorithm and context awareness technology to adapt to different environmental noises and user pronunciation habits to obtain a text output, including:
[0015] For the speech signal in the multi-modal feature vector, a preprocessing method based on a dynamic noise suppression algorithm is used. Combining the background noise information in the environmental sensor data, the noise suppression parameters are adjusted in real time. Through adaptive filtering technology, a preliminary denoised speech signal is generated;
[0016] For the preliminary denoised speech signal, a speech recognition model based on an adaptive deep neural network is adopted. Combining the semantic features in the context text information, speech-to-text processing is performed. Through context awareness technology, the recognition parameters are dynamically adjusted to adapt to the user's pronunciation habits, and a preliminary text output is generated;
[0017] For the preliminary text output, a correction method based on a language model is adopted. Combining the context semantic features and grammar rules, the recognition result is optimized. Through dynamic threshold adjustment technology, the accuracy and consistency of the recognition result are ensured, and a final text output is generated.
[0018] Optionally, the text output is input into a semantic parsing model based on a graph neural network for context semantic analysis and user intention recognition. Among them, the semantic parsing model captures the implicit semantic relationships and long-distance dependencies in the text through dynamic graph structure learning and multi-hop reasoning technology, and obtains the semantic representation of the user's intention and its confidence score, including:
[0019] For the text output, a semantic parsing model based on a graph neural network is adopted. The words in the text are abstracted as nodes in the graph structure, and the semantic relationships between the words are abstracted as edges. Through dynamic graph structure learning technology, a preliminary semantic graph structure is generated;
[0020] For the preliminary semantic graph structure, a semantic analysis method based on multi-hop reasoning technology is adopted to capture the implicit semantic relationships and long-distance dependencies in the text. Through the attention technology, the selection of the reasoning path is optimized, and a preliminary semantic representation is generated;
[0021] For the preliminary semantic representation, an intention recognition method based on a classification model is adopted. Combining the context semantic features and the user's historical behavior data, the user's intention is recognized. Through the confidence score technology, the accuracy and reliability of the intention recognition are calculated, and the semantic representation of the user's intention and its confidence score are generated.
[0022] Optionally, according to the semantic representation and its confidence score, the parameters of the speech recognition model and the semantic parsing model are dynamically adjusted by using a feedback optimization technology based on reinforcement learning. Among them, the feedback optimization technology generates a model optimization strategy by combining user interaction feedback data and real-time performance monitoring, and continuously improves the speech recognition accuracy and semantic understanding accuracy, including:
[0023] According to the semantic representation and its confidence score, a feedback data collection method based on user interaction is adopted. Combining the user's confirmation or correction behavior of the recognition result, user interaction feedback data is generated. Through lightweight data caching technology, the real-time and continuity of the feedback data are ensured;
[0024] For the running states of the speech recognition model and the semantic parsing model, an analysis method based on real-time performance monitoring is adopted. Combining the recognition accuracy, response time, and user satisfaction indicators, a performance monitoring report is generated. Through dynamic threshold adjustment technology, the accuracy and timeliness of the monitoring results are ensured;
[0025] For the user interaction feedback data and the performance monitoring report, an optimization strategy generation method based on reinforcement learning is adopted. Combining the parameter space of the speech recognition and semantic parsing models, the model parameters are dynamically adjusted. Through a multi-objective optimization algorithm, the recognition accuracy and response time are balanced to generate a preliminary optimization strategy;
[0026] For the preliminary optimization strategy, a parameter adjustment method based on an adaptive learning algorithm is adopted. Combining the real-time performance monitoring data, the parameters of the speech recognition model and the semantic parsing model are dynamically adjusted. Through feedback correction technology, the model performance is continuously optimized to generate the final optimization strategy.
[0027] Another embodiment of the present application provides a speech recognition and natural language processing integrated system, and the system includes:
[0028] A fusion module, configured to perform multi-modal data fusion according to the user's speech signal, context text information, and environmental sensor data. Among them, the multi-modal data fusion adopts a multi-modal alignment network based on deep learning. Through cross-modal attention technology, the correlation between the acoustic features and context semantic features in the speech signal is extracted to obtain a fused multi-modal feature vector;
[0029] A processing module, configured to input the multi-modal feature vector into a speech recognition model based on an adaptive deep neural network for speech-to-text processing. Among them, the speech recognition model adjusts the recognition parameters in real time by combining a dynamic noise suppression algorithm and context awareness technology to adapt to different environmental noises and user pronunciation habits to obtain a text output;
[0030] An identification module, configured to input the text output into a semantic parsing model based on a graph neural network for context semantic analysis and user intention recognition. Among them, the semantic parsing model captures the implicit semantic relationships and long-distance dependencies in the text through dynamic graph structure learning and multi-hop reasoning technology to obtain the semantic representation of the user intention and its confidence score;
[0031] An adjustment module, configured to dynamically adjust the parameters of the speech recognition model and the semantic parsing model according to the semantic representation and its confidence score by using a feedback optimization technology based on reinforcement learning. Among them, the feedback optimization technology generates a model optimization strategy by combining user interaction feedback data and real-time performance monitoring, and continuously improves the speech recognition accuracy and semantic understanding accuracy.
[0032] Another embodiment of the present application provides a storage medium, in which a computer program is stored, and the computer program is configured to execute the method described in any one of the above when running.
[0033] Another embodiment of the present application provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the method described in any one of the above.
[0034] Compared with the prior art, a speech recognition and natural language processing integration method provided by the present invention performs multi-modal data fusion based on a user's speech signal, context text information, and environmental sensor data to obtain a fused multi-modal feature vector; inputs the multi-modal feature vector into a speech recognition model based on an adaptive deep neural network for speech-to-text processing to obtain a text output; inputs the text output into a semantic parsing model based on a graph neural network for context semantic analysis and user intention recognition to obtain a semantic representation of the user intention and its confidence score; and dynamically adjusts the parameters of the speech recognition model and the semantic parsing model using a feedback optimization technique based on reinforcement learning to generate a model optimization strategy, thereby being able to improve the accuracy of speech recognition and the semantic understanding ability of natural language processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 It is a hardware structure block diagram of a computer terminal for a speech recognition and natural language processing integration method provided by an embodiment of the present invention;
[0036] Figure 2 It is a flowchart of a speech recognition and natural language processing integration method provided by an embodiment of the present invention;
[0037] Figure 3 It is a structural diagram of a speech recognition and natural language processing integration system provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0039] An embodiment of the present invention first provides a speech recognition and natural language processing integration method, which can be applied to an electronic device, such as a computer terminal, specifically, a general computer, etc.
[0040] The following takes running on a computer terminal as an example to describe it in detail. Figure 1 It is a hardware structure block diagram of a computer terminal for a speech recognition and natural language processing integration method provided by an embodiment of the present invention. AsFigure 1 As shown, the computer device includes a processor, a memory, and a network interface connected via a system bus. Among them, the memory may include a non-volatile storage medium and an internal memory.
[0041] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions that, when executed, can cause the processor to execute any one of the integrated methods for speech recognition and natural language processing.
[0042] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.
[0043] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, it can cause the processor to execute any one of the integrated methods for speech recognition and natural language processing.
[0044] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that Figure 1 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0045] It should be understood that the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0046] See Figure 2 , an embodiment of the present invention provides an integrated method for speech recognition and natural language processing, which may include the following steps:
[0047] S201, perform multimodal data fusion based on the user's voice signal, context text information, and environmental sensor data. Among them, the multimodal data fusion adopts a multimodal alignment network based on deep learning. Through cross-modal attention technology, the correlation between the acoustic features and context semantic features in the voice signal is extracted to obtain a fused multimodal feature vector.
[0048] Performing multimodal data fusion based on the user's voice signal, context text information, and environmental sensor data means that the system improves the accuracy of speech recognition and natural language processing by aggregating and integrating information from different sources. For example, the user's voice signal contains the acoustic features of pronunciation, while the context text information provides semantic information about the topic being discussed. In addition, the data from environmental sensors (such as noise level, temperature, light, etc.) provides additional context for the system to better understand the user's voice and semantics. For example, in a noisy environment, the acoustic features may be affected, while the context information can help the system infer the user's intention. By applying deep learning methods, using a multimodal alignment network and cross-modal attention technology, the system can extract the correlation between these multimodal information, thus generating a more rich multimodal feature vector, laying the foundation for subsequent processing steps.
[0049] This way of multimodal data fusion significantly enhances the effects of speech recognition and natural language processing. Since a single data source may be affected by noise, unclear pronunciation, or incomplete context, it is difficult to accurately understand the user's intention relying solely on the voice signal or text information. By integrating multiple data, the system can more comprehensively capture the correlation between the acoustic features and context semantic features in the voice signal, improving the ability to understand the user's intention in complex contexts. For example, on a noisy street, the user expresses "looking for a restaurant" in a vague voice. By combining context information (such as the recent geographical location or previous conversation content), the system can more accurately identify the user's needs. Therefore, the finally generated multimodal feature vector not only contains the user's voice features, but also rich context information related to the user's intention, greatly improving the intelligence level of the system.
[0050] Specifically, according to the user's voice signal, context text information, and environmental sensor data, a data acquisition framework based on edge computing can be adopted to obtain multimodal data in real time. Through an adaptive data cleaning algorithm, the data is filtered for noise and unified in format to generate a preliminary standardized data set.
[0051] In this step, first of all, the edge-computing-based data acquisition framework is responsible for collecting voice signals, context text information from users, and environmental sensor data in real time. The advantage of edge computing is that it can process data closer to the data source, thus reducing latency and improving real-time performance. For example, when the user is outdoors, environmental sensors can capture the noise level to help the system understand the noise level of the surrounding environment. At the same time, the voice signal will also be captured and transmitted to the processing unit. Then, the adaptive data cleaning algorithm will be applied to the collected data to filter out background noise and irrelevant information, and unify data in different formats into a standardized data set. For example, convert different audio formats to a unified sampling rate and bit rate, so as to prepare for subsequent processing.
[0052] Through this data acquisition and adaptive cleaning method, it can be ensured that the data input into the system is both high-quality and real-time. This not only improves the accuracy of the data, but also provides a good foundation for subsequent speech recognition and semantic parsing. Using edge computing reduces the latency of data transmission, thus responding to the user's voice requests in real time. For example, in an application scenario where only short-time voice input is allowed, a fast and efficient data cleaning process can avoid unnecessary delays and ensure a smooth user experience. In addition, the standardized data set can ensure the comparability between different modal data, laying the foundation for the next step of feature extraction and greatly improving the effect of multimodal fusion.
[0053] In this step, first, it is necessary to establish an edge-computing-based data acquisition framework for the purpose of collecting multimodal data in real time near the user device. These data include the user's voice signal, context text information from social media, and sensor data from the environment (such as noise level, temperature, etc.). For example, when the user says "Search for nearby coffee shops" while walking, during this process, the system will capture their voice signal and the noise data of the surrounding environment in real time, and obtain environmental information such as weather conditions and light changes through the camera or sensors. By performing preprocessing at the device end, the latency of data transmission is reduced, laying a foundation for subsequent analysis.
[0054] Then, the system will use the adaptive data cleaning algorithm to process the collected data. The goal of this algorithm is to filter out background noise and unify the data format to ensure data quality. For example, the voice signal may contain background noise, and the context text information may have various formats. The adaptive data cleaning algorithm will analyze the quality of these data in real time according to preset criteria, automatically identify and remove low-quality data, and introduce a format conversion mechanism to standardize all data into a unified format, which is convenient for subsequent feature extraction and analysis. Finally, after cleaning and format unification, a preliminary standardized data set will be generated to ensure the accuracy and usability of the data.
[0055] The significance of this process lies in ensuring the efficiency and accuracy of subsequent feature extraction and analysis steps by obtaining and cleaning multi-modal data in real time. Since data processing occurs on edge devices, it can reduce the latency caused by network transmission, making data processing more time-sensitive. For example, in a high-noise environment, real-time noise filtering can ensure that the system captures the clear user intent, reducing the risk of mis-identification caused by background noise, thereby improving the overall performance and accuracy of subsequent speech recognition and semantic analysis.
[0056] For the preliminary standardized dataset, a deep learning-based feature extraction method is adopted to extract the acoustic features in the speech signal and the semantic features in the context text information respectively. Through the multi-head attention technique, the feature expressions of different modalities are captured to generate a preliminary feature representation.
[0057] In this stage, the system will apply a deep learning-based feature extraction method to extract the key acoustic features and semantic features from the standardized dataset. For the speech signal, models such as convolutional neural network (CNN) or recurrent neural network (RNN) are used to extract acoustic features, such as features based on Mel-frequency cepstral coefficients (MFCC), which can effectively characterize the time-frequency characteristics of speech. At the same time, the text information will be processed through pre-trained language models (such as BERT, GPT, etc.) to extract the semantic features of the sentences. This process will provide a rich feature basis for multi-modal data fusion. The multi-head attention technique can help the system simultaneously focus on different feature dimensions during the feature extraction process, enhancing the expressive ability of the model.
[0058] Through deep learning feature extraction, the system can understand and represent data from different modalities more deeply. This feature extraction can not only improve the system's recognition effect of user intent but also capture the complex relationship between acoustic features and semantic features. For example, for the expression "I want to eat pizza" in speech, the acoustic features can capture the intent of "want to eat", while the semantic features provided by the text information can identify the specific object "pizza". The combination of the two will greatly improve the understanding accuracy of user needs, providing richer information for subsequent multi-modal fusion.
[0059] In this step, deep learning feature extraction is first performed on the preliminary standardized dataset. Specifically, for speech signals, the system uses a convolutional neural network (CNN) or a recurrent neural network (RNN) to extract acoustic features. These features can include Mel-frequency cepstral coefficients (MFCC) and the time-frequency features of the audio signal, etc., which effectively characterize the sound characteristics and emotional color of the speech. At the same time, for the context text information, the system uses a pre-trained language model, such as BERT, to extract the semantic features of the text. This can better understand the implicit meaning in the text, especially when the text information is closely related to the meaning of the speech signal.
[0060] Secondly, the multi-head attention technique is introduced in the feature extraction process. By setting multiple attention heads, the system can, when processing features, pay attention to the importance of different parts in the feature vector. For example, when analyzing the user's speech signal, one attention head may focus on capturing the user's intonation changes, while another attention head can focus on the emotional color of the keywords. In this way, the system can generate a richer and more multi-dimensional preliminary feature representation. Through the multi-head attention mechanism, the acoustic features and semantic features are combined to ensure that the user's intention can be fully reflected.
[0061] The significance of this feature extraction process lies in converting multi-modal information into high-dimensional feature representations that can be used for subsequent analysis. This high-level feature extraction can effectively improve the performance of speech recognition and natural language processing. For example, through the feature extraction of speech signals and text information, the system can more accurately recognize the user's speech commands, thereby providing more personalized and accurate services for them. Through the application of multi-head attention, the performance of the model can also be effectively improved, making the feature representation more powerful.
[0062] For the preliminary feature representation, a feature fusion method based on cross-modal attention technology is adopted to perform weighted fusion of acoustic features and semantic features. Through context-aware technology, the correlation between acoustic features and semantic features is captured to generate a preliminary fused feature representation;
[0063] In this step, the system performs weighted fusion of the preliminary acoustic features and semantic features to better capture the relationship between the two in the context. By applying cross-modal attention technology, the system can learn which acoustic features are more important in a specific context and generate a fused feature representation based on the weighted features. For example, when the user says "I want to go to a restaurant", the system can recognize that "going to a restaurant" is the key intention, and the acoustic features and semantic features will be associated through the attention mechanism, and then a more representative fused feature will be generated. This feature fusion process will provide sufficient information for the subsequent model input.
[0064] This feature fusion method greatly improves the accuracy of multimodal understanding. Through correlation and combination, the system can not only utilize the time-domain information of acoustic features but also refer to the context information of semantic features, thereby capturing more complex user intentions. For example, in a noisy environment, when a user says "go to the restaurant", the acoustic features may be interfered with, but after fusion, the semantic features can ensure that the system can still accurately identify the user's true intention. This feature fusion makes the system more adaptable to complex environments in practical applications and improves the overall user experience.
[0065] In this step, the system will perform feature fusion on the preliminary feature representations. Specifically, a cross-modal attention mechanism will be used to weight and fuse the acoustic and semantic features. The system first weights the feature representations of each modality and assigns corresponding weights to different features according to their importance in a specific context. For example, in certain specific contexts, acoustic features may be more important than semantic features, and vice versa. Through this weighted fusion method, the system can integrate the features of the two modalities into a preliminary fused feature representation.
[0066] In addition, through context-aware technology, the system can further capture the deep relationship between acoustic and semantic features. For example, by analyzing the corresponding features of the user saying "I'm in the coffee shop" in terms of voice and text, the system can deduce the impact of "coffee shop" on acoustic features in this scenario, helping to strengthen the interaction relationship between features. Such a process not only improves the system's understanding ability of different voices and intentions but also provides more abundant information for subsequent processing.
[0067] The role of this feature fusion is to provide a unified and multi-dimensional feature vector for the subsequent model input, enabling the model to more comprehensively understand the user's intention and its context. By fusing acoustic and semantic information, the system can more accurately identify the user's intention, especially in complex interaction scenarios, which helps to improve the intelligence level of voice assistants or artificial intelligence systems. For example, when processing multi-turn conversations, the fused features can ensure that the system captures the subtle changes in the conversation, thereby better understanding the user's meaning and making more appropriate responses.
[0068] For the preliminary fused feature representation, a feature integration method based on matrix factorization is adopted to map the fused features into a unified feature space. Through dynamic weight adjustment technology, the timeliness and consistency of the feature vector are ensured to generate the final multi-modal feature vector.
[0069] At this stage, the system will use matrix factorization techniques to integrate the preliminarily fused feature representations and map the multi-modal features into a unified feature space. This process mainly includes performing decomposition operations on the fused feature matrix to obtain lower-dimensional and more representative feature representations. By mapping the features to a common space, the system can ensure the comparability of features between different modalities and is conducive to subsequent processing steps, such as the inputs of speech recognition models and semantic parsing models. In addition, the dynamic weight adjustment technique ensures the consistency and real-time nature of feature vectors in different environments and scenarios.
[0070] Through matrix factorization and feature integration, the system can effectively improve the quality and expressive power of feature vectors. This processing method ensures the timeliness of features, enabling the system to adapt to users' pronunciation habits and environmental changes. For example, in a complex noisy environment, the feature vectors may need to re-adjust their weights to highlight important information, which will help the system perform more accurately and efficiently in speech recognition and intent understanding. In addition, the unified feature space can make the subsequent multi-modal fusion process smoother, reduce interference between different modalities, and promote the improvement of speech recognition accuracy.
[0071] In the last step, the system will perform matrix factorization on the preliminary fused feature representations. This process mainly involves decomposing the fused feature matrix into several low-dimensional feature matrices. In this way, the system can extract the potential structures and relationships between features. For example, through the singular value decomposition (SVD) technique, the information most important for user intent recognition can be effectively retained while reducing redundant features. The key to this step is to ensure that the mapped features can reflect their temporal characteristics and effectiveness in all aspects through the analysis of feature confidence.
[0072] In this process, the dynamic weight adjustment technique will also be introduced to ensure the adaptability and consistency of features in different application scenarios. By analyzing users' feedback and performance data in real time, the system can automatically adjust the weights of different features, thus better balancing the impact of different modality information on user experience and processing accuracy. For example, in some cases, users may rely more on speech input, and at this time, the system will automatically increase the weight of the corresponding acoustic features in the model.
[0073] The significance of this feature integration process lies in generating the final multi-modal feature vector, which not only integrates rich acoustic and semantic information but also has good timeliness and consistency, ensuring high-quality input for the next model processing. Through the generated multi-modal feature vector, the system can greatly improve the understanding of the user's speech and enhance the integration effect of speech recognition and natural language processing. For example, this feature vector will provide strong support for subsequent speech recognition and intent recognition, enabling the overall intelligent speech assistant or natural language processing system to respond to user needs more flexibly and accurately in complex environments.
[0074] S202, input the multi-modal feature vector into a speech recognition model based on an adaptive deep neural network for speech-to-text processing. Among them, the speech recognition model adjusts the recognition parameters in real time by combining a dynamic noise suppression algorithm and context awareness technology to adapt to different environmental noises and user pronunciation habits, and obtains a text output.
[0075] In this step, the system first inputs the fused multi-modal feature vector into a speech recognition model constructed by an adaptive deep neural network (DNN). This model has been pre-trained and can adapt to various audio input situations, especially in cases where the environmental noise is different or the user's pronunciation habits are diverse. When the user gives an instruction in a noisy environment, such as "Find a nearby restaurant", the system's model can effectively recognize the user's speech and accurately convert it into text by combining the previously extracted multi-modal features. Through the dynamic noise suppression algorithm, the model can analyze the background noise in real time and automatically adjust its own recognition parameters to ensure higher recognition accuracy in different environments.
[0076] The significance of this process is that by adjusting the recognition parameters in the speech-to-text process in real time, the system can significantly improve the accuracy and robustness of speech recognition. Especially in complex or noisy environments, traditional speech recognition technologies may fail due to background noise. However, by combining dynamic noise suppression technology, this model can not only significantly reduce unnecessary interference but also ensure capturing the user's core intention. For example, even in a noisy coffee shop environment, the system can successfully recognize the user's request, ensuring that speech recognition not only depends on the acoustic features themselves but also effectively combines context information to enhance the user experience.
[0077] Specifically, for the speech signal in the multi-modal feature vector, a preprocessing method based on a dynamic noise suppression algorithm can be adopted, combined with the background noise information in the environmental sensor data, to adjust the noise suppression parameters in real time, and through adaptive filtering technology, a preliminary denoised speech signal is generated.
[0078] In this step, the system first analyzes the input multi-modal feature vector, extracts the speech signal part, and preprocesses it using a dynamic noise suppression algorithm. This algorithm can monitor the ambient noise in real time and dynamically adjust the noise suppression parameters by combining data from environmental sensors, such as the surrounding background volume and frequency. For example, if the system detects that the user is in a noisy coffee shop, the noise suppression algorithm will automatically strengthen the suppression of low-frequency background noise to ensure that the user's speech signal can be clearly captured.
[0079] In addition, by introducing adaptive filtering technology, the system can continuously adjust the parameters of the filter according to the real-time changes in the environment. This means that if the ambient noise changes over a period of time, the system will automatically update the filter settings to optimize the processing effect. This real-time feedback mechanism can significantly improve the denoising effect, making the processed speech signal clearer and laying a solid foundation for subsequent speech recognition algorithms.
[0080] The purpose of this process is to ensure that the signal input to the speech recognition model is as pure as possible, greatly improving the system's recognition accuracy of user commands. For users, clear speech input means a smoother interaction experience. In a noisy environment, the effective application of denoising technology ensures that even in a complex background, the system can still capture the essence of the voice command, thereby enhancing the reliability of speech recognition.
[0081] In this implementation process, it is first necessary to extract the speech signal part from the multi-modal feature vector, and then use the dynamic noise suppression algorithm for preprocessing. The core of the dynamic noise suppression algorithm is to dynamically adjust the noise suppression parameters according to the background noise information collected by the environmental sensors. For example, if the user is speaking on the street, the environmental sensor may detect higher-frequency traffic noise. The system will enhance the suppression of this type of noise based on this data and focus on the user's speech signal to ensure the clarity of the speech signal.
[0082] Next, the adaptive filtering technology is used to adjust the filter parameters in the denoising process in real time. The adaptive filter can update its parameters in real time according to the input signal characteristics, so that the filter can better adapt to the changing ambient noise. For example, in a noisy restaurant environment, this technology can automatically optimize the processing to ensure that the user's speech signal can be effectively distinguished from the background noise, generating a clearer denoised speech signal. By combining these technologies, the system can generate a high-quality preliminary denoised speech signal, laying a foundation for subsequent speech-to-text processing.
[0083] Finally, to further enhance the denoising effect, the system can also adopt multi-channel signal processing technology. This technology uses the speech signals captured by multiple microphones for phase difference processing, thereby further reducing the noise from the environment. In specific implementation, the system will perform synchronous processing on the signals of each microphone, combined with the previous dynamic noise suppression algorithm, to jointly generate an optimized denoised speech signal. This process ensures that the signal input into the speech recognition model is as clean as possible, significantly improving the accuracy of subsequent processing.
[0084] For the preliminary denoised speech signal, a speech recognition model based on an adaptive deep neural network is adopted. Combining the semantic features in the context text information, speech-to-text processing is carried out. Through context-aware technology, the recognition parameters are dynamically adjusted to adapt to the user's pronunciation habits, generating a preliminary text output.
[0085] At this stage, the system inputs the processed denoised speech signal into a speech recognition model constructed by an adaptive deep neural network (DNN). Different from traditional models, this model can dynamically adapt to the user's pronunciation habits and context. Specifically, the model quickly learns and adjusts its recognition strategy by comparing the user's real-time speech signal with the previously collected speech data. For example, if the system finds that the user is accustomed to speaking quickly, the model will appropriately adjust the parameters to adapt to this pronunciation pattern and ensure the accurate capture of the speech content.
[0086] Combined with the context text information, the speech recognition model will also consider the previous user instructions and the theme of the current conversation. Through context-aware technology, the model can analyze the user's intention in real time during the recognition process. For example, when the user says "I want to eat pizza", the model not only analyzes the speech but also reviews the previous conversation to identify the relevant context of "eat", thereby enhancing the understanding and capture of the user's intention. Finally, after a series of processes, the model will generate a preliminary text output.
[0087] The key significance of this process lies in that by dynamically adjusting the recognition parameters, the system can adapt to different pronunciation methods and styles of users, thereby ensuring the accuracy of the text output. This flexible adaptability is an important factor in enhancing the user interaction experience, especially in a complex and changing natural language environment. It can be said that the high accuracy of speech input directly affects whether the user is willing to interact with the system again.
[0088] At this stage, the system inputs the denoised speech signal into a speech recognition model based on an adaptive deep neural network (DNN) for processing. First, the model extracts acoustic features from the input speech signal, using techniques such as convolutional neural networks (CNNs) to extract the time-frequency features of the audio signal. These features include the frequency components of the speech and the transient information of the time-domain waveform, etc. The purpose of this process is to transform the original speech signal into a feature representation for subsequent processing by the deep learning model.
[0089] Next, the system combines the semantic features in the context text information to enhance the accuracy of the recognition result. By introducing context-aware technology, the model not only relies on the acoustic features of the audio input but can also refer to the previous conversation history or relevant information to better understand the user's intention. For example, when the user is saying "I want to play a song" through the voice assistant, the system will recognize the two keywords "play" and "song" and combine them with the context information to infer the user's specific needs and generate a preliminary text output.
[0090] Finally, to adapt to the pronunciation habits of different users, the system dynamically adjusts the recognition parameters. The adaptive deep neural network has the ability to learn and remember the pronunciation patterns of specific users. The model will optimize the recognition algorithm by monitoring the speech changes of the user during input. For example, if a user is accustomed to speaking quickly, the system will automatically increase the time window to capture each syllable in the rapid pronunciation, thereby improving the accuracy of text generation. The dynamic adjustment of this process ensures that the speech recognition model can always provide high-quality output in different environments and backgrounds.
[0091] For the preliminary text output, a correction method based on a language model is adopted. Combining the context semantic features and grammar rules, the recognition result is optimized. Through dynamic threshold adjustment technology, the accuracy and consistency of the recognition result are ensured to generate the final text output.
[0092] In this step, the system inputs the preliminary text output into the language model for correction. The introduction of the language model enables the system to not only rely on the already generated text result but also combine the context semantic features and grammar rules to further optimize the recognition result. Specifically, the system analyzes the preliminarily generated text, compares it with known language patterns, and identifies possible spelling mistakes and grammar irregularities. For example, if the user's instruction is "I want to find a restaurant", the system can anticipate the user's intention and correct it to "I want to go to find a restaurant" to improve the fluency and accuracy of the text.
[0093] In addition, the introduction of dynamic threshold adjustment technology enables the system to flexibly adjust the confidence in the recognition results based on context and user interaction feedback. For example, when the system detects that most of the content in the conversation is related to "restaurant", it can appropriately reduce the error tolerance for ambiguous words to ensure that the final output text is not only rigorous in terms of words and sentences, but also more in line with the actual intentions of the users. During the correction process, the system will monitor these parameters in real time and quickly adjust according to previous user behaviors to ensure the accuracy and consistency of recognition.
[0094] The significance of this process is that through the correction of the language model, the system can generate text outputs that are more in line with user intentions and natural language norms, improving the effectiveness and fluency of subsequent processing. The intuitive feedback of the user experience is particularly important at this time. Only by ensuring the high quality of the recognition results can user satisfaction be improved and good interaction between users and the system be promoted.
[0095] When performing text output correction, the system will first evaluate the preliminary text results and perform grammar and semantic corrections in combination with the existing language model. The construction of the language model relies on a large amount of text data and can help identify certain word combinations or grammar structures that do not conform to language habits. For example, if the preliminary text output is "I want to go to the store to buy a bottle of water", but the user's true intention is to go to the supermarket to buy water, the system will use context semantic features for optimization and finally correct it to "I want to go to the supermarket to buy water".
[0096] On this basis, the system introduces dynamic threshold adjustment technology to dynamically adjust the correction criteria according to the current context and user interaction feedback. By analyzing the user's confirmation or negation behavior of the text output, the model can self-learn and optimize the amplitude and direction of its changes. For example, if the user often confirms "go to the supermarket" instead of "go to the store", the system will record this habit and give more priority in future processing to ensure that the text output is more in line with the user's habits.
[0097] Finally, the corrected text output will be compared with the context information again to ensure its accuracy and logical consistency. The system will build a feedback loop, record the confirmed text output as the user's new preference, and form a closed-loop learning mechanism. At this time, if the user uses a similar expression in the subsequent conversation, the system will generate a more accurate text output based on the previous correction results, forming a smoother interaction experience. Through this effective correction and optimization mechanism, the finally generated text output will have higher accuracy, consistency, and user satisfaction.
[0098] S203. Input the text output into a semantic parsing model based on a graph neural network for context semantic analysis and user intention recognition. The semantic parsing model captures implicit semantic relationships and long-distance dependencies in the text through dynamic graph structure learning and multi-hop reasoning techniques, and obtains the semantic representation of the user intention and its confidence score.
[0099] In specific operations, first, each word in the text output is abstracted as a node in the graph structure, and the semantic relationships between these nodes are represented by edges. Through continuous updating of the dynamic graph structure, the model can flexibly handle context information that changes over time, thereby enhancing the understanding of semantic relationships. At the same time, the multi-hop reasoning technique enables the model to perform reasoning at multiple levels, uncovering implicit meanings or potential user intentions. Therefore, the final model output will include the semantic representation of the user intention and its confidence score, providing the necessary information basis for subsequent operations such as response generation and task execution.
[0100] This step, by combining advanced techniques of graph neural networks, enables the semantic parsing model to not only understand the surface meaning of the text but also deeply excavate the implicit relationships in the text. Through in-depth analysis of the context, the model can accurately identify the user's intention, laying a foundation for providing personalized services and efficient interaction experiences. For example, if a voice assistant can accurately understand the user's needs: when the user says "I want to order food", the system can not only recognize the intention of "ordering food" but also combine historical conversation information to judge the user's preferences and possible choices, thus providing services that better meet the user's expectations. This precise intention recognition ability is the key to achieving more advanced natural language processing.
[0101] Specifically, for the text output, a semantic parsing model based on a graph neural network can be adopted to abstract the words in the text as nodes in the graph structure and the semantic relationships between the words as edges, and generate a preliminary semantic graph structure through dynamic graph structure learning techniques.
[0102] In this stage, the system first converts each word in the text output into a node in the graph structure. This step includes word segmentation of the text to ensure that the grammatical and semantic information of each word is retained. Subsequently, the semantic relationships between the word nodes are established, and edges are used to represent the connections between different words, such as synonyms, antonyms, or dependencies in the context. Through this graph structure, the system can flexibly represent the information in the text, and then through dynamic graph structure learning techniques, each node can interact with other nodes to obtain context information and implicit relationships. In this way, the reasoning ability of the model is significantly enhanced, and it can capture more complex semantic structures.
[0103] This process abstracts the text into a graph structure, providing a more intuitive and flexible framework for understanding and analysis. The introduction of graph neural networks enables the model to analyze semantics at different levels, more effectively handle long-distance dependencies and context information. For example, when the user's input contains multiple layers of meaning or implicit information, the system will analyze it layer by layer through the graph structure, extract the relationships between keywords, and thus more accurately understand the user's intention. This ability significantly improves the depth and accuracy of semantic parsing, laying a good foundation for subsequent user intention recognition.
[0104] When implementing this process, the system first performs word segmentation on the user's text input to extract the keywords in the text. Taking the sentence "I want to book a flight for tomorrow" as an example, after word segmentation, the system takes words such as "I", "want", "book", "a", "tomorrow", "flight" as nodes in the graph. Next, the system analyzes the relationships between these words. For example, the relationship between the action of "booking" and "flight" is connected by establishing edges, and these edges can represent semantic interrelationships, such as the relationship between an action and an object. At the same time, if there is a time description of "tomorrow" in the context, an appropriate time marker can be added to it in the graph.
[0105] After constructing the initial graph structure, the system will use dynamic graph structure learning technology to allow the representations between nodes to be dynamically updated according to context information. Specifically, as the user input changes, the graph structure will continuously adjust to timely reflect new semantic relationships and context information. This is particularly important for capturing long-distance dependencies. For example, if information such as "night flight" or "business class" is mentioned in the user's subsequent input, the system can timely integrate it into the original graph structure to form a more complete semantic graph. The generation of this initial semantic graph structure lays a foundation for subsequent in-depth semantic analysis, enabling the model to better capture the key intentions and details in the text.
[0106] For the initial semantic graph structure, a semantic analysis method based on multi-hop reasoning technology is adopted to capture the implicit semantic relationships and long-distance dependencies in the text, and through attention technology, optimize the selection of the reasoning path to generate an initial semantic representation;
[0107] In this step, the system performs in-depth semantic analysis based on the preliminary semantic graph structure. Through multi-hop reasoning technology, the model can perform multiple reasoning in the graph structure, exploring implicit semantic relationships and long-distance dependencies layer by layer. For example, when analyzing sentences, the system not only relies on directly adjacent words, but also deeply understands the slightly scattered key information in the text through the hierarchical structure of the graph. In addition, in order to improve the reasoning effect, the system combines the attention mechanism to focus on those nodes that are most important for intent recognition, thereby selecting the reasoning path more effectively. This dynamically optimized reasoning process enables the model to more comprehensively capture the intrinsic meaning of the text and generate a preliminary semantic representation.
[0108] Through the application of multi-hop reasoning technology, the model can more deeply understand the complex semantic structure in the text, such as implicit meaning and long-distance grammatical relations. This ability significantly improves the intelligence level of the system, enabling it to accurately grasp the user's real needs when faced with ambiguous or ambiguous user expressions. For example, when a user says "I want to go to that very special restaurant", the system may not be able to accurately grasp the user's intention if it relies solely on superficial vocabulary, but through multi-hop reasoning analysis, the system can identify the meaning of "special" and associate it with the user's historical preferences, providing a more targeted basis for further responses.
[0109] In this step, the system will conduct an in-depth analysis of the generated preliminary semantic graph structure and use multi-hop reasoning technology to capture more complex implicit semantic relationships in the text. For example, in the context of a user expressing "I want to book a flight ticket for tomorrow", the model not only focuses on the intuitive relationship between "book" and "flight ticket", but also makes multiple inferences to understand the impact of "tomorrow" on the entire behavior. Through multi-hop reasoning, the system can iteratively deduce from one node to another, gradually extracting more contextual information and potential user intentions.
[0110] To improve the efficiency of the reasoning process, the system uses the attention mechanism to emphasize nodes that are more important in a specific context. Continuing with the above example, the system may recognize the importance of the "air ticket" node in the user's intent, and therefore allocate more computing resources to nodes directly related to "air tickets" during reasoning, such as "reservation" and "travel date". This targeted reasoning allows the system to more accurately capture the user's true intent in complex sentences and generate preliminary semantic representations.
[0111] The preliminary semantic representation finally obtained through multi-hop reasoning will contain the core information of the user's intention and its constituent elements. The system can output the corresponding structured representation, such as "the user wants to order a plane ticket tomorrow", and mark the key nodes involved and the relationship between them, thereby providing a solid foundation for subsequent intention recognition and feedback mechanism.
[0112] For the preliminary semantic representation, an intent recognition method based on a classification model is adopted. Combining context semantic features and user historical behavior data, user intent is recognized. Through confidence scoring technology, the accuracy and reliability of intent recognition are calculated, and a semantic representation of the user intent and its confidence score are generated.
[0113] The key in this process is to input the preliminarily generated semantic representation into the classification model to identify the user's intent. This model will use context semantic features and the user's historical behavior data for training and prediction. First, the system will convert the semantic representation into feature vectors suitable for processing by the classification model. These feature vectors not only contain information related to the current conversation but also incorporate the user's past interaction data to identify intents that better match the user's preferences. Then, the model will perform intent classification according to the set intent categories (such as querying information, performing tasks, providing suggestions, etc.). Finally, the system will calculate the confidence score for each recognition result, reflecting the model's confidence in the recognition result. For example, when the user says "I want to book a flight ticket for tomorrow", after analysis, the system may recognize the intent of "booking a ticket" and give a relatively high confidence score, indicating that the model is very certain about this recognition result.
[0114] The key to this step is to clarify the user's intent through the classification model, enabling the system to accurately respond to the user's needs. The confidence score provides a basis for further decision-making, helping the system evaluate the reliability of the recognition result and avoid misjudgment. For example, in the case where the user's expressed intent is unclear, the system can choose to ask for confirmation instead of directly performing an operation. By combining context semantics and historical behavior data, the model can not only improve the accuracy of user intent recognition but also provide a more personalized user experience. If the system can adapt to the user's preferences in a timely manner, it can effectively improve user satisfaction and usage loyalty.
[0115] In this stage, the system inputs the preliminarily generated semantic representation into the classification model for intent recognition. By combining context semantic features and the user's historical behavior data, the system can construct a set of multi-dimensional feature vectors that not only describe the intent of the current sentence but also incorporate the user's past behavior patterns. For example, the system can analyze whether the user has often booked flight tickets in the past, their preferred departure and destination locations, and their preferences for flight times, etc., which will be reflected in the semantic features through historical data.
[0116] During the intent recognition process, the classification model uses these feature vectors for training and prediction. For example, the model may recognize "I want to book a flight ticket for tomorrow" as the intent of "booking a ticket", and at the same time calculate the confidence score of this recognition result, indicating the model's confidence in this classification result. By using algorithms such as Softmax regression or neural networks, the system can normalize the probabilities of different intents, thus assigning a confidence score to each intent.
[0117] Finally, based on the intent recognition result and its confidence score, the system can generate a more accurate semantic representation of the user's intent. For example, if the model predicts the confidence of "booking a ticket" as 0.92, the system can surely provide the user with relevant information about flight reservation; if the confidence is below a certain preset threshold, the system can choose to adopt a strategy of confirmation questions, such as asking the user "Do you want to book a flight ticket?" to ensure that it can accurately respond to the user's needs. This feedback mechanism based on accurate recognition not only improves the user experience but also enhances the intelligence level of the system.
[0118] S204, according to the semantic representation and its confidence score, use the feedback optimization technology based on reinforcement learning to dynamically adjust the parameters of the speech recognition model and the semantic parsing model, where the feedback optimization technology generates a model optimization strategy by combining user interaction feedback data and real-time performance monitoring, and continuously improves the speech recognition accuracy and semantic understanding accuracy.
[0119] This process first collects the user's interaction feedback data, including the user's confirmation, correction, or negation of the speech recognition result and other behaviors. These feedback messages constitute a valuable data source, which can reflect the user's satisfaction and recognition accuracy with the system's recognition result. Subsequently, the system combines the data of real-time performance monitoring to conduct a real-time evaluation of the model, such as monitoring the recognition accuracy rate, response time, and the user's overall experience, so as to form a comprehensive understanding of the model performance. Based on these data, the system will generate an optimization strategy and dynamically adjust the parameters to improve the performance of the model in different situations, and finally achieve the goal of improving the speech recognition accuracy and semantic understanding accuracy.
[0120] The importance of this process lies in ensuring that the system has the ability to self-improve through continuous feedback loops, thereby enhancing the user experience. The direct feedback from users provides real data-driven input for the system, ensuring that the system can adaptively adjust itself to handle different voice inputs and environments, thus effectively reducing the occurrence of recognition errors and misunderstandings. The application of reinforcement learning in this process enables the system to enhance its intelligence level during continuous interactions, achieve refined adjustments, and then provide users with a more accurate and smooth voice recognition experience. This optimization mechanism not only improves the intelligence of the system but also lays a foundation for future development, enabling it to continuously learn and adapt to user needs in practical applications.
[0121] Specifically, according to the semantic representation and its confidence score, a feedback data collection method based on user interaction can be adopted. By combining the user's confirmation or correction behavior of the recognition result, user interaction feedback data can be generated. Through lightweight data caching technology, the real-time and continuity of the feedback data can be ensured.
[0122] This step involves collecting feedback data from user interactions for the system's self-optimization. The system monitors the user's explicit feedback on the voice recognition results, whether to accept, correct, or reject these results. For example, when the user requests the voice assistant to make a setting, if the assistant recognizes it as "increase the volume" and the user corrects it to "decrease the volume", this correction information is valuable feedback data. At the same time, the system uses lightweight data caching technology to ensure the real-time and continuity of this feedback data during user interactions, avoiding affecting the rapid learning and adjustment of the model due to data latency.
[0123] By capturing the user's feedback on the system results in real-time, the system can immediately identify the deficiencies in recognition accuracy and quickly reflect the true needs of users. This process enables the system to have good adaptability and optimize according to the usage habits of specific user groups. By collecting a large amount of feedback data, the learning ability of the system is continuously improved, and it can reduce errors and improve user satisfaction and experience in subsequent interactions.
[0124] In this process, the system first constructs a user interaction monitoring module that can track and record all interactions between the user and the voice recognition system. Whenever the user issues a voice command, the system captures the user's reaction. For example, the user may say, "Set the alarm for 7 am." If the system misrecognizes it as "Set the alarm for 8 am", the user may immediately correct it to "Not 8 am, 7 am." Such confirmation or correction behaviors are valuable feedback information, which will be marked and saved.
[0125] To ensure the real-time and continuous nature of feedback data, the system adopts lightweight data caching technology. When users interact, feedback information is temporarily stored in memory instead of being directly written to the database, which can avoid the sluggish response caused by data write latency. When designing the system, message queue technology can be used to instantly push each user feedback to the processing module to avoid losing important interaction information. In addition, even when users have not used the system for a long time, the system will retain previous feedback for subsequent analysis and optimization.
[0126] When generating user interaction feedback data, the system not only records users' correction behaviors but also combines semantic representation and confidence scores to form a series of data features. For example, if a user corrects a specific command multiple times, the system marks this command as a high-error command and focuses on it during subsequent model training. Through these means, the system can build a multi-dimensional user feedback data warehouse, providing a basis for subsequent analysis and optimization.
[0127] Regarding the operating states of the speech recognition model and the semantic parsing model, an analysis method based on real-time performance monitoring is adopted. Combining recognition accuracy, response time, and user satisfaction metrics, a performance monitoring report is generated. Through dynamic threshold adjustment technology, the accuracy and timeliness of the monitoring results are ensured.
[0128] In this step, the system monitors the overall performance of the speech recognition and semantic parsing models to ensure their efficient operation in actual operations. Through real-time performance monitoring, the system can continuously track multiple metrics of the models, including the accuracy of recognition, the response time for processing requests, and user satisfaction. For example, if user feedback shows a decrease in the recognition rate during a certain period, the system can immediately obtain the data and conduct analysis. These data are integrated to form a performance monitoring report, which will help the development team identify potential problems and formulate improvement measures.
[0129] The importance of real-time performance monitoring lies in ensuring that the system is always in the best operating state. The system can react quickly and adjust strategies to avoid user loss caused by long-term performance degradation. The dynamic threshold adjustment technology enables the system to flexibly adjust the evaluation criteria according to different business requirements and usage scenarios. For example, in a high-demand scenario, the system may adjust the model's parameters to improve the recognition speed, while in an occasion where high recognition accuracy is required, it may tend to comprehensively optimize the model. This flexibility ensures that the system can always meet users' needs and enhance the user experience.
[0130] When implementing real-time performance monitoring, the system needs to establish a multi-level monitoring framework that can track multiple performance metrics simultaneously. The system will first set key metrics, such as recognition accuracy, response time, and user satisfaction. For example, recognition accuracy can be calculated by comparing with the actual user input to obtain the correct recognition rate; the response time is the time difference between when the user issues a command and when the system returns the result; and user satisfaction can be measured through the feedback from the user in subsequent conversations or the ratings given actively.
[0131] To generate a performance monitoring report, the system will regularly (e.g., every hour) summarize and analyze the monitoring data. The monitoring module will input the collected data into the data analysis model to generate a report with detailed metrics. For example, if it is detected that the dissatisfaction of users with the recognition results has increased during a certain period, the system will mark the data of that period and analyze the potential reasons, such as whether there is environmental noise interference or changes in user pronunciation, etc.
[0132] In the process of using dynamic threshold adjustment technology, the system can automatically adjust the performance detection standard according to the changes in real-time monitoring data. For example, when the recognition accuracy of the system is lower than the preset threshold, the system can issue an alarm and quickly start a self-correction program to adjust the hyperparameters of the speech recognition model to ensure that the monitoring results still maintain high accuracy and timeliness in abnormal situations. This technology not only improves the stability of the model but also provides a smoother experience for users.
[0133] For the user interaction feedback data and performance monitoring report, adopt an optimization strategy generation method based on reinforcement learning, combine the parameter spaces of the speech recognition and semantic parsing models, dynamically adjust the model parameters, and balance the recognition accuracy and response time through a multi-objective optimization algorithm to generate a preliminary optimization strategy;
[0134] In this process, the system will combine the user interaction feedback data and the performance monitoring report and use reinforcement learning technology to generate an optimization strategy. The system regards indicators such as recognition accuracy and response time as states, and the corresponding adjustment strategies as actions. By analyzing the historical interaction data and monitoring reports, the system can determine which parameter adjustments have brought positive improvements. For example, if a certain user group reports insufficient recognition rate at a fast speaking speed, the system will use reinforcement learning algorithms to derive appropriate adjustment strategies to improve the model's adaptability to speech in this situation.
[0135] By combining reinforcement learning with feedback data, the system can continuously learn and improve, achieving self-optimization in different scenarios. Such an optimization strategy can not only enhance the accuracy of speech recognition but also optimize the response speed, making the system more flexible under multi-objective requirements and truly meeting the personalized needs of users. The introduction of reinforcement learning adds impetus to the system's adaptive ability, enabling it to always maintain good performance in a dynamically changing environment.
[0136] In this step, the system comprehensively analyzes user interaction feedback and performance monitoring reports, and uses reinforcement learning techniques to generate optimization strategies. To implement this process, the system first needs to build a reinforcement learning framework that can model user feedback data to achieve learning and adaptation. By treating the user's feedback as a reward signal, the system can use Q-learning or other reinforcement learning algorithms to gradually optimize the parameters of the speech recognition and semantic parsing models.
[0137] For example, when analyzing user feedback, the system finds that the recognition accuracy significantly decreases when the user uses voice commands in a noisy environment. Through reinforcement learning model training, the system marks this environment as a "low recognition scenario" and applies an adaptive noise suppression algorithm in subsequent model adjustments. For example, the system may adjust the weight of noise suppression to enhance the recognition ability in this scenario.
[0138] In addition, to achieve multi-objective optimization, the system considers two metrics, recognition accuracy and response time, simultaneously during the optimization process. The algorithm continuously explores various parameter combinations and evaluates their impact on recognition performance, such as whether increasing the recognition accuracy will lead to an extended response time. Finally, the system generates a preliminary optimization strategy based on this feedback to ensure efficient response ability while meeting user needs.
[0139] For the preliminary optimization strategy, a parameter adjustment method based on an adaptive learning algorithm is adopted. Combining real-time performance monitoring data, the parameters of the speech recognition model and the semantic parsing model are dynamically adjusted. Through feedback correction technology, the model performance is continuously optimized to generate the final optimization strategy.
[0140] In this step, the system further refines the initially generated optimization strategy. Through a parameter adjustment method based on an adaptive learning algorithm, combined with real-time performance monitoring data, the parameters of the speech recognition model and the semantic parsing model are dynamically adjusted. The adaptive learning algorithm will evaluate the current model's performance and user feedback based on the real-time monitored data, and further adjust the hyperparameters or network structure of the model to adapt to the changing environment and user needs. For example, when user feedback shows that the recognition rate in a specific situation is not as expected, the system will activate the adaptive learning technology to make necessary parameter corrections to the model to ensure better meeting user expectations.
[0141] The importance of this process lies in that it enables the system to continuously reflect on and optimize its parameter settings during operation, ensuring that the model can maintain optimal performance in various usage scenarios. By combining real-time performance monitoring, the system can flexibly respond to unexpected situations, such as adjusting the recognition algorithm under different environmental noises, thereby continuously improving the accuracy of speech recognition and the depth of semantic parsing. The application of this feedback correction technology endows the system with the ability of persistent self-improvement, providing a better interaction experience for users and enhancing the intelligence level of the system.
[0142] When performing this step, the system will use adaptive learning algorithms (such as adaptive gradient algorithm, AdaBoost, etc.) to refine and adjust the preliminary optimization strategy. First, the system will analyze the real-time performance monitoring data to identify the performance of the current model in different scenarios, such as the recognition effect in a high-noise environment and a quiet environment. This analysis will help the system determine the specific parameters that need to be adjusted.
[0143] For example, if in a noisy coffee shop, the user feedbacks that the recognition accuracy is low, the system can, through the adaptive learning algorithm, preferentially adjust the parameters of noise suppression to enhance the clarity of the speech signal. At this time, the system can gradually adjust the noise parameters and perform self-calibration through online test feedback to ensure that the best recognition effect can be provided in these specific environments.
[0144] Continuing with the application of the feedback correction technology, the system will continuously collect the user's speech input and correction feedback. Whenever the user gives feedback, the system immediately records it and incorporates it into the next round of model adjustment. For example, in the case where the user continuously corrects a specific command, the system will update the model parameters according to these feedback records, thereby improving the recognition rate of this specific command. Eventually, through the above process, the system will generate the final optimization strategy to ensure that it has higher recognition accuracy and response speed in a variety of usage scenarios, significantly enhancing the user interaction experience. Through such a continuous optimization process, the system can continuously evolve during long-term operation to adapt to the changing needs of users.
[0145] It can be seen that, based on the user's voice signal, context text information, and environmental sensor data, multi-modal data fusion is performed to obtain a fused multi-modal feature vector; the multi-modal feature vector is input into a speech recognition model based on an adaptive deep neural network for speech-to-text processing to obtain a text output; the text output is input into a semantic parsing model based on a graph neural network for context semantic analysis and user intention recognition to obtain the semantic representation of the user intention and its confidence score; according to the semantic representation and its confidence score, the parameters of the speech recognition model and the semantic parsing model are dynamically adjusted using a feedback optimization technique based on reinforcement learning to generate a model optimization strategy, thereby improving the accuracy of speech recognition and the semantic understanding ability of natural language processing.
[0146] Another embodiment of the present invention provides a speech recognition and natural language processing integrated system. Refer to Figure 3 , the system may include:
[0147] A fusion module 301, configured to perform multi-modal data fusion based on the user's voice signal, context text information, and environmental sensor data. Among them, the multi-modal data fusion adopts a multi-modal alignment network based on deep learning, and through cross-modal attention technology, extracts the correlation between the acoustic features and context semantic features in the voice signal to obtain a fused multi-modal feature vector;
[0148] A processing module 302, configured to input the multi-modal feature vector into a speech recognition model based on an adaptive deep neural network for speech-to-text processing. Among them, the speech recognition model combines a dynamic noise suppression algorithm and context awareness technology to adjust the recognition parameters in real time to adapt to different environmental noises and user pronunciation habits to obtain a text output;
[0149] An identification module 303, configured to input the text output into a semantic parsing model based on a graph neural network for context semantic analysis and user intention recognition. Among them, the semantic parsing model captures the implicit semantic relationships and long-distance dependencies in the text through dynamic graph structure learning and multi-hop reasoning technology to obtain the semantic representation of the user intention and its confidence score;
[0150] An adjustment module 304, configured to dynamically adjust the parameters of the speech recognition model and the semantic parsing model according to the semantic representation and its confidence score using a feedback optimization technique based on reinforcement learning. Among them, the feedback optimization technique generates a model optimization strategy by combining user interaction feedback data and real-time performance monitoring to continuously improve the accuracy of speech recognition and the accuracy of semantic understanding.
[0151] It can be seen that, based on the user's voice signal, context text information, and environmental sensor data, multi-modal data fusion is performed to obtain a fused multi-modal feature vector; the multi-modal feature vector is input into a speech recognition model based on an adaptive deep neural network for speech-to-text processing to obtain a text output; the text output is input into a semantic parsing model based on a graph neural network for context semantic analysis and user intent recognition to obtain a semantic representation of the user intent and its confidence score; according to the semantic representation and its confidence score, the parameters of the speech recognition model and the semantic parsing model are dynamically adjusted using a feedback optimization technique based on reinforcement learning to generate a model optimization strategy, thereby improving the accuracy of speech recognition and the semantic understanding ability of natural language processing.
[0152] An embodiment of the present invention also provides a storage medium, in which a computer program is stored, and the computer program is configured to execute the steps in any one of the above method embodiments when running.
[0153] Specifically, in this embodiment, the above storage medium may be configured to store a computer program for executing the following steps:
[0154] S201, perform multi-modal data fusion based on the user's voice signal, context text information, and environmental sensor data. Among them, the multi-modal data fusion uses a multi-modal alignment network based on deep learning, and through cross-modal attention technology, extracts the correlation between the acoustic features and context semantic features in the voice signal to obtain a fused multi-modal feature vector;
[0155] S202, input the multi-modal feature vector into a speech recognition model based on an adaptive deep neural network for speech-to-text processing. Among them, the speech recognition model combines a dynamic noise suppression algorithm and context awareness technology to adjust the recognition parameters in real time to adapt to different environmental noises and user pronunciation habits to obtain a text output;
[0156] S203, input the text output into a semantic parsing model based on a graph neural network for context semantic analysis and user intent recognition. Among them, the semantic parsing model captures the implicit semantic relationships and long-distance dependencies in the text through dynamic graph structure learning and multi-hop reasoning technology to obtain a semantic representation of the user intent and its confidence score;
[0157] S204, according to the semantic representation and its confidence score, dynamically adjust the parameters of the speech recognition model and the semantic parsing model using a feedback optimization technique based on reinforcement learning. Among them, the feedback optimization technique generates a model optimization strategy by combining user interaction feedback data and real-time performance monitoring, and continuously improves the accuracy of speech recognition and the accuracy of semantic understanding.
[0158] It can be seen that, based on the user's voice signal, context text information, and environmental sensor data, multi-modal data fusion is performed to obtain a fused multi-modal feature vector; the multi-modal feature vector is input into a speech recognition model based on an adaptive deep neural network for speech-to-text processing to obtain a text output; the text output is input into a semantic parsing model based on a graph neural network for context semantic analysis and user intent recognition to obtain the semantic representation of the user intent and its confidence score; according to the semantic representation and its confidence score, the parameters of the speech recognition model and the semantic parsing model are dynamically adjusted using a feedback optimization technique based on reinforcement learning to generate a model optimization strategy, thereby improving the accuracy of speech recognition and the semantic understanding ability of natural language processing.
[0159] An embodiment of the present invention also provides an electronic device, including a memory and a processor, where a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0160] Specifically, the above electronic device may further include a transmission device and an input / output device, where the transmission device is connected to the above processor, and the input / output device is connected to the above processor.
[0161] Specifically, in this embodiment, the above processor may be configured to execute the following steps through a computer program:
[0162] S201, perform multi-modal data fusion based on the user's voice signal, context text information, and environmental sensor data. Among them, the multi-modal data fusion adopts a multi-modal alignment network based on deep learning, and through cross-modal attention technology, extracts the correlation relationship between the acoustic features and context semantic features in the voice signal to obtain a fused multi-modal feature vector;
[0163] S202, input the multi-modal feature vector into a speech recognition model based on an adaptive deep neural network for speech-to-text processing. Among them, the speech recognition model adjusts the recognition parameters in real time by combining a dynamic noise suppression algorithm and context awareness technology to adapt to different environmental noises and user pronunciation habits to obtain a text output;
[0164] S203, input the text output into a semantic parsing model based on a graph neural network for context semantic analysis and user intent recognition. Among them, the semantic parsing model captures the implicit semantic relationship and long-distance dependence in the text through dynamic graph structure learning and multi-hop reasoning technology to obtain the semantic representation of the user intent and its confidence score;
[0165] S204. According to the semantic representation and its confidence score, use the feedback optimization technology based on reinforcement learning to dynamically adjust the parameters of the speech recognition model and the semantic parsing model. Among them, the feedback optimization technology generates a model optimization strategy by combining user interaction feedback data and real-time performance monitoring, and continuously improves the speech recognition accuracy and semantic understanding accuracy.
[0166] It can be seen that according to the user's voice signal, context text information and environmental sensor data, multi-modal data fusion is performed to obtain a fused multi-modal feature vector; the multi-modal feature vector is input into a speech recognition model based on an adaptive deep neural network for speech-to-text processing to obtain a text output; the text output is input into a semantic parsing model based on a graph neural network for context semantic analysis and user intention recognition to obtain a semantic representation of the user intention and its confidence score; according to the semantic representation and its confidence score, use the feedback optimization technology based on reinforcement learning to dynamically adjust the parameters of the speech recognition model and the semantic parsing model, and generate a model optimization strategy, so as to improve the accuracy of speech recognition and the semantic understanding ability of natural language processing.
[0167] The above has detailed the structure, features and effects of the present invention according to the illustrated embodiments. The above is only a preferred embodiment of the present invention, but the present invention is not limited to the scope of implementation shown in the drawings. Any changes made according to the concept of the present invention, or equivalent embodiments modified into equivalent changes, still within the spirit covered by the specification and the drawings, shall be within the protection scope of the present invention.
Claims
1. A method for integrating speech recognition and natural language processing, characterized in that: The method comprises: Performing multimodal data fusion according to the user's voice signal, contextual text information and environmental sensor data, wherein the multimodal data fusion adopts a multimodal alignment network based on deep learning, and extracts the correlation between the acoustic features in the voice signal and the contextual semantic features through cross-modal attention technology to obtain a fused multimodal feature vector; Inputting the multimodal feature vector into a speech recognition model based on an adaptive deep neural network to perform speech-to-text processing, wherein the speech recognition model adjusts recognition parameters in real time to adapt to different environmental noises and user pronunciation habits by combining a dynamic noise suppression algorithm and context-aware technology, and obtains text output; Input the text output into a semantic parsing model based on a graph neural network to perform contextual semantic analysis and user intent recognition, wherein the semantic parsing model captures implicit semantic relationships and long-distance dependencies in the text through dynamic graph structure learning and multi-hop reasoning technology to obtain a semantic representation of the user's intent and its confidence score; According to the semantic representation and its confidence score, the parameters of the speech recognition model and the semantic parsing model are dynamically adjusted using feedback optimization technology based on reinforcement learning. The feedback optimization technology generates a model optimization strategy by combining user interaction feedback data and real-time performance monitoring, thereby continuously improving the speech recognition accuracy and semantic understanding accuracy.
2. The method according to claim 1, characterized in that The multimodal data fusion is performed according to the user's voice signal, contextual text information and environmental sensor data, wherein the multimodal data fusion adopts a multimodal alignment network based on deep learning, and extracts the correlation between the acoustic features and the contextual semantic features in the voice signal through cross-modal attention technology to obtain a fused multimodal feature vector, including: According to the user's voice signal, contextual text information and environmental sensor data, a data collection framework based on edge computing is used to obtain multimodal data in real time. The data is filtered and formatted through an adaptive data cleaning algorithm to generate a preliminary standardized data set. For the preliminary standardized data set, a feature extraction method based on deep learning is used to extract the acoustic features in the speech signal and the semantic features in the contextual text information. The multi-head attention technology is used to capture the feature expressions of different modalities and generate preliminary feature representations. For the preliminary feature representation, a feature fusion method based on cross-modal attention technology is used to weightedly fuse acoustic features and semantic features. The correlation between acoustic features and semantic features is captured through context-aware technology to generate a preliminary fused feature representation. For the preliminary fused feature representation, a feature integration method based on matrix decomposition is adopted to map the fused features into a unified feature space. The dynamic weight adjustment technology is used to ensure the timeliness and consistency of the feature vector to generate the final multimodal feature vector.
3. The method according to claim 2, characterized in that The multimodal feature vector is input into a speech recognition model based on an adaptive deep neural network to perform speech-to-text processing, wherein the speech recognition model adjusts recognition parameters in real time to adapt to different environmental noises and user pronunciation habits by combining a dynamic noise suppression algorithm and context-aware technology to obtain text output, including: For the speech signal in the multimodal feature vector, a preprocessing method based on a dynamic noise suppression algorithm is used. In combination with the background noise information in the environmental sensor data, the noise suppression parameters are adjusted in real time. Through the adaptive filtering technology, a preliminary denoised speech signal is generated. For the preliminary denoised speech signal, a speech recognition model based on an adaptive deep neural network is used to convert speech to text in combination with the semantic features in the contextual text information. Through context-aware technology, the recognition parameters are dynamically adjusted to adapt to the user's pronunciation habits and generate preliminary text output; For the preliminary text output, a correction method based on a language model is adopted, combining the contextual semantic features and grammatical rules to optimize the recognition results. The dynamic threshold adjustment technology is used to ensure the accuracy and consistency of the recognition results and generate the final text output.
4. The method according to claim 3, characterized in that The text output is input into a semantic parsing model based on a graph neural network to perform contextual semantic analysis and user intent recognition, wherein the semantic parsing model captures implicit semantic relationships and long-distance dependencies in the text through dynamic graph structure learning and multi-hop reasoning technology to obtain a semantic representation of the user intent and its confidence score, including: For text output, a semantic parsing model based on graph neural network is used to abstract the words in the text into nodes in the graph structure, and the semantic relationship between words into edges. Through dynamic graph structure learning technology, a preliminary semantic graph structure is generated; For the preliminary semantic graph structure, a semantic analysis method based on multi-hop reasoning technology is used to capture the implicit semantic relations and long-distance dependencies in the text. The selection of reasoning paths is optimized through attention technology to generate preliminary semantic representations. For the preliminary semantic representation, an intention recognition method based on a classification model is adopted to identify user intentions by combining contextual semantic features and user historical behavior data. The accuracy and reliability of intention recognition are calculated through confidence scoring technology to generate a semantic representation of user intentions and its confidence score.
5. The method according to claim 4, characterized in that The method dynamically adjusts the parameters of the speech recognition model and the semantic parsing model according to the semantic representation and its confidence score by using the feedback optimization technology based on reinforcement learning, wherein the feedback optimization technology generates a model optimization strategy by combining user interaction feedback data and real-time performance monitoring to continuously improve the speech recognition accuracy and semantic understanding accuracy, including: According to the semantic representation and its confidence score, a feedback data collection method based on user interaction is adopted, combined with the user's confirmation or correction behavior of the recognition result, to generate user interaction feedback data. Through lightweight data caching technology, the real-time and continuity of feedback data are ensured. For the running status of speech recognition model and semantic parsing model, an analysis method based on real-time performance monitoring is adopted, combining recognition accuracy, response time and user satisfaction indicators to generate performance monitoring reports. The dynamic threshold adjustment technology is used to ensure the accuracy and timeliness of monitoring results. For user interaction feedback data and performance monitoring reports, we use a reinforcement learning-based optimization strategy generation method, combined with the parameter space of speech recognition and semantic parsing models, to dynamically adjust model parameters. Through a multi-objective optimization algorithm, we balance recognition accuracy and response time to generate a preliminary optimization strategy. For the preliminary optimization strategy, a parameter adjustment method based on an adaptive learning algorithm is adopted. In combination with real-time performance monitoring data, the parameters of the speech recognition model and the semantic parsing model are dynamically adjusted. Through feedback correction technology, the model performance is continuously optimized to generate the final optimization strategy.
6. A speech recognition and natural language processing integrated system, characterized in that: The system comprises: A fusion module is used to perform multimodal data fusion based on the user's voice signal, contextual text information and environmental sensor data, wherein the multimodal data fusion adopts a multimodal alignment network based on deep learning, and extracts the correlation between the acoustic features in the voice signal and the contextual semantic features through cross-modal attention technology to obtain a fused multimodal feature vector; A processing module, used for inputting the multimodal feature vector into a speech recognition model based on an adaptive deep neural network to perform speech-to-text processing, wherein the speech recognition model adjusts recognition parameters in real time to adapt to different environmental noises and user pronunciation habits by combining a dynamic noise suppression algorithm and context-aware technology, and obtains text output; A recognition module is used to input the text output into a semantic parsing model based on a graph neural network to perform contextual semantic analysis and user intent recognition, wherein the semantic parsing model captures implicit semantic relationships and long-distance dependencies in the text through dynamic graph structure learning and multi-hop reasoning technology to obtain a semantic representation of the user's intent and its confidence score; An adjustment module is used to dynamically adjust the parameters of the speech recognition model and the semantic parsing model according to the semantic representation and its confidence score using a feedback optimization technology based on reinforcement learning, wherein the feedback optimization technology generates a model optimization strategy by combining user interaction feedback data and real-time performance monitoring to continuously improve the speech recognition accuracy and semantic understanding accuracy.
7. The system according to claim 6, characterized in that The fusion module is specifically used for: According to the user's voice signal, contextual text information and environmental sensor data, a data collection framework based on edge computing is used to obtain multimodal data in real time. The data is filtered and formatted through an adaptive data cleaning algorithm to generate a preliminary standardized data set. For the preliminary standardized data set, a feature extraction method based on deep learning is used to extract the acoustic features in the speech signal and the semantic features in the contextual text information. The multi-head attention technology is used to capture the feature expressions of different modalities and generate preliminary feature representations. For the preliminary feature representation, a feature fusion method based on cross-modal attention technology is used to weightedly fuse acoustic features and semantic features. The correlation between acoustic features and semantic features is captured through context-aware technology to generate a preliminary fused feature representation. For the preliminary fused feature representation, a feature integration method based on matrix decomposition is adopted to map the fused features into a unified feature space. The dynamic weight adjustment technology is used to ensure the timeliness and consistency of the feature vector to generate the final multimodal feature vector.
8. The system according to claim 7, characterized in that The processing module is specifically used for: For the speech signal in the multimodal feature vector, a preprocessing method based on a dynamic noise suppression algorithm is used. In combination with the background noise information in the environmental sensor data, the noise suppression parameters are adjusted in real time. Through the adaptive filtering technology, a preliminary denoised speech signal is generated. For the preliminary denoised speech signal, a speech recognition model based on an adaptive deep neural network is used to convert speech to text in combination with the semantic features in the contextual text information. Through context-aware technology, the recognition parameters are dynamically adjusted to adapt to the user's pronunciation habits and generate preliminary text output; For the preliminary text output, a correction method based on a language model is adopted, combining the contextual semantic features and grammatical rules to optimize the recognition results. The dynamic threshold adjustment technology is used to ensure the accuracy and consistency of the recognition results and generate the final text output.
9. A storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program is configured to execute the method according to any one of claims 1 to 5 when executed.
10. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Intelligent customer service intention understanding method based on text and voice information fusion
CN112287675A
Multi-field and multi-intention oral language semantic understanding model training method
CN116682419A
Multimodal fusion speech translation method, system and equipment
CN118692446A
Automatic voice processing method based on language large model and electronic equipment
CN119091864A
Customer voice analysis system based on large model
CN119541499A
Cited By
Speech recognition authentication method and system based on multi-modal features and dynamic evaluation
CN120748413A
Speech recognition authentication method and system based on multi-modal features and dynamic evaluation
CN120748413B
Self-evolution method and system of speech recognition model
CN120808760A
User classification system based on voice intention recognition
CN120977338A
User classification system based on speech intent recognition
CN120977338B