Diversified interaction system based on AIGC intelligent model
By adopting technology based on AIGC intelligent model in a multimodal interactive system, integrating environment and user behavior data, dynamically generating multimodal content, the problem of limitations in the existing technology is solved, and a personalized and real-time interactive experience is achieved.
Patent Information
- Application Number
- CN202411995190.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-02
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the multimodal interactive system in the fields of cultural tourism, study and night tours, it is difficult to fully integrate environmental and user behavior data, resulting in limited interaction effects and unable to provide an immersive and interactive experience.
Using a diversified interactive system based on the AIGC intelligent model, multimodal content that meets the needs of the scenario and users is dynamically generated through technologies such as data collection, feature extraction, classification evaluation, content generation, speech analysis, semantic analysis and speech synthesis to achieve a personalized and real-time interactive experience.
It realizes a personalized and real-time multi-modal interactive experience in museums, caves and other scenarios, improves the flexibility and accuracy of interaction, and is suitable for cultural tourism and educational applications that require highly customized interaction.
Smart Images

Figure CN119918003A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent interaction and multimodal technology, and in particular to a diversified interactive system based on an AIGC intelligent model. Background Art
[0002] With the continuous advancement of intelligent technology, multimodal interactive systems have gradually been introduced into the fields of cultural tourism, study tours, and night tours to enhance user experience, especially in special scenarios such as museums and caves. Traditional interactive systems usually rely on voice recognition or visual recognition technology, and the processing of single-modal data cannot fully integrate multiple information from the environment, user behavior, etc., making it difficult to achieve accurate scene adaptability and personalized response. In addition, the existing system lacks a deep understanding of user emotions and needs, resulting in limited interaction effects and difficulty in providing an immersive and highly interactive experience. Especially in the cultural education and tourism industries, the real-time and customization levels of intelligent responses are low, and users' emotional feedback and behavioral patterns are not effectively captured and utilized.
[0003] The diversified interactive system based on the AIGC intelligent model proposed in the present invention effectively overcomes the shortcomings of the existing technology by integrating technologies such as multimodal data collection, feature extraction, classification evaluation, content generation, speech analysis, semantic parsing and speech synthesis. The system can dynamically generate multimodal content that meets the scene and user needs based on environmental data and user behavior characteristics in specific cultural and tourism scenes such as museums and caves, and realize personalized and real-time interactive experience. Through in-depth analysis of the user's voice and emotional information, the system can adjust the generated response in real time according to feedback, which greatly improves the flexibility and accuracy of the interaction, and is particularly suitable for cultural and tourism and educational applications that require highly customized interaction. Summary of the invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a diversified interactive system based on the AIGC intelligent model to solve the shortcomings of traditional technologies in terms of interactive effects and user experience.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, the present invention provides a diversified interactive system based on an AIGC intelligent model, which comprises:
[0008] Data collection module, classification and evaluation module, content generation module, speech analysis module, semantic parsing module, speech synthesis module, feedback adjustment module;
[0009] The data collection module collects environmental data and user behavior data and performs pre-processing;
[0010] The classification and evaluation module classifies the scenes using a random forest algorithm based on the preprocessed environmental data to obtain scene classification results, and at the same time uses a support vector machine (SVM) to evaluate user status results based on the preprocessed user behavior data;
[0011] The content generation module inputs the scene classification results and the user status evaluation results into the AIGC intelligent model according to the scene classification results and the user status evaluation results to generate multimodal content that adapts to the current scene and user needs;
[0012] The speech analysis module, when the user makes a speech response to the generated multimodal content, uses automatic speech recognition ASR technology to extract the speech content input by the user and converts it into text, then uses language identification LID and speech emotion recognition SER technology to extract the language information and emotion information of the user's speech content, generates a preliminary speech analysis result, uses acoustic event classification AEC and acoustic event detection AED technology to parse the non-language information in the speech content input by the user, and generates a final speech analysis result in combination with the preliminary speech analysis result;
[0013] The semantic parsing module calls the large language model LLM to parse the semantics according to the final speech analysis result, combined with the scene classification result and the user status evaluation result, and dynamically adjusts the generated response text;
[0014] The speech synthesis module converts the adjusted response text into speech output through speech synthesis technology;
[0015] The feedback adjustment module adjusts the generated response text based on the user's voice feedback after the user receives the voice response.
[0016] As a preferred solution of the diversified interactive system based on the AIGC intelligent model described in the present invention, the environmental data and user behavior data are collected and pre-processed, and the specific steps are as follows:
[0017] Collect environmental data and user behavior data through microphones, cameras, touch screens, and keyboards;
[0018] Clean the collected environmental data and user behavior data, and process missing data through interpolation and filling;
[0019] The cleaned environmental data and user behavior data are subjected to noise reduction processing using wavelet transform;
[0020] Normalize the noise-reduced environmental data and user behavior data;
[0021] Perform feature extraction on the normalized environmental data and user behavior data.
[0022] As a preferred solution of the diversified interactive system based on the AIGC intelligent model described in the present invention, the specific steps of extracting features from the normalized environmental data and user behavior data are as follows:
[0023] Use Mel-frequency cepstral coefficient (MFCC) method to extract audio feature vectors from speech data;
[0024] Use a convolutional neural network (CNN) model to extract visual feature vectors from facial expressions and gestures in visual data;
[0025] Convert text data into word vector representation and use the BERT model to generate text embedding feature vectors;
[0026] The principal component analysis PCA is used to reduce the dimension of environmental data to obtain the environmental feature vector set, which is expressed as:
[0027] E = PCA(a,b,c);
[0028] Where E represents the environmental feature vector set, PCA represents principal component analysis, a, b and c represent location coordinate data, time stamp data and climate condition data respectively;
[0029] The audio feature vector, visual feature vector, and text embedding feature vector are mapped to the same dimension through dimension alignment in a fully connected layer;
[0030] The audio feature vector, visual feature vector, and text embedding feature vector are weightedly fused using the weighted average method to obtain the multimodal feature vector of user behavior, which is expressed as follows:
[0031] X = αF + βV + γW;
[0032] Among them, X represents the multimodal feature vector of user behavior, α, β and γ represent the weighting coefficients of audio feature vector, visual feature vector and text feature vector respectively, and F, V and W represent audio feature vector, visual feature vector and text embedding feature vector respectively.
[0033] As a preferred solution of the diversified interactive system based on the AIGC intelligent model described in the present invention, the random forest algorithm is used to classify the scene based on the pre-processed environmental data to obtain the scene classification result, and the support vector machine SVM is used to evaluate the user status result based on the pre-processed user behavior data. The specific steps are as follows:
[0034] For the environmental feature vector set E, the random forest RF algorithm is used to perform the scene classification task;
[0035] By integrating multiple decision trees built based on different subsets of environmental data, the scene classification result A is obtained;
[0036] Use support vector machine (SVM) to evaluate user emotions and intentions;
[0037] Input the user behavior multimodal feature vector X into the support vector machine SVM model to obtain the user status evaluation result, which is expressed as:
[0038] d = sign(·X+b1);
[0039] f = sign(·X+b2);
[0040] Wherein, d represents the emotion evaluation result, f represents the intention evaluation result, represents the weight vector of emotion evaluation, represents the weight vector of intention evaluation, b1 and b2 are the bias terms of emotion and intention respectively, and sign represents the sign function;
[0041] The user state evaluation result refers to a collection of emotion evaluation results and intention evaluation results.
[0042] As a preferred solution of the diversified interactive system based on the AIGC smart model of the present invention, wherein: according to the scene classification results and the user status evaluation results, the scene classification results and the user status evaluation results are input into the AIGC smart model to generate multimodal content that adapts to the current scene and user needs, the specific steps are:
[0043] The scene classification result A and the user state evaluation result (d, f) are weightedly fused to generate the input vector, which is expressed as:
[0044] Z = A + (d, f);
[0045] Where Z represents the weighted input vector, and represents the weighting coefficients of the scene classification result and the user state evaluation result respectively;
[0046] Input the weighted input vector into the generator of the AIGC smart model;
[0047] First, the generator generates text content based on the input vector using the BERT model, expressed as:
[0048] = BERT(Z;);
[0049] Among them, represents the generated text content, and represents the parameters of the BERT model;
[0050] Secondly, the generator uses the TTS model to convert text content into speech content. The expression is:
[0051] =TTS(;);
[0052] Among them, represents the generated speech content, represents the parameters of the TTS model
[0053] Next, the generator uses the generative adversarial network GAN model to generate video content, expressed as:
[0054] =GAN(Z; );
[0055] Among them, represents the generated video content, represents the parameters of the GAN model;
[0056] Finally, the generated text, voice, and video are combined into a multimodal content, expressed as:
[0057] Y=[,,];
[0058] Among them, Y represents the complete multimodal content that adapts to the current scenario and user needs.
[0059] As a preferred solution of the diversified interactive system based on the AIGC intelligent model described in the present invention, when the user makes a voice response to the generated multimodal content, the automatic speech recognition ASR technology is used to extract the voice content input by the user and convert it into text, and then the language identification LID and speech emotion recognition SER technology are used to extract the language information and emotion information of the user's voice content to generate a preliminary speech analysis result. The specific steps are:
[0060] When a user responds to the generated multimodal content, the user’s voice input is first converted into text using automatic speech recognition (ASR) technology;
[0061] Identify the language category of the user's voice input through language identification LID technology;
[0062] After speech content recognition, speech emotion recognition (SER) technology is used to analyze the emotional information in the user's speech;
[0063] Introduce the speech analysis formula to generate preliminary speech analysis results. The expression is:
[0064] R=σ(ε·O(t)+ζ·P(t)+η·Q(t)+δ·ψ(h,n,m));
[0065] Where R represents the preliminary speech analysis result, σ represents the activation function, ε, ζ and η represent the weighting coefficients of ASR, LID and SER respectively, h, n and m represent the text information output by ASR, the language information output by LID and the emotional information output by SER respectively, ψ represents the fusion function, O(t) is the text signal extracted by ASR, which represents the text information converted from the speech at time t, P(t) is the language signal extracted by LID, which represents the emotional state recognized from the speech at time t, and Q(t) is the emotional signal extracted by SER, which represents the emotional state recognized from the speech at time t.
[0066] As a preferred solution of the diversified interactive system based on the AIGC intelligent model described in the present invention, the acoustic event classification AEC and acoustic event detection AED technology are used to analyze the non-language information in the voice content input by the user, and the final voice analysis result is generated in combination with the preliminary voice analysis result. The specific steps are as follows:
[0067] For the original audio signal, the time-frequency conversion is performed through the Mel-frequency cepstral coefficient MFCC to convert the audio signal into a spectrum graph;
[0068] Use convolutional neural network (CNN) to extract features and classify spectrograms to identify different acoustic events in audio.
[0069] Performing time sequence detection on the audio signal to identify the occurrence time, duration and intensity of each event;
[0070] Use recurrent neural network (RNN) to further analyze time series information and extract non-language signal features;
[0071] The preliminary speech analysis results are combined with the non-language signal features obtained through acoustic event classification and detection to generate the final speech analysis results, which are expressed as:
[0072] r=σ(ε·O(t)+ζ·P(t)+η·Q(t)+δ·ψ(h,n,m)+λ·(N(t)+L(t)+
[0073] D(t));
[0074] Among them, r represents the final speech analysis result, N(t) is the background noise intensity, which means that at time t, the intensity of the background noise is calculated based on the input audio signal, L(t) is the laughter intensity, which means the obviousness of the laughter in the speech at time t, D(t) is the speech pause duration, which means the pause duration in the speech is calculated at time t, and λ represents the weighting coefficient of the non-language signal.
[0075] As a preferred solution of the diversified interactive system based on the AIGC intelligent model described in the present invention, the specific steps are as follows: according to the final speech analysis result, combined with the scene classification result and the user status evaluation result, the large language model LLM is called to parse the semantics and generate the response text.
[0076] Use the Qwen2 basic language model as the basis of the large language model LLM;
[0077] The final speech analysis results, scene classification results and user status evaluation results are weighted and fused to form a comprehensive input vector;
[0078] Pass the synthesized input vector as input to the large language model LLM, and use the large language model LLM internal encoder to convert the synthesized input vector into a contextual embedding;
[0079] After obtaining the context embedding, the large language model LLM uses its internal neural network structure to deeply process the context embedding and obtain the hidden state;
[0080] Based on the hidden state, the large language model LLM generates the response text through the internal decoder, which is expressed as:
[0081] =Decoder LLM (;);
[0082] Among them, represents the generated response text, Decoder LLM Represents the large language model LLM through the internal decoder and represents the parameter set of the decoder.
[0083] As a preferred solution of the diversified interactive system based on the AIGC intelligent model of the present invention, the adjusted response text is converted into speech output through speech synthesis technology, and the specific steps are as follows:
[0084] Select TTS, an end-to-end speech synthesis method based on enhanced variational inference and normalized flow;
[0085] The generated response text is converted into input text features through word segmentation and word embedding processes;
[0086] For the input text features, the potential speech features are generated through the enhanced variational inference network;
[0087] Use normalization flow to transform the latent features to adjust the pitch, rhythm and intonation of the speech;
[0088] Based on the latent features adjusted by the normalized flow, speech waveforms with diverse rhythms and pitches are generated through a random duration predictor;
[0089] The generated speech waveform is output through an audio playback device.
[0090] As a preferred solution of the diversified interactive system based on the AIGC intelligent model described in the present invention, when the user receives a voice response, the generated response text is adjusted based on the user's voice feedback. The specific steps are:
[0091] After the user receives the voice response generated by the speech synthesis technology, the user can directly comment and ask questions to provide voice feedback and receive the user's voice feedback information;
[0092] Based on the received voice feedback, the user's voice feedback information is converted into text information using automatic speech recognition ASR technology;
[0093] Combine language identification LID and speech emotion recognition SER technology to extract language and emotion information from user feedback;
[0094] Generate preliminary speech feedback analysis results through speech analysis formula;
[0095] Perform acoustic event classification (AEC) and acoustic event detection (AED) on speech feedback to extract non-language information, and generate the final speech feedback analysis results in combination with the preliminary speech feedback analysis results;
[0096] The final feedback speech analysis results, scene classification results, and user status assessment results are input into the large language model (LLM) for semantic analysis to adjust the generated response text.
[0097] The beneficial effects of the present invention are as follows: the present invention realizes accurate response to user needs through a diversified interactive system based on the AIGC intelligent model, combining environmental data and user behavior data; through data collection and preprocessing steps, the clarity and effectiveness of the data are guaranteed; the scene classification and user state evaluation module is used, combined with random forest and support vector machine algorithms, to accurately identify user emotions and intentions, thereby generating personalized, multimodal content; the speech analysis and semantic parsing module adjusts the generated response text in real time to ensure that the interactive content is consistent with the user's emotional state; the speech synthesis technology makes the response voice more diverse through an end-to-end TTS method; finally, the feedback adjustment module is optimized according to user feedback, so that the system continuously improves the response accuracy and user experience; in summary, the present invention provides a flexible and efficient interactive method through data fusion and intelligent analysis, which significantly improves the adaptability and intelligence level of the interactive system. BRIEF DESCRIPTION OF THE DRAWINGS
[0098] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0099] Figure 1 This is a schematic diagram of a diversified interactive system based on the AIGC intelligent model in Example 1.
[0100] Figure 2 This is a schematic diagram of generating multimodal content based on the AIGC intelligent model in Example 1. DETAILED DESCRIPTION
[0101] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.
[0102] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0103] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.
[0104] Example 1, reference Figure 1 and Figure 2 , which is the first embodiment of the present invention, and provides a diversified interactive system based on the AIGC intelligent model, comprising the following steps:
[0105] Data collection module, classification and evaluation module, content generation module, speech analysis module, semantic parsing module, speech synthesis module, feedback adjustment module;
[0106] The data collection module collects environmental data and user behavior data and performs pre-processing;
[0107] The classification and evaluation module classifies the scenes using a random forest algorithm based on the preprocessed environmental data to obtain scene classification results, and at the same time uses a support vector machine (SVM) to evaluate user status results based on the preprocessed user behavior data;
[0108] The content generation module inputs the scene classification results and the user status evaluation results into the AIGC intelligent model according to the scene classification results and the user status evaluation results to generate multimodal content that adapts to the current scene and user needs;
[0109] The speech analysis module, when the user makes a speech response to the generated multimodal content, uses automatic speech recognition ASR technology to extract the speech content input by the user and converts it into text, then uses language identification LID and speech emotion recognition SER technology to extract the language information and emotion information of the user's speech content, generates a preliminary speech analysis result, uses acoustic event classification AEC and acoustic event detection AED technology to parse the non-language information in the speech content input by the user, and generates a final speech analysis result in combination with the preliminary speech analysis result;
[0110] The semantic parsing module calls the large language model LLM to parse the semantics according to the final speech analysis result, combined with the scene classification result and the user status evaluation result, and dynamically adjusts the generated response text;
[0111] The speech synthesis module converts the adjusted response text into speech output through speech synthesis technology;
[0112] The feedback adjustment module adjusts the generated response text based on the user's voice feedback after the user receives the voice response.
[0113] Collect environmental data and user behavior data and perform pre-processing;
[0114] Collect environmental data and user behavior data through microphones, cameras, touch screens, and keyboards;
[0115] Environmental data are location coordinate data, time stamp data and climate condition data;
[0116] User behavior data is voice data, visual data, and text data;
[0117] Through multi-channel data collection, it is possible to fully capture multimodal information about user behavior and external environment, providing a rich and diverse data source for subsequent analysis;
[0118] Clean the collected environmental data and user behavior data, and process missing data through interpolation and filling;
[0119] The cleaned environmental data and user behavior data are subjected to noise reduction processing using wavelet transform;
[0120] Normalize the noise-reduced environmental data and user behavior data;
[0121] By cleaning and processing the data, the noise and anomalies in the original data are reduced, the quality and validity of the data are improved, and the accuracy of subsequent feature extraction and analysis is ensured;
[0122] Extract features from normalized environmental data and user behavior data;
[0123] Use Mel-frequency cepstral coefficient (MFCC) method to extract audio feature vectors from speech data;
[0124] Use a convolutional neural network (CNN) model to extract visual feature vectors from facial expressions and gestures in visual data;
[0125] Convert text data into word vector representation and use the BERT model to generate text embedding feature vectors;
[0126] The principal component analysis PCA is used to reduce the dimension of environmental data to obtain the environmental feature vector set, which is expressed as:
[0127] E = PCA(a,b,c);
[0128] Where E represents the environmental feature vector set, PCA represents principal component analysis, a, b and c represent location coordinate data, time stamp data and climate condition data respectively;
[0129] The audio feature vector, visual feature vector, and text embedding feature vector are mapped to the same dimension through dimension alignment in a fully connected layer;
[0130] The audio feature vector, visual feature vector, and text embedding feature vector are weightedly fused using the weighted average method to obtain the multimodal feature vector of user behavior, which is expressed as follows:
[0131] X = αF + βV + γW;
[0132] Among them, X represents the obtained multimodal feature vector of user behavior, α, β and γ represent the weighting coefficients of audio feature vector, visual feature vector and text feature vector respectively, F, V and W represent audio feature vector, visual feature vector and text embedding feature vector respectively;
[0133] By weighted fusion of features from different modalities, we can fully integrate information from various types of data and ensure that the system can make more accurate responses and decisions based on multiple input dimensions.
[0134] Based on the preprocessed environmental data, the random forest algorithm is used to classify the scenes to obtain the scene classification results. At the same time, the support vector machine (SVM) is used to evaluate the user status results based on the preprocessed user behavior data.
[0135] For the environmental feature vector set E, the random forest RF algorithm is used to perform the scene classification task;
[0136] By integrating multiple decision trees built based on different subsets of environmental data, the scene classification result A is obtained;
[0137] The scenario categories are cultural tourism, caves, museums, research study travel and industry professional travel;
[0138] Use support vector machine (SVM) to evaluate user emotions and intentions;
[0139] Input the user behavior multimodal feature vector X into the support vector machine SVM model to obtain the user status evaluation result, which is expressed as:
[0140] d = sign(·X+b1);
[0141] f = sign(·X+b2);
[0142] Wherein, d represents the emotion evaluation result, f represents the intention evaluation result, represents the weight vector of emotion evaluation, represents the weight vector of intention evaluation, b1 and b2 are the bias terms of emotion and intention respectively, and sign represents the sign function;
[0143] The user state evaluation result refers to the collection of emotion evaluation results and intention evaluation results.
[0144] According to the scene classification results and user status evaluation results, the scene classification results and user status evaluation results are input into the AIGC intelligent model to generate multimodal content that adapts to the current scene and user needs;
[0145] The scene classification result A and the user state evaluation result (d, f) are weightedly fused to generate the input vector, which is expressed as:
[0146] Z = A + (d, f);
[0147] Where Z represents the weighted input vector, and represents the weighting coefficients of the scene classification result and the user state evaluation result respectively;
[0148] Input the weighted input vector into the generator of the AIGC smart model;
[0149] First, the generator generates text content based on the input vector using the BERT model, expressed as:
[0150] = BERT(Z;);
[0151] Among them, represents the generated text content, and represents the parameters of the BERT model;
[0152] Secondly, the generator uses the TTS model to convert text content into speech content. The expression is:
[0153] =TTS(;);
[0154] Among them, represents the generated speech content, represents the parameters of the TTS model
[0155] Next, the generator uses the generative adversarial network GAN model to generate video content, expressed as:
[0156] =GAN(Z; );
[0157] Among them, represents the generated video content, represents the parameters of the GAN model;
[0158] Finally, the generated text, voice, and video are combined into a multimodal content, expressed as:
[0159] Y=[,,];
[0160] Among them, Y represents the complete multimodal content adapted to the current scenario and user needs;
[0161] Through dynamically generated multimodal content, the response content can be adjusted in real time according to user needs and scenario changes, ensuring the smoothness and naturalness of the interaction.
[0162] When the user responds to the generated multimodal content by voice, the automatic speech recognition ASR technology is used to extract the voice content input by the user and convert it into text. Then, the language identification LID and speech emotion recognition SER technologies are used to extract the language information and emotion information of the user's voice content to generate preliminary voice analysis results. The acoustic event classification AEC and acoustic event detection AED technologies are used to parse the non-language information in the voice content input by the user. Combined with the preliminary voice analysis results, the final voice analysis results are generated.
[0163] When a user responds to the generated multimodal content, the user’s voice input is first converted into text using automatic speech recognition (ASR) technology;
[0164] Identify the language category of the user's voice input through language identification LID technology;
[0165] The language categories are Chinese, English and French;
[0166] After speech content recognition, speech emotion recognition (SER) technology is used to analyze the emotional information in the user's speech;
[0167] The emotional messages are joy, anger, and sadness;
[0168] Introduce the speech analysis formula to generate preliminary speech analysis results. The expression is:
[0169] R=σ(ε·O(t)+ζ·P(t)+η·Q(t)+δ·ψ(h,n,m));
[0170] Where R represents the preliminary speech analysis result, σ represents the activation function, ε, ζ and η represent the weighting coefficients of ASR, LID and SER respectively, h, n and m represent the text information output by ASR, the language information output by LID and the emotional information output by SER respectively, ψ represents the fusion function, O(t) is the text signal extracted by ASR, which represents the text information converted from the speech at time t, P(t) is the language signal extracted by LID, which represents the emotional state recognized from the speech at time t, Q(t) is the emotional signal extracted by SER, which represents the emotional state recognized from the speech at time t;
[0171] For the original audio signal, the time-frequency conversion is performed through the Mel-frequency cepstral coefficient MFCC to convert the audio signal into a spectrum graph;
[0172] Use convolutional neural network (CNN) to extract features and classify spectrograms to identify different acoustic events in audio.
[0173] Acoustic events are sound bursts, prolonged silences, changes in intonation, and increased speech rate;
[0174] Performing time sequence detection on the audio signal to identify the occurrence time, duration and intensity of each event;
[0175] Use recurrent neural network (RNN) to further analyze time series information and extract non-language signal features;
[0176] Nonverbal signal features were background noise intensity, laughter, and speech pause length;
[0177] The preliminary speech analysis results are combined with the non-language signal features obtained through acoustic event classification and detection to generate the final speech analysis results, which are expressed as:
[0178] r=σ(ε·O(t)+ζ·P(t)+η·Q(t)+δ·ψ(h,n,m)+λ·(N(t)+L(t)+
[0179] D(t));
[0180] Where r represents the final speech analysis result, N(t) is the background noise intensity, which means that at time t, the intensity of the background noise is calculated based on the input audio signal, L(t) is the laughter intensity, which means the obviousness of the laughter in the speech at time t, D(t) is the pause duration of the speech, which means the pause duration in the speech is calculated at time t, and λ represents the weighting coefficient of the non-language signal;
[0181] This step can comprehensively analyze the language and emotional information in the user's voice, so as to more accurately understand the user's intentions and emotions, and provide detailed voice feedback for the generation of response content.
[0182] Based on the final speech analysis results, combined with the scene classification results and user status assessment results, the large language model LLM is called to parse the semantics and generate the response text;
[0183] Use the Qwen2 basic language model as the basis of the large language model LLM;
[0184] The final speech analysis results, scene classification results and user status evaluation results are weighted and fused to form a comprehensive input vector, which is expressed as:
[0185]
[0186] Among them, represents the comprehensive input vector, and represent the weighting coefficients of the final speech analysis result, scene classification result and user status assessment result respectively;
[0187] The synthesized input vector is passed as input to the large language model LLM, and the synthesized input vector is converted into a context embedding using the internal encoder of the large language model LLM, expressed as:
[0188] =Encoder LLM (;);
[0189] Among them, represents context embedding, Encoder LLM Represents the encoder function inside the large language model LLM, which represents the parameter set of the encoder;
[0190] After obtaining the context embedding, the large language model LLM uses its internal neural network structure to deeply process the context embedding and obtain the hidden state, which is expressed as:
[0191] = Transformer(;);
[0192] Among them, represents the hidden state, Transformer represents a multi-layer neural network based on the Transformer architecture, and represents the parameter set of the Transformer model;
[0193] Based on the hidden state after deep processing, the large language model LLM generates the response text through the internal decoder, which is expressed as:
[0194] =Decoder LLM (;);
[0195] Among them, represents the generated response text, DecoderLLM Represents the internal decoder of the large language model LLM and the parameter set of the decoder.
[0196] Through the semantic parsing capabilities of the large language model (LLM), text content can be generated that responds to the user's current needs and emotional state.
[0197] The adjusted response text is converted into speech output through speech synthesis technology;
[0198] Select TTS, an end-to-end speech synthesis method based on enhanced variational inference and normalized flow;
[0199] The generated response text is converted into input text features through word segmentation and word embedding processes;
[0200] For the input text features, the potential speech features are generated through the enhanced variational inference network;
[0201] Use normalization flow to transform the latent features to adjust the pitch, rhythm and intonation of the speech;
[0202] Based on the latent features adjusted by the normalized flow, speech waveforms with diverse rhythms and pitches are generated through a random duration predictor;
[0203] Output the generated speech waveform through an audio playback device;
[0204] By flexibly adjusting the pitch and rhythm of the voice, the voice response can be made more in line with the user's emotions and needs, enhancing the naturalness and affinity of the voice.
[0205] When the user receives a voice response, the generated response text is adjusted based on the user's voice feedback;
[0206] After the user receives the voice response generated by the speech synthesis technology, the user can directly comment and ask questions to provide voice feedback and receive the user's voice feedback information;
[0207] Based on the received voice feedback, the user's voice feedback information is converted into text information using automatic speech recognition ASR technology;
[0208] Combine language identification LID and speech emotion recognition SER technology to extract language and emotion information from user feedback;
[0209] Generate preliminary speech feedback analysis results through speech analysis formula;
[0210] Perform acoustic event classification (AEC) and acoustic event detection (AED) on speech feedback to extract non-language information, and generate the final speech feedback analysis results in combination with the preliminary speech feedback analysis results;
[0211] The final feedback speech analysis results, scene classification results, and user status assessment results are input into the large language model (LLM) for semantic analysis, and the generated response text is adjusted;
[0212] Through users' instant feedback, not only can the response text be adjusted dynamically, but the system's response strategy can also be continuously optimized, making the interaction more personalized and accurate.
[0213] In summary, the present invention achieves accurate response to user needs through: a diversified interactive system based on the AIGC intelligent model, combining environmental data and user behavior data, ensuring the clarity and effectiveness of the data through data collection and preprocessing steps, and using scene classification and user state evaluation modules, combined with random forest and support vector machine algorithms, to accurately identify user emotions and intentions, thereby generating personalized, multimodal content; a speech analysis and semantic parsing module to adjust the generated response text in real time to ensure that the interactive content is consistent with the user's emotional state; speech synthesis technology uses an end-to-end TTS method to make the response voice more diverse; finally, the feedback adjustment module is optimized according to user feedback, so that the system continuously improves response accuracy and user experience; in summary, the present invention provides a flexible and efficient interactive method through data fusion and intelligent analysis, which significantly improves the adaptability and intelligence level of the interactive system.
[0214] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A diversified interactive system based on the AIGC intelligent model, characterized by: Including data collection module, classification and evaluation module, content generation module, speech analysis module, semantic parsing module, speech synthesis module, feedback adjustment module; The data collection module collects environmental data and user behavior data and performs pre-processing; The classification and evaluation module classifies the scenes using a random forest algorithm based on the preprocessed environmental data to obtain scene classification results, and at the same time uses a support vector machine (SVM) to evaluate user status results based on the preprocessed user behavior data; The content generation module inputs the scene classification results and the user status evaluation results into the AIGC intelligent model according to the scene classification results and the user status evaluation results to generate multimodal content that adapts to the current scene and user needs; The speech analysis module, when the user makes a speech response to the generated multimodal content, uses automatic speech recognition ASR technology to extract the speech content input by the user and converts it into text, then uses language identification LID and speech emotion recognition SER technology to extract the language information and emotion information of the user's speech content, generates a preliminary speech analysis result, uses acoustic event classification AEC and acoustic event detection AED technology to parse the non-language information in the speech content input by the user, and generates a final speech analysis result in combination with the preliminary speech analysis result; The semantic parsing module, based on the final speech analysis result, combines the scene classification result and the user status evaluation result, calls the large language model LLM to parse the semantics and generate a response text; The speech synthesis module converts the adjusted response text into speech output through speech synthesis technology; The feedback adjustment module adjusts the generated response text based on the user's voice feedback after the user receives the voice response.
2. The diversified interactive system based on the AIGC intelligent model as claimed in claim 1, characterized in that: The environmental data and user behavior data are collected and preprocessed, and the specific steps are as follows: Collect environmental data and user behavior data through microphones, cameras, touch screens, and keyboards; Clean the collected environmental data and user behavior data, and process missing data through interpolation and filling; The cleaned environmental data and user behavior data are subjected to noise reduction processing using wavelet transform; Normalize the noise-reduced environmental data and user behavior data; Perform feature extraction on the normalized environmental data and user behavior data.
3. The diversified interactive system based on the AIGC intelligent model as claimed in claim 2, characterized in that: The specific steps of extracting features from the normalized environment data and user behavior data are as follows: Use Mel-frequency cepstral coefficient (MFCC) method to extract audio feature vectors from speech data; Use a convolutional neural network (CNN) model to extract visual feature vectors from facial expressions and gestures in visual data; Convert text data into word vector representation and use the BERT model to generate text embedding feature vectors; The principal component analysis PCA is used to reduce the dimension of environmental data to obtain the environmental feature vector set, which is expressed as: E = PCA(a,b,c); Where E represents the environmental feature vector set, PCA represents principal component analysis, a, b and c represent location coordinate data, time stamp data and climate condition data respectively; The audio feature vector, visual feature vector, and text embedding feature vector are mapped to the same dimension through dimension alignment in a fully connected layer; The audio feature vector, visual feature vector, and text embedding feature vector are weightedly fused using the weighted average method to obtain the multimodal feature vector of user behavior, which is expressed as follows: X = αF + βV + γW; Among them, X represents the multimodal feature vector of user behavior, α, β and γ represent the weighting coefficients of audio feature vector, visual feature vector and text feature vector respectively, and F, V and W represent audio feature vector, visual feature vector and text embedding feature vector respectively.
4. The diversified interactive system based on the AIGC intelligent model as claimed in claim 3, characterized in that: The random forest algorithm is used to classify the scenes based on the preprocessed environmental data to obtain the scene classification results, and the support vector machine SVM is used to evaluate the user status results based on the preprocessed user behavior data. The specific steps are: For the environmental feature vector set E, the random forest RF algorithm is used to perform the scene classification task; By integrating multiple decision trees built based on different subsets of environmental data, the scene classification result A is obtained; Use support vector machine (SVM) to evaluate user emotions and intentions; Input the user behavior multimodal feature vector X into the support vector machine SVM model to obtain the user status evaluation result, which is expressed as: d = sign(·X+b1); f = sign(·X+b2); Wherein, d represents the emotion evaluation result, f represents the intention evaluation result, represents the weight vector of emotion evaluation, represents the weight vector of intention evaluation, b1 and b2 are the bias terms of emotion and intention respectively, and sign represents the sign function; The user state evaluation result refers to a collection of emotion evaluation results and intention evaluation results.
5. The diversified interactive system based on the AIGC intelligent model as claimed in claim 4, characterized in that: According to the scene classification results and the user status evaluation results, the scene classification results and the user status evaluation results are input into the AIGC intelligent model to generate multimodal content that adapts to the current scene and user needs. The specific steps are as follows: The scene classification result A and the user state evaluation result (d, f) are weightedly fused to generate the input vector, which is expressed as: Z = A + (d, f); Where Z represents the weighted input vector, and represents the weighting coefficients of the scene classification result and the user state evaluation result respectively; Input the weighted input vector into the generator of the AIGC smart model; First, the generator generates text content based on the input vector using the BERT model, expressed as: = BERT(Z;); Among them, represents the generated text content, and represents the parameters of the BERT model; Secondly, the generator uses the TTS model to convert text content into speech content. The expression is: =TTS(;); Among them, represents the generated speech content, represents the parameters of the TTS model Next, the generator uses the generative adversarial network GAN model to generate video content, expressed as: =GAN(Z; ); Among them, represents the generated video content, represents the parameters of the GAN model; Finally, the generated text, voice, and video are combined into a multimodal content, expressed as: Y=[,,]; Among them, Y represents the complete multimodal content that adapts to the current scenario and user needs.
6. The diversified interactive system based on the AIGC intelligent model as claimed in claim 5, characterized in that: When the user responds to the generated multimodal content by voice, the automatic speech recognition ASR technology is used to extract the voice content input by the user and convert it into text, and then the language identification LID and speech emotion recognition SER technology are used to extract the language information and emotion information of the user's voice content to generate a preliminary speech analysis result. The specific steps are as follows: When a user responds to the generated multimodal content, the user’s voice input is first converted into text using automatic speech recognition (ASR) technology; Identify the language category of the user's voice input through language identification LID technology; After speech content recognition, speech emotion recognition (SER) technology is used to analyze the emotional information in the user's speech; Introduce the speech analysis formula to generate preliminary speech analysis results. The expression is: R=σ(ε·O(t)+ζ·P(t)+η·Q(t)+δ·ψ(h,n,m)); Where R represents the preliminary speech analysis result, σ represents the activation function, ε, ζ and η represent the weighting coefficients of ASR, LID and SER respectively, h, n and m represent the text information output by ASR, the language information output by LID and the emotional information output by SER respectively, ψ represents the fusion function, O(t) is the text signal extracted by ASR, which represents the text information converted from the speech at time t, P(t) is the language signal extracted by LID, which represents the emotional state recognized from the speech at time t, and Q(t) is the emotional signal extracted by SER, which represents the emotional state recognized from the speech at time t.
7. The diversified interactive system based on the AIGC intelligent model as claimed in claim 6, characterized in that: The acoustic event classification AEC and acoustic event detection AED technologies are used to analyze the non-language information in the voice content input by the user, and the final voice analysis results are generated in combination with the preliminary voice analysis results. The specific steps are as follows: For the original audio signal, the time-frequency conversion is performed through the Mel-frequency cepstral coefficient MFCC to convert the audio signal into a spectrum graph; Use convolutional neural network (CNN) to extract features and classify spectrograms to identify different acoustic events in audio. Performing time sequence detection on the audio signal to identify the occurrence time, duration and intensity of each event; Use recurrent neural network (RNN) to further analyze time series information and extract non-language signal features; The preliminary speech analysis results are combined with the non-language signal features obtained through acoustic event classification and detection to generate the final speech analysis results, which are expressed as: r=σ(ε·O(t)+ζ·P(t)+η·Q(t)+δ·ψ(h,n,m)+λ·(N(t)+L(t)+ D(t)); Among them, r represents the final speech analysis result, N(t) is the background noise intensity, which means that at time t, the intensity of the background noise is calculated based on the input audio signal, L(t) is the laughter intensity, which means the obviousness of the laughter in the speech at time t, D(t) is the speech pause duration, which means the pause duration in the speech is calculated at time t, and λ represents the weighting coefficient of the non-language signal.
8. The diversified interactive system based on the AIGC intelligent model as claimed in claim 7, characterized in that: According to the final speech analysis result, combined with the scene classification result and the user status evaluation result, the large language model LLM is called to parse the semantics and generate the response text. The specific steps are as follows: Use the Qwen2 basic language model as the basis of the large language model LLM; The final speech analysis results, scene classification results and user status evaluation results are weighted and fused to form a comprehensive input vector; Pass the synthesized input vector as input to the large language model LLM, and use the large language model LLM internal encoder to convert the synthesized input vector into a contextual embedding; After obtaining the context embedding, the large language model LLM uses its internal neural network structure to deeply process the context embedding and obtain the hidden state; Based on the hidden state, the large language model LLM generates the response text through the internal decoder, which is expressed as: =Decoder LLM (;); Among them, represents the generated response text, Decoder LLM Represents the large language model LLM through the internal decoder and represents the parameter set of the decoder.
9. The diversified interactive system based on the AIGC intelligent model as claimed in claim 8, characterized in that: The adjusted response text is converted into speech output through speech synthesis technology, and the specific steps are as follows: Select TTS, an end-to-end speech synthesis method based on enhanced variational inference and normalized flow; The generated response text is converted into input text features through word segmentation and word embedding processes; For the input text features, the potential speech features are generated through the enhanced variational inference network; Use normalization flow to transform the latent features to adjust the pitch, rhythm and intonation of the speech; Based on the latent features adjusted by the normalized flow, speech waveforms with diverse rhythms and pitches are generated through a random duration predictor; The generated speech waveform is output through an audio playback device.
10. The diversified interactive system based on the AIGC intelligent model as claimed in claim 9, characterized in that: When the user receives the voice response, the generated response text is adjusted based on the user's voice feedback. The specific steps are: After the user receives the voice response generated by the speech synthesis technology, the user can directly comment and ask questions to provide voice feedback information, and receive the user's voice feedback information; Based on the received voice feedback, the user's voice feedback information is converted into text information using automatic speech recognition ASR technology; Combine language identification LID and speech emotion recognition SER technology to extract language and emotion information from user feedback; Generate preliminary speech feedback analysis results through speech analysis formula; Perform acoustic event classification (AEC) and acoustic event detection (AED) on speech feedback to extract non-language information, and generate the final speech feedback analysis results in combination with the preliminary speech feedback analysis results; The final feedback speech analysis results, scene classification results, and user status assessment results are input into the large language model (LLM) for semantic analysis to adjust the generated response text.
Citation Information
Cited By
Adaptive scene intelligent interaction system based on AI
CN120704532A