Real-time voice interaction and adaptive content generation system based on RTC and AIGC
By integrating AIGC technology in the RTC system, capturing user voice input for natural language understanding and content generation, and optimizing content strategies in real-time, the stability and delay control problems of RTC systems in large-scale concurrent user processing in the existing technology, as well as the improvement space for the AIGC system in real-time and interactivity, real-time voice interaction and personalized content generation are achieved with efficient and low-latency real-time voice interaction and personalized content generation.
Patent Information
- Application Number
- CN202510366647.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-27
AI Technical Summary
When existing RTC systems deal with large-scale concurrent users, there is still room for improvement in stability and delay control, while AIGC systems have room for improvement in real-time and interactiveness, which is difficult to meet the needs of users for real-time interactive dissemination and automatic content generation.
The real-time voice interaction and adaptive content generation system based on RTC and AIGC are adopted to capture user voice input through the device's microphone, and use voice recognition technology to convert it into text information, perform natural language understanding and semantic analysis, dynamically generate personalized content, and optimize content generation strategies in real time based on user feedback.
It realizes an RTC system that efficiently handles large-scale concurrent users, provides a low-latency real-time voice interaction experience, and generates personalized and dynamic interactive content through AIGC technology to meet the needs of different types of users and improves the quality and efficiency of real-time communication.
Smart Images

Figure CN120220686A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and specifically to a real-time voice interaction and adaptive content generation system based on RTC and AIGC. Background Art
[0002] With the rapid development of technology, Real-Time Communication (RTC) and AI-Generated Content (AIGC) have become key technologies driving modern communication and creation. The real-time communication technology enables users to exchange information in real time by providing low-latency audio and video transmission functions. However, in dealing with a large number of concurrent users, there is still room for improvement in the stability and latency control of existing RTC systems. Against this background, we need an RTC system that can efficiently handle a large number of concurrent users while maintaining low latency.
[0003] On the other hand, AI-generated content technology can automatically generate high-quality text, images, videos, etc. by leveraging advanced technologies such as deep learning. This not only greatly improves the efficiency of content generation but also provides users with a more personalized and rich content experience. However, current AIGC systems are mainly used for the generation of static content, and there is still much room for improvement in terms of real-time performance and interactivity. For example, in social media or online live broadcasts, users not only passively receive content but also hope to generate or modify content through interaction, which poses challenges to traditional AIGC systems.
[0004] Combining RTC and AIGC technologies can realize a real-time voice interaction and adaptive content generation system based on RTC and AIGC. This system can not only provide a real-time voice interaction experience but also automatically generate adaptive content according to users' real-time feedback, thus greatly enriching the user experience. The technical integration and function expansion of such a system can not only improve the quality and efficiency of real-time communication but also achieve more personalized and dynamic interactive content, thereby meeting the needs of different types of users and promoting the development of real-time interactive communication and content automatic generation technologies.
[0005] In view of the above technical deficiencies, a solution for a real-time voice interaction and adaptive content generation system based on RTC and AIGC is proposed. Summary of the Invention
[0006] To solve the above problems, the present invention provides the following technical solutions:
[0007] A real-time voice interaction and adaptive content generation method based on RTC and AIGC, comprising:
[0008] Capture the user's voice input through the device's microphone and use speech recognition technology to convert it into text information;
[0009] Pass the recognized text input to the natural language understanding module for semantic parsing to identify the user's intent and information;
[0010] According to the user's intent and requirements, use AIGC technology to dynamically generate personalized content;
[0011] According to the feedback in the interaction, optimize the generated content in real time to adapt to the changes in the user's needs, convert the generated text content into voice output, and provide it to the user.
[0012] Furthermore, the speech acquisition uses a deep learning model for acoustic modeling, combines a language model for decoding and searching to generate a candidate text sequence, outputs the recognition result in chunks through streaming processing technology, performs a fast Fourier transform on each frame of the signal to obtain the spectrum, maps it to the Mel scale through a Mel filter bank to simulate the non-linear perception of the human ear. Take the logarithmic energy and perform a discrete cosine transform, extract the first 13-dimensional coefficients as features, omit the DCT step, retain the output of the Mel filter bank, and parameterize the speech signal based on the vocal tract model.
[0013] Furthermore, the speech acquisition includes removing the device IMEI and location-sensitive information in the recording, completing wake-word detection at the device end, not uploading non-wake-up speech, encrypting the speech stream using TLS1.3, and using ECDHE-ECDSA for key exchange to ensure end-to-end security. Identify microphone open circuit / short circuit through impedance detection or white noise injection, and automatically switch to the backup microphone when the main microphone fails.
[0014] Furthermore, the natural language processing includes extracting named entities in the user input based on a conditional random field or BiLSTM model, mapping the user's statement to a preset intent label through a pre-trained language model, storing the historical conversation state to support multi-turn interaction, calling a pre-trained large language model to generate a text response, linking a multi-modal generation model to generate image or video content, and adjusting the generation style according to the user profile, including language complexity, sentiment tendency, and domain term adaptation.
[0015] Furthermore, the feedback optimization includes directly adjusting the generated content through user ratings or correction instructions, analyzing the user's interaction behavior to optimize the generation strategy, and updating the AIGC model parameters with the user satisfaction as the reward function.
[0016] The calculation formula is as follows:
[0017]
[0018] Where, θ t+1To optimize the generated strategy, R(τ) is the cumulative reward based on the interaction trajectory τ, and α is the learning rate.
[0019] Furthermore, the generated content includes recording the user's negative / affirmative operations on the generated content, monitoring interactive behaviors, monitoring the interruption rate, frequency of repeated questions, and response waiting time, combining emotion recognition technology to judge user emotions, online fine-tuning of generation model parameters, updating the model based on user correction data, using lightweight technology to reduce computing overhead, ensuring response speed, caching the most recent 3-5 rounds of conversation history, building a dynamic context vector, retrieving similar historical cases to guide current generation, balancing accuracy, diversity and security, shielding sensitive words through constraint decoding, adjusting the randomness of answers through temperature parameters, using reinforcement learning to optimize long-term user satisfaction indicators, selecting timbre according to user portraits, adjusting speech speed according to emotions, using VITS and Tacotron models to generate waveforms, supporting mixed synthesis of Chinese and English, streaming processing to achieve sentence-by-sentence playback, using the WebRTC protocol to transmit audio streams, adaptively adjusting the bit rate to combat network fluctuations, deploying edge nodes to process TTS nearby, and reducing cross-regional transmission delays.
[0020] According to one aspect of the present invention, there is provided a real-time voice interaction and adaptive content generation system based on RTC and AIGC, comprising: a real-time voice interaction module for collecting user voice input, transmitting it through the RTC protocol and converting it into text data;
[0021] The AIGC content generation module dynamically generates multimodal response content based on the text data and contextual information input by the user;
[0022] Adaptive optimization module, which adjusts content generation strategies in real time based on user behavior data and feedback;
[0023] The multimodal output module returns the generated text, voice or image content to the user terminal through the RTC protocol.
[0024] Furthermore, the generation system includes an integrated noise reduction algorithm and echo cancellation function, supports concurrent input from multiple devices; uses WebRTC or a custom UDP protocol to achieve end-to-end low-latency transmission, converts voice streams into text in real time based on a deep learning model, converts the generated text content into natural speech output, stores and analyzes the semantic relevance of historical conversations, calls a pre-trained large language model to generate text, and links the image / video generation model to output composite content, double-checks the legitimacy of the generated content through a rule engine and an AI model, builds dynamic user tags based on operation history, device usage patterns, and sensor data, optimizes the weight parameters of the generation model through explicit scoring and implicit behavior data, and switches the interaction strategy according to environmental data;
[0025] The adaptive echo cancellation formula is as follows:
[0026]
[0027] Among them, W(n + 1) is the output after elimination, W(n) is the current output, ||X(n)|| is the filter coefficient vector at the nth moment, and ε is the regularization constant to prevent the denominator from being zero;
[0028] By building a connectionist temporal classification loss function, speech recognition learning is performed on the data, and the calculation is as follows:
[0029]
[0030] Among them, π is all possible phoneme alignment paths, B is the path compression function, P(π|x) is the probability of the current path, and y is the true label sequence.
[0031] According to one aspect of the present invention, there is provided a computer device including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above-mentioned real-time voice interaction and adaptive content generation method based on RTC and AIGC are implemented.
[0032] According to one aspect of the present invention, there is provided a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the above-mentioned real-time voice interaction and adaptive content generation method based on RTC and AIGC are implemented.
[0033] Compared with the prior art, the beneficial effects of the present invention are:
[0034] 1. In the real-time voice interaction and adaptive content generation method based on RTC and AIGC of the present invention, the voice input of the user is captured by the microphone of the device and converted into text information using speech recognition technology; the recognized text input is passed to the natural language understanding module for semantic parsing to identify the user's intention and information; according to the user's intention and requirements, AIGC technology is used to dynamically generate personalized content; the generated content is optimized in real time according to the feedback in the interaction to adapt to the changes in the user's needs, and the generated text content is converted into voice output and provided to the user, which has the effect of meeting the needs of different types of users and promoting real-time interactive dissemination and automatic content generation.
[0035] 2. In a real-time voice interaction and adaptive content generation system based on RTC and AIGC of the present invention, a real-time voice interaction module is used to collect user voice input, transmit it through the RTC protocol and convert it into text data; an AIGC content generation module dynamically generates multimodal response content based on the text data and context information input by the user; an adaptive optimization module adjusts the content generation strategy in real time according to user behavior data and feedback; a multimodal output module returns the generated text, voice or image content to the user terminal through the RTC protocol, which improves the quality and efficiency of real-time communication. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] For the convenience of those skilled in the art to understand, the present invention will be further described below with reference to the accompanying drawings;
[0037] Figure 1 It is a schematic diagram of the overall framework of a real-time voice interaction and adaptive content generation method based on RTC and AIGC of the present invention;
[0038] Figure 2 It is a schematic diagram of the process of a real-time voice interaction and adaptive content generation system based on RTC and AIGC of the present invention;
[0039] Figure 3 It is a schematic diagram of the structure of a computer device in a real-time voice interaction and adaptive content generation system based on RTC and AIGC of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0041] As Figures 1-3 shown, the present application provides a real-time voice interaction and adaptive content generation method based on RTC and AIGC, including:
[0042] S1: Capture the user's voice input through the microphone of the device and convert it into text information using speech recognition technology;
[0043] S2: Transmit the recognized text input to the natural language understanding module for semantic parsing to identify the user's intention and information;
[0044] S3: Dynamically generate personalized content using AIGC technology according to the user's intention and requirements;
[0045] S4: Optimize the generated content in real time according to the feedback in the interaction, adapt to the changing needs of users, convert the generated text content into voice output, and provide it to users.
[0046] In one embodiment, the front-end device, multi-microphone array: supports beamforming and sound source localization, and captures user speech directionally. Low-power processor: runs local VAD, noise reduction, and wake word detection.
[0047] RTC communication layer, WebRTC protocol: realizes low-latency (<200ms) audio stream transmission. Edge node: deploys TURN / STUN servers close to users to optimize the network transmission path.
[0048] AIGC generation layer, intent recognition and entity extraction: parses user semantics based on BERT or GPT models. Content generation: uses GPT-4 or similar models to generate dynamic responses. Real-time optimization: online fine-tunes the generation strategy based on user feedback.
[0049] TTS voice synthesis layer, neural network TTS: such as VITS or Tacotron2, supports multi-language and emotional voice synthesis. Streaming playback: generates and plays sentence by sentence, with the first sentence delay <100ms.
[0050] Step 1: Speech acquisition and preprocessing
[0051] Microphone array capture:
[0052] User says: "Book me a flight to Beijing tomorrow."
[0053] The 6-microphone circular array focuses on the user's direction (azimuth 45°) through beamforming, with a sampling rate of 16kHz.
[0054] Signal enhancement;
[0055] Noise reduction: uses a deep noise reduction model (such as RNNoise) to increase the SNR from -5dB to 15dB.
[0056] Echo cancellation: eliminates the background music echo played by the device through the NLMS algorithm.
[0057] Voice activity detection (VAD);
[0058] Detects a valid speech segment (lasting 2.5 seconds), and retains 2.0 seconds of valid audio after silence compression.
[0059] Step 2: Real-time transmission and intent parsing
[0060] Low-latency encoding and transmission;
[0061] The audio stream is transmitted to the cloud ASR module via WebRTC using Opus encoding (16 kbps), with an end-to-end latency of 120 ms.
[0062] Automatic Speech Recognition (ASR):
[0063] The recognition result is: "Book me a flight to Beijing tomorrow."
[0064] Recognition confidence: 0.92.
[0065] Natural Language Understanding (NLU);
[0066] Intent classification: Recognized as "flight_booking" (confidence 0.89).
[0067] Entity extraction:
[0068] Departure time: Tomorrow (parsed as 2024-03-20).
[0069] Destination: Beijing.
[0070] Context association: Query the user's historical orders and find a preference for "economy class".
[0071] Step 3: Adaptive content generation:
[0072] Initial response generation;
[0073] Generated by the AIGC model: "We have found an economy class flight to Beijing tomorrow for 800 yuan. Do you want to book it?"
[0074] Generation time: 200 ms.
[0075] User feedback collection:
[0076] User interruption: "Too expensive. Do you have anything cheaper?" (Explicit feedback + implicit interruption signal).
[0077] Real-time optimization:
[0078] Online fine-tuning: Fine-tune the generation template on the data corrected by the user to reduce the generation weight of "800 yuan".
[0079] Retrieval enhancement: Retrieve "low-price flight" related cases from the knowledge base and generate a new response: "We have found a special offer flight to Beijing tomorrow for 600 yuan. Do you want to book it?"
[0080] Optimization time: 150 ms.
[0081] Step 4: Speech synthesis and playback
[0082] TTS parameter control, adjust the voice style according to the user's mood (slightly dissatisfied): increase the speaking speed by 10% and reduce the pitch by 5Hz.
[0083] Streaming speech generation, using the VITS model to generate speech sentence by sentence, with a delay of 80ms for the first sentence. Synthesized content: "I found a special air ticket to Beijing tomorrow, the price is 600 yuan, do you want to book it?"
[0084] Real-time playback, streaming audio to user devices via WebRTC, with a playback delay of <200ms.
[0085] Step 5: Continuous Optimization and Learning:
[0086] Feedback data records, the user's final choice: "Okay, let's book it." Record the complete process of this interaction: user input, system response, user feedback, and final decision.
[0087] Update the offline model, add the interaction data to the training set, fine-tune the generation strategy of the AIGC model, and optimize the goal: to increase the recommendation priority of low-cost air tickets.
[0088] Specifically, the speech acquisition adopts a deep learning model for acoustic modeling, combines a language model for decoding search, generates a candidate text sequence, outputs the recognition result in blocks through streaming processing technology, performs fast Fourier transform on each frame signal to obtain a spectrum, maps it to the Mel scale through a Mel filter group, and simulates the nonlinear perception of the human ear. Take the logarithmic energy and perform discrete cosine transform, extract the first 13 dimensional coefficients as features, omit the DCT step, retain the Mel filter group output, and parameterize the speech signal based on the vocal tract model.
[0089] Specifically, the voice collection includes removing the device IMEI and geographical location sensitive information in the recording, completing wake-up word detection on the device side, not uploading non-wake-up voice, using TLS1.3 to encrypt the voice stream, and using ECDHE-ECDSA for key exchange to ensure end-to-end security. The microphone open circuit / short circuit is identified through impedance detection or white noise injection, and the main microphone automatically switches to the backup microphone when it fails.
[0090] Specifically, the natural language processing includes extracting named entities in user input based on conditional random fields or BiLSTM models, mapping user sentences to preset intent labels through pre-trained language models, storing historical dialogue states to support multiple rounds of interaction, calling pre-trained large language models to generate text responses, linking multimodal generation models to generate images or video content, and adjusting the generation style according to user portraits, including language complexity, emotional tendency and domain terminology adaptation.
[0091] In one embodiment, real-time feedback data collection:
[0092] Explicit feedback collection:
[0093] Users directly rate (e.g., 1 - 5 stars) or make text corrections (e.g., "Change to 23°C").
[0094] Record users' negative / positive operations on the generated content (e.g., clicking the "Disagree" button).
[0095] Implicit feedback analysis:
[0096] Monitor interaction behaviors: interruption rate (number of times users interrupt the speech), frequency of repeated questions, response waiting duration. Combine with emotion recognition technology to judge users' emotions (e.g., analyze the degree of anger / satisfaction through speech intonation or text keywords).
[0097] Real - time model adjustment:
[0098] Online fine - tune the parameters of the generation model: update the model based on users' correction data (e.g., perform small - scale gradient descent on the last layer of the GPT model).
[0099] Use lightweight technologies (such as LoRA adapters) to reduce computational overhead and ensure response speed.
[0100] Context memory enhancement:
[0101] Cache the last 3 - 5 rounds of conversation history and construct a dynamic context vector.
[0102] Retrieve similar historical cases to guide the current generation (e.g., if the user emphasizes "concise answer" multiple times, suppress redundant content).
[0103] Multi - objective optimization strategy:
[0104] Balance accuracy, diversity, and security: mask sensitive words through constrained decoding, and adjust the randomness of answers through the temperature parameter.
[0105] Adopt reinforcement learning (such as the PPO algorithm) to optimize the long - term user satisfaction index.
[0106] Voice synthesis parameter control:
[0107] Select the voice color according to the user profile (e.g., use a high pitch in children's mode), and adjust the speech rate according to the emotion (speed up by 20% when angry).
[0108] Predict text prosody: automatically add pauses and stresses at "important notifications".
[0109] High - quality voice generation:
[0110] Use models such as VITS and Tacotron to generate waveforms, supporting mixed Chinese - English synthesis.
[0111] Streaming processing enables sentence-by-sentence playback with the delay of the first sentence controlled within 200 ms.
[0112] Network transmission optimization:
[0113] The WebRTC protocol is used to transmit the audio stream, and the bitrate (6 - 64 kbps) is adaptively adjusted to combat network fluctuations.
[0114] Edge nodes are deployed to process TTS locally, reducing cross-region transmission latency.
[0115] Real-time playback synchronization;
[0116] Dynamic mixing processing: Voice playback is started within 50 ms after the user stops speaking.
[0117] Common response audio (such as "Okay") is pre-loaded to achieve instant feedback within 100 ms.
[0118] AB testing verification;
[0119] The new and old generation strategies are run in parallel, and the differences in metrics such as click-through rate and task completion rate are statistically analyzed.
[0120] The improvement in experience is quantified through user surveys (such as NPS scores).
[0121] Offline model re-training;
[0122] The accumulated feedback data is added to the training set daily to fully update the generation model.
[0123] The capabilities of the large model are periodically distilled into a lightweight version, taking into account both effectiveness and inference speed.
[0124] Fault tolerance mechanism:
[0125] When the model times out (>1 second), switch to a fast template engine (such as generating fixed answers through rule matching).
[0126] When a speech synthesis failure is detected, it automatically switches to text display.
[0127] When two consecutive recognition failures occur, it actively asks: "Do you mean...?" and provides option buttons.
[0128] During the optimization process, a progress prompt is given: "Improving according to your feedback, expected to take effect in 10 seconds".
[0129] Example process;
[0130] The user says, "Play Jay Chou's 'Qi Li Xiang'." → The system plays the song. → The user interrupts: "The volume is too low." → The system immediately increases the volume and responds: "The volume has been increased to 80% for you." → Subsequent similar requests default to increasing the volume. The entire process realizes a closed-loop of "feedback collection (insufficient volume) → real-time optimization (adjusting parameters) → voice confirmation (announcing the operation result)", and the end-to-end delay is less than 500 ms.
[0131] Specifically, the feedback optimization includes directly adjusting the generated content through user ratings or correction instructions, analyzing user interaction behaviors to optimize the generation strategy, and updating the AIGC model parameters with user satisfaction as the reward function.
[0132] The calculation formula is as follows:
[0133]
[0134] where θ t+1 is the optimized generation strategy, R(τ) is the cumulative reward based on the interaction trajectory τ, and α is the learning rate.
[0135] Specifically, the generated content includes recording the user's negative / positive operations on the generated content, monitoring interaction behaviors, monitoring the interruption rate, repeated question frequency, and response waiting duration, combining emotion recognition technology to judge the user's emotion, online fine-tuning the generated model parameters, updating the model based on user correction data, using lightweight technology to reduce computational overhead to ensure response speed, caching the last 3-5 rounds of conversation history, constructing a dynamic context vector, retrieving similar historical cases to guide the current generation, balancing accuracy, diversity, and security, shielding sensitive words through constrained decoding, adjusting the randomness of answers through the temperature parameter, optimizing the long-term user satisfaction index using reinforcement learning, selecting the voice color according to the user profile, adjusting the speech rate according to the emotion, using the VITS and Tacotron models to generate waveforms, supporting mixed Chinese and English synthesis, realizing sentence-by-sentence playback through streaming processing, transmitting the audio stream using the WebRTC protocol, adaptively adjusting the bit rate to counter network fluctuations, and deploying edge nodes to process TTS nearby to reduce cross-regional transmission delay.
[0136] According to one aspect of the present invention, there is provided a real-time voice interaction and adaptive content generation system based on RTC and AIGC, including: a real-time voice interaction module for collecting user voice input, transmitting it through the RTC protocol and converting it into text data;
[0137] an AIGC content generation module for dynamically generating multi-modal response content based on the text data and context information input by the user;
[0138] an adaptive optimization module for adjusting the content generation strategy in real time according to user behavior data and feedback;
[0139] The multimodal output module returns the generated text, voice or image content to the user terminal through the RTC protocol.
[0140] Specifically, the generation system includes an integrated noise reduction algorithm and echo cancellation function, supports concurrent input from multiple devices; uses WebRTC or a custom UDP protocol to achieve end-to-end low-latency transmission, converts voice streams into text in real time based on a deep learning model, converts the generated text content into natural voice output, stores and analyzes the semantic relevance of historical conversations, calls a pre-trained large language model to generate text, and links the image / video generation model to output composite content, double-checks the legitimacy of the generated content through a rule engine and an AI model, builds dynamic user tags based on operation history, device usage patterns, and sensor data, optimizes the weight parameters of the generation model through explicit scoring and implicit behavior data, and switches the interaction strategy according to environmental data;
[0141] The adaptive echo cancellation formula is as follows:
[0142]
[0143] Among them, W(n+1) is the output after elimination, W(n) is the current output, ||X(n)|| is the filter coefficient vector at the nth moment, and ε is the regularization constant to prevent the denominator from being zero;
[0144] By building a connection time series classification loss function, the data is used for speech recognition learning, and the calculation is as follows:
[0145]
[0146] Among them, π is all possible phoneme alignment paths, B is the path compression function, P(π|x) is the probability of the current path, and y is the true label sequence.
[0147] In one embodiment, the microphone array: a 6-8 microphone circular array supports beamforming and sound source localization.
[0148] Local processor: runs pre-processing algorithms such as VAD, noise reduction, and wake-up word detection.
[0149] Speaker: Plays system-generated voice responses.
[0150] WebRTC protocol: enables low-latency (<200ms) audio streaming.
[0151] Edge nodes: deploy TURN / STUN servers to optimize network transmission paths.
[0152] Streaming media server: manages the encoding, decoding and forwarding of audio streams.
[0153] Automatic Speech Recognition (ASR): Converts user speech into text.
[0154] Natural Language Understanding (NLU): Parses user intents and entities.
[0155] Content Generation: Generates dynamic responses based on GPT-4 or similar models.
[0156] Feedback Optimization: Adjusts the generation strategy in real-time according to user feedback.
[0157] Neural Network Text-to-Speech (TTS): Such as VITS or Tacotron2, supports multi-language and emotional speech synthesis.
[0158] Streaming Playback: Generates and plays sentences one by one, with the first sentence delay < 100ms.
[0159] Interaction Log: Records data such as user input, system responses, and user feedback.
[0160] Model Training: Offline updates the AIGC and TTS models based on cumulative data.
[0161] The present invention also provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above real-time voice interaction and adaptive content generation system based on RTC and AIGC are implemented.
[0162] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above real-time voice interaction and adaptive content generation system based on RTC and AIGC are implemented.
[0163] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium provided in this application and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0164] It should be noted that in this article, the terms "include", "comprise", or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, device, article, or method including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, device, article, or method. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, device, article, or method including that element.
[0165] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the present invention to only the specific embodiments. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the present invention, so that those skilled in the art can understand and utilize the present invention well. The present invention is only limited by the claims and their full scope and equivalents.
Claims
1. A real-time voice interaction and adaptive content generation method based on RTC and AIGC, characterized in that: include: Capture user voice input through the device's microphone and convert it into text information using speech recognition technology; Pass the recognized text input to the natural language understanding module for semantic analysis and identification of user intent and information; Dynamically generate personalized content using AIGC technology based on user intent and needs; The generated content is optimized in real time based on the feedback from the interaction, adapting to changes in user needs, and the generated text content is converted into voice output and provided to users.
2. According to claim 1, a real-time voice interaction and adaptive content generation method based on RTC and AIGC is characterized in that: The speech acquisition adopts a deep learning model for acoustic modeling, combines a language model for decoding search, generates a candidate text sequence, outputs the recognition result in blocks through streaming processing technology, performs fast Fourier transform on each frame signal to obtain a spectrum, maps it to the Mel scale through a Mel filter group, and simulates the nonlinear perception of the human ear. Take the logarithmic energy and perform discrete cosine transform, extract the first 13 dimensional coefficients as features, omit the DCT step, retain the Mel filter group output, and parameterize the speech signal based on the vocal tract model.
3. According to claim 2, a method for real-time voice interaction and adaptive content generation based on RTC and AIGC is characterized in that: The voice collection includes removing the device IMEI and geographical location sensitive information in the recording, completing wake-up word detection on the device side, not uploading non-wake-up voice, using TLS1.3 to encrypt the voice stream, and using ECDHE-ECDSA for key exchange to ensure end-to-end security. The microphone is identified as open / short-circuited through impedance detection or white noise injection, and automatically switches to the backup microphone when the main microphone fails.
4. The method for real-time voice interaction and adaptive content generation based on RTC and AIGC according to claim 1, characterized in that: The natural language processing includes extracting named entities in user input based on conditional random fields or BiLSTM models, mapping user sentences to preset intent labels through pre-trained language models, storing historical dialogue states to support multiple rounds of interaction, calling pre-trained large language models to generate text responses, linking multi-modal generation models to generate images or video content, and adjusting the generation style according to user portraits, including language complexity, emotional tendency and domain terminology adaptation.
5. The method for real-time voice interaction and adaptive content generation based on RTC and AIGC according to claim 4, characterized in that: The feedback optimization includes directly adjusting the generated content through user ratings or correction instructions, analyzing user interaction behaviors to optimize the generation strategy, and updating the AIGC model parameters with user satisfaction as the reward function. The calculation formula is as follows: Among them, θ t+1 To optimize the generated strategy, R(τ) is the cumulative reward based on the interaction trajectory τ, and α is the learning rate.
6. The method for real-time voice interaction and adaptive content generation based on RTC and AIGC according to claim 5, characterized in that: The generated content includes recording the user's negative / affirmative operations on the generated content, monitoring interactive behaviors, monitoring the interruption rate, frequency of repeated questions, and response waiting time, combining emotion recognition technology to judge user emotions, online fine-tuning of generation model parameters, updating the model based on user correction data, using lightweight technology to reduce computing overhead, ensuring response speed, caching the most recent 3-5 rounds of conversation history, building a dynamic context vector, retrieving similar historical cases to guide current generation, balancing accuracy, diversity and security, shielding sensitive words through constraint decoding, adjusting the randomness of answers through temperature parameters, using reinforcement learning to optimize long-term user satisfaction indicators, selecting timbre according to user portraits, adjusting speech speed according to emotions, using VITS and Tacotron models to generate waveforms, supporting mixed synthesis of Chinese and English, streaming processing to achieve sentence-by-sentence playback, using the WebRTC protocol to transmit audio streams, adaptively adjusting the bit rate to combat network fluctuations, deploying edge nodes to process TTS nearby, and reducing cross-regional transmission delays.
7. A real-time voice interaction and adaptive content generation system based on RTC and AIGC, characterized in that: include: Real-time voice interaction module, used to collect user voice input, transmit it through RTC protocol and convert it into text data; The AIGC content generation module dynamically generates multimodal response content based on the text data and contextual information input by the user; Adaptive optimization module, which adjusts content generation strategies in real time based on user behavior data and feedback; The multimodal output module returns the generated text, voice or image content to the user terminal through the RTC protocol.
8. The real-time voice interaction and adaptive content generation system based on RTC and AIGC according to claim 7, characterized in that: The generation system includes an integrated noise reduction algorithm and echo cancellation function, supports concurrent input from multiple devices; uses WebRTC or a custom UDP protocol to achieve end-to-end low-latency transmission, converts voice streams into text in real time based on a deep learning model, converts the generated text content into natural voice output, stores and analyzes the semantic relevance of historical conversations, calls a pre-trained large language model to generate text, and links the image / video generation model to output composite content, double-checks the legitimacy of the generated content through a rule engine and an AI model, builds dynamic user tags based on operation history, device usage patterns, and sensor data, optimizes the weight parameters of the generation model through explicit scoring and implicit behavior data, and switches the interaction strategy according to environmental data; The adaptive echo cancellation formula is as follows: Among them, W(n+1) is the output after elimination, W(n) is the current output, ||X(n)|| is the filter coefficient vector at the nth moment, and ε is the regularization constant to prevent the denominator from being zero; By building a connection time series classification loss function, the data is used for speech recognition learning, and the calculation is as follows: Among them, π is all possible phoneme alignment paths, B is the path compression function, P(π|x) is the probability of the current path, and y is the true label sequence.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the real-time voice interaction and adaptive content generation method based on RTC and AIGC described in any one of claims 1 to 6 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the real-time voice interaction and adaptive content generation method based on RTC and AIGC described in any one of claims 1 to 6 are implemented.
Citation Information
Cited By
Real-time voice interaction method and system based on large model
CN120853551A
Experimental commentary real-time generation method
CN121260160A
AI-Based Real-Time Multilingual Translation Performance Support System
KR103004571B1