Voice interaction intelligent processing method and real-time communication equipment
By adopting a smart voice interaction processing method in real-time communication devices, combining VAD and AEC algorithms, and using WebRTC and UDP protocols for low-latency transmission, the problem of difficult separation of user voice and device sound interference in the voice interaction system is solved, and an efficient and low-latency voice interaction experience is achieved.
Patent Information
- Application Number
- CN202510391797.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-06-13
AI Technical Summary
When existing voice interaction systems are deployed in real-time communication devices, it is difficult to effectively separate the interference components of user voice and device sound, resulting in problems such as difficult to eliminate echo interference, low recognition rate, and lag in interaction.
The voice interaction intelligent processing method is adopted to accurately identify effective voice signals through the voice activity detection (VAD) module of the real-time communication device, and the "double-speaking" scenario is processed in combination with the optimized echo cancellation (AEC) algorithm to ensure that the voice signal is pure. The WebRTC transmission network and UDP protocol are used to realize low-latency voice data streaming, and signaling interaction and link optimization are carried out in conjunction with the WebSocket protocol. The cloud-based big model adopts streaming processing technology to realize real-time speech recognition, semantic understanding and speech synthesis.
It significantly reduces the delay of voice interaction, improves interaction fluency and accuracy, provides a natural and efficient voice interaction experience, and is suitable for terminal fields such as AI social companionship and oral learning.
Smart Images

Figure CN120148508A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of real-time communication, and in particular to a voice interaction intelligent processing method and a real-time communication device. Background Art
[0002] In recent years, the rapid development of artificial intelligence technology has promoted the wide application of voice interaction in fields such as smart homes and social companion robots, and it has gradually become the core mode of human-computer interaction. However, there are still technical problems in the actual deployment of existing voice interaction systems in real-time communication devices.
[0003] For example, the audio signals collected by the device microphone in the traditional solution usually contain the mixed superposition of user voices and the device's own playback sounds (such as music and reminder sounds). When such mixed superposition signals are transmitted to the cloud for processing, due to the insufficient preprocessing ability of the device side, it is difficult to effectively separate the interference components of the user voice and the device sound. Problems that occur include difficult elimination of echo interference, low recognition rate, and interaction lag. Therefore, there is a need in the industry to design an intelligent solution for voice interaction processing to solve the above technical problems. Summary of the Invention
[0004] The technical problem solved by the present invention is how to design a voice interaction intelligent processing method and a real-time communication device that can significantly reduce the latency of voice interaction.
[0005] In a first aspect, the present application proposes a voice interaction intelligent processing method, which realizes low latency of large model voice interaction; the method is used for communication between a cloud large model, a real-time communication device, and a user, the real-time communication device is used to receive a user instruction voice from the user, the method is used for a double-talk scenario where the user instruction voice and the device playback sound exist simultaneously, and the method includes: S1, collecting a mixed audio signal and extracting an effective voice signal from the mixed audio signal; S2, converting the effective voice signal into an echo-cancelled voice signal, and then uploading it to the cloud large model, and generating a feedback voice stream after processing; S3, receiving the feedback voice stream transmitted from the cloud large model and playing it by a local rendering module.
[0006] The inventor has found through research that an important bottleneck in the increasing popularity of human-computer interaction is that in the "double-talk" scenario where the user and the device speak simultaneously, the accuracy of speech recognition drops significantly, restricting the interactive experience. This bottleneck leads to an unsatisfactory market promotion speed for real-time communication devices with "voice interaction" embedded, and even a relatively high device return rate. The solution of the present invention accurately identifies effective voice signals through the voice activity detection (VAD) module of the real-time communication device, and processes the "double-talk" scenario by combining an optimized echo cancellation (AEC) algorithm to ensure pure voice signals. Using the WebRTC transmission network, low-latency voice data stream transmission is achieved through the UDP protocol, and signaling interaction and link optimization are carried out in cooperation with the WebSocket protocol. The cloud large model adopts streaming processing technology to achieve real-time speech recognition, semantic understanding, and speech synthesis. The present invention significantly reduces the end-to-end latency, improves the interactive fluency and accuracy, and is applicable to terminal fields such as AI social companionship and oral English learning, providing users with a natural and efficient voice interaction experience.
[0007] A further technical solution thereof is that the step S1 includes: S11, collecting a mixed audio signal through a microphone array in the real-time communication device, where the mixed audio signal includes a user instruction voice signal and also includes a device playback sound superposition signal corresponding to the device playback sound; S12, separating the device playback sound component from the mixed audio signal and calculating its energy mean as a reference energy baseline; S13, extracting the spectral features of the user instruction voice signal, and when the energy peak of the user instruction voice signal exceeds a preset multiple of the reference energy baseline and the spectral features match, extracting the effective voice signal therein.
[0008] A further technical solution thereof is that the step S2 includes: S21, preprocessing the effective voice signal using an adaptive echo cancellation algorithm to obtain an echo-cancelled voice signal; S22, in the WebRTC transmission environment, encapsulating the echo-cancelled voice signal into an RTP data packet, and uploading it to the cloud large model based on the UDP protocol after encapsulation; S23, the cloud large model performs real-time parsing on the streaming input RTP data packets to generate a feedback voice stream.
[0009] A further technical solution thereof is that the step S21 includes: S31, collecting the device playback sound superposition signal in real time; S32, using an adaptive echo cancellation algorithm, taking the device playback sound superposition signal as a reference signal to generate a reverse compensation signal matching the propagation path of the device playback sound; S33, performing phase-aligned digital domain superposition processing on the reverse compensation signal and the mixed audio signal to cancel the echo interference in the device playback sound superposition signal and obtain an echo-cancelled voice signal.
[0010] A further technical solution thereof is that the step S22 includes: S41, in a WebRTC transmission environment, encoding the echo-cancelled voice signal into a sequence of RTP data packets according to dynamic bitrate parameters, where the dynamic bitrate parameters are generated based on the average signal energy within a preset time window; S42, detecting the energy peak of the user voice in the RTP data packet sequence, and when the energy peak exceeds the reference energy baseline and the spectral characteristics match a preset template, generating a VAD start instruction carrying a timestamp and an encoding frame rate; S43, transmitting the VAD start instruction and the dynamic bitrate parameters to the cloud through the WebSocket protocol, and synchronously associating and storing the real-time collected bandwidth, latency, and packet loss rate parameters; S44, calculating the encapsulation interval of the RTP data packet sequence according to the ratio of the target bitrate of the dynamic bitrate parameters to the bandwidth, and adjusting the redundancy encoding ratio based on the packet loss rate difference; S45, uploading the RTP data packet sequence to the cloud large model according to the encapsulation interval and the redundancy ratio based on the UDP protocol, where when the average voice energy is lower than a preset reference threshold for N consecutive cycles, a VAD stop instruction is generated and the upload is aborted until a new instruction is triggered.
[0011] A further technical solution thereof is that the step S23 includes: S51, performing real-time parsing on the streaming input RTP data packets at a preset gap time window, and outputting a sequence of text fragments; S52, generating a final intent instruction based on the context relevance of the text fragment sequence; S53, generating a feedback voice stream according to the intent instruction, where the parameters of the feedback voice stream match the audio sampling rate of the real-time communication device.
[0012] A further technical solution thereof is that the step S3 includes: S71, receiving the feedback voice stream transmitted from the cloud large model through the UDP protocol, and playing it for the user to listen to by the local rendering module, so as to achieve low-latency interaction.
[0013] In a second aspect, the present application proposes a real-time communication device, which is communicatively connected to the cloud large model, and the real-time communication device is used to receive an instruction voice from a user, and the real-time communication device is used to implement the voice interaction intelligent processing method as described in the first aspect.
[0014] In summary, for the voice interaction intelligent processing method and the real-time communication device of the present invention, by optimizing the audio signal processing, data transmission, and cloud computing processes, low latency and high accuracy are achieved, and a resource-efficient voice interaction experience is also realized. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments that conform to the present invention, and are used together with the specification to explain the principles of the present invention.
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0017] Figure 1 It is a flowchart of the voice interaction intelligent processing method provided by the embodiment of the present invention.
[0018] Figure 2 It is another flowchart of the voice interaction intelligent processing method provided by the embodiment of the present invention.
[0019] Figure 3 It is yet another flowchart of the voice interaction intelligent processing method provided by the embodiment of the present invention.
[0020] Figure 4 It is a schematic diagram of the framework of the electronic device provided by the embodiment of the present invention. Detailed implementation manners
[0021] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0022] It should be understood that when used in this specification and the appended claims, the terms "comprises" and "comprising" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or other features, wholes, steps, operations, elements, components, and / or their combinations.
[0023] It should also be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0024] It should be further understood that the term "and / or" used in the specification of the present invention and the appended claims refers to any one or any combination of the related listed items and all possible combinations, and includes these combinations.
[0025] As used in this specification and the appended claims, the term "if" may be construed, depending on the context, as "when" or "once" or "in response to determining" or "in response to detecting". Similarly, the phrase "if determined" or "if [the described condition or event] is detected" may be construed, depending on the context, to mean "once determined" or "in response to determining" or "once [the described condition or event] is detected" or "in response to detecting [the described condition or event]".
[0026] In this specification and the appended claims, there may be multiple ways of expressing the same technical feature or technical term. For example, different forms of expression such as upper-level generalization, lower-level limitation, or synonymous substitution may be adopted. Those skilled in the art can clearly understand the substantially same technical meaning pointed to by different expressions based on their professional knowledge and in combination with the overall content of the specification and the drawings. The differences in different expressions only lie in the diversity at the literal level, do not constitute a substantial modification or limitation of the technical solution, and do not affect the certainty of the protection scope of the patent claims and the full disclosure of the technical content of the specification.
[0027] Embodiment 1
[0028] Please refer to Figures 1 to 3 As shown, it is a voice interaction intelligent processing method proposed in an embodiment of the present invention. Specifically, refer to Figure 1 As shown, the voice interaction intelligent processing method described in this application is used for communication between a cloud large model, a real-time communication device, and a user. The real-time communication device is used to receive a user instruction voice from the user. The method is used for a dual-talk scenario where the user instruction voice and the device-played sound coexist. The method includes: S1, collecting a mixed audio signal and extracting the effective voice signal from the mixed audio signal; S2, converting the effective voice signal into an echo-canceled voice signal and then uploading it to the cloud large model, and generating a feedback voice stream after processing; S3, receiving the feedback voice stream transmitted from the cloud large model and playing it by a local rendering module. Among them, in the dual-talk scenario, the user instruction voice and the device-played sound will be mixed with each other. If this mutual mixing is not processed, the user can distinguish it, but sometimes the real-time communication device cannot distinguish it.
[0029] A further technical solution thereof is that the real-time communication device includes a microphone array, and the step S1 includes: S11, collecting a mixed audio signal through the microphone array by the real-time communication device, where the mixed audio signal includes a user instruction voice signal and also includes a device playback sound superposition signal corresponding to the sound played by the device; S12, separating the component of the device playback sound superposition signal from the mixed audio signal and calculating its energy mean value as a reference energy baseline; S13, extracting the spectral features of the user instruction voice signal, and when the energy peak value of the user instruction voice signal exceeds a preset multiple of the reference energy baseline and the spectral features match, extracting the valid voice signal therein.
[0030] A further technical solution thereof is that the step S2 includes: S21, preprocessing the valid voice signal by using an adaptive echo cancellation algorithm to obtain an echo-cancelled voice signal; S22, in a WebRTC transmission environment, encapsulating the echo-cancelled voice signal into an RTP data packet, and uploading it to the cloud large model based on the UDP protocol after encapsulation; S23, the cloud large model performing real-time parsing on the streaming input RTP data packets to generate a feedback voice stream.
[0031] A further technical solution thereof is that the step S21 includes: S31, collecting the device playback sound superposition signal in real time; S32, using an adaptive echo cancellation algorithm, taking the device playback sound superposition signal as a reference signal to generate a reverse compensation signal matching the propagation path of the device playback sound; S33, performing a phase-aligned digital domain superposition process on the reverse compensation signal and the mixed audio signal to cancel the echo interference in the device playback sound superposition signal and obtain an echo-cancelled voice signal.
[0032] Its further technical solution is that the step S22 includes: S41, in the WebRTC transmission environment, encoding the echo-cancelled voice signal into a sequence of RTP data packets according to the dynamic bitrate parameter, where the dynamic bitrate parameter is generated based on the mean value of the signal energy within a preset time window; S42, detecting the energy peak of the user voice in the sequence of RTP data packets, and generating a VAD start instruction carrying a timestamp and an encoding frame rate when the energy peak exceeds the reference energy baseline and the spectral characteristics match a preset template; S43, transmitting the VAD start instruction and the dynamic bitrate parameter to the cloud through the WebSocket protocol, and synchronously associating and storing the parameters of the bandwidth, latency, and packet loss rate collected in real time; S44, calculating the encapsulation interval of the sequence of RTP data packets according to the ratio of the target bitrate of the dynamic bitrate parameter to the bandwidth, and adjusting the redundancy encoding ratio based on the packet loss rate difference; S45, uploading the sequence of RTP data packets to the cloud large model according to the encapsulation interval and the redundancy ratio based on the UDP protocol, where when the mean value of the voice energy is lower than the preset reference threshold for N consecutive periods, a VAD stop instruction is generated and the upload is aborted until a new instruction is triggered. Among them, when the mean value of the voice energy is lower than the preset reference threshold for N consecutive periods, the specific situation of the N periods can be defined by the engineer, such as two consecutive periods or three consecutive periods.
[0033] Its further technical solution is that the step S23 includes: S51, performing real-time parsing on the streaming input RTP data packets with a preset gap time window, and outputting a sequence of text fragments; S52, generating a final intent instruction based on the context relevance of the sequence of text fragments; S53, generating a feedback voice stream according to the intent instruction, where the parameters of the feedback voice stream match the audio sampling rate of the real-time communication device. Among them, the preset gap time window, that is, a small time interval, and the specific value can be defined by the engineer, such as any value between 100 milliseconds and 300 milliseconds.
[0034] Its further technical solution is that the step S3 includes: S71, receiving the feedback voice stream transmitted from the cloud large model through the UDP protocol, and playing it for the user to listen to by the local rendering module, so as to achieve low-latency interaction. Among them, the local rendering module is arranged in the real-time communication device, and the playback sound after playback can be heard by the user.
[0035] In a second aspect, the present application proposes a real-time communication device, which is communicatively connected to a cloud large model, and the real-time communication device is used to receive an instruction voice from a user, and the real-time communication device is used to implement the voice interaction intelligent processing method as described in any of the above embodiments.
[0036] In summary, in the field of real-time communication, due to the use of the TCP protocol retransmission mechanism and the non-streaming batch processing mode in traditional solutions, a perceptible dialogue gap occurs between user questions and device feedback, especially causing a split in the dialogue rhythm during "continuous multi-round interactions". In terms of environmental adaptability, real-time communication devices are often deployed in complex acoustic environments, and there is an acoustic coupling phenomenon caused by background noise and the device's own speaker. Traditional echo cancellation algorithms are prone to misjudging the effective sound source in the double-talk scenario of "user speaking overlapping with device broadcasting", resulting in the voice recognition engine receiving distorted signals, which in turn leads to semantic parsing deviations or even incorrect command executions. The issue of resource efficiency is reflected in the inefficient occupation of the limited computing power and network bandwidth of the platform. For example, the continuous transmission of invalid voice segments consumes the uplink bandwidth, redundant audio data increases the decoding burden on the cloud, and the high memory occupation of local preprocessing algorithms affects the multi-task concurrency performance. In addition, in cross-regional deployment or mobile network scenarios, fixed transmission strategies are difficult to adapt to the dynamically changing network quality, and traditional solutions lack the ability to perform real-time optimization of the transmission path and coding rate, seriously affecting the reliability of interactions. Based on this, the solution described in this application can significantly reduce the latency of voice interaction and improve the user experience during the process of chatting with the device by voice in a mixed audio signal scenario.
[0037] Embodiment 2
[0038] Please refer to Figure 4 , Figure 4 which is a block diagram of an electronic device provided by the present invention. The electronic device can be a terminal or a server. Among them, the terminal can be an electronic device with a communication function such as a smart phone, a tablet computer, a notebook computer, a desktop computer, a personal digital assistant, and a wearable device. It includes a processor 111, a communication interface 112, a memory 113, and a communication bus 114. Among them, the processor 111, the communication interface 112, and the memory 113 complete mutual communication through the communication bus 114.
[0039] The memory 113 is used to store computer programs.
[0040] In an embodiment of the present invention, when the processor 111 is used to execute the program stored on the memory 113, it implements the method provided by any one of the foregoing method embodiments.
[0041] It should be understood that in the embodiments of the present application, the processor 111 may be a central processing unit (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0042] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a storage medium, and the storage medium is a computer-readable storage medium. The computer program is executed by at least one processor in the computer system to implement the process steps of the above method embodiments.
[0043] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.
[0044] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of each unit is only a logical function division, and there may be other division methods in actual implementation. For example, a unit or component can be combined or integrated into another system, or some features can be ignored or not executed.
[0045] The steps in the method of the embodiments of the present invention can be adjusted, combined, and deleted according to actual needs. The units in the device of the embodiments of the present invention can be combined, divided, and deleted according to actual needs. In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0046] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that makes a contribution to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0047] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0048] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, provided that these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include these changes and modifications.
[0049] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A voice interaction intelligent processing method, characterized in that: The method is used for communication between a cloud-based large model, a real-time communication device and a user, the real-time communication device is used to receive a user instruction voice from the user, and the method is used for a dual-speaking scenario where the user instruction voice and the device playback sound exist simultaneously, and the method includes: S1, collecting mixed audio signals and extracting effective speech signals from the mixed audio signals; S2, converts the effective voice signal into an echo-cancelled voice signal, which is then uploaded to the cloud-based large model and processed to generate a feedback voice stream; S3 receives the feedback voice stream transmitted from the large model in the cloud and plays it through the local rendering module.
2. The method for intelligent voice interaction processing according to claim 1, characterized in that: The real-time communication device includes a microphone array, and step S1 includes: S11, collecting a mixed audio signal through a microphone array on a real-time communication device, where the mixed audio signal includes a user command voice signal and a device playback sound superposition signal corresponding to a device playback sound; S12, separating a component of a device playback sound superposition signal from the mixed audio signal, and calculating an energy mean thereof as a reference energy baseline; S13, extracting the frequency spectrum characteristics of the user command voice signal, and when the energy peak of the user command voice signal exceeds the preset multiple of the reference energy baseline and the frequency spectrum characteristics match, extracting the effective voice signal therein.
3. The method for intelligent voice interaction processing according to claim 2, characterized in that: The step S2 comprises: S21, preprocessing the valid voice signal using an adaptive echo cancellation algorithm to obtain an echo-cancelled voice signal; S22, in the WebRTC transmission environment, encapsulate the echo-cancelled voice signal into an RTP data packet, and upload it to the cloud-based large model based on the UDP protocol after encapsulation; S23, the cloud-based large model parses the streaming input RTP data packets in real time to generate a feedback voice stream.
4. The method for intelligent voice interaction processing according to claim 3, characterized in that: The step S21 comprises: S31, collecting the superimposed signal of the sound played by the device in real time; S32, using an adaptive echo cancellation algorithm, taking the device playback sound superposition signal as a reference signal, and generating a reverse compensation signal that matches the device playback sound propagation path; S33, performing phase-aligned digital domain superposition processing on the reverse compensation signal and the mixed audio signal to cancel the echo interference in the device playback sound superposition signal to obtain an echo-cancelled voice signal.
5. The method for intelligent voice interaction processing according to claim 4, characterized in that: The step S22 comprises: S41, in a WebRTC transmission environment, encoding the echo-cancelled voice signal into an RTP data packet sequence according to a dynamic bit rate parameter, wherein the dynamic bit rate parameter is generated based on a signal energy mean value within a preset time window; S42, detecting the energy peak of the user voice in the RTP data packet sequence, and when the energy peak exceeds the reference energy baseline and the spectrum feature matches the preset template, generating a VAD start instruction carrying a timestamp and a coding frame rate; S43, transmitting the VAD start instruction and the dynamic bit rate parameter to the cloud through the WebSocket protocol, and synchronously associating and storing the bandwidth, delay and packet loss rate parameters collected in real time; S44, calculating the encapsulation interval of the RTP data packet sequence according to the ratio of the target bit rate of the dynamic bit rate parameter to the bandwidth, and adjusting the redundant coding ratio based on the packet loss rate difference; S45, based on the UDP protocol, the RTP data packet sequence is uploaded to the cloud-based large model according to the encapsulation interval and redundancy ratio, wherein when the voice energy mean is lower than the preset reference threshold for N consecutive cycles, a VAD stop instruction is generated and the upload is terminated until a new instruction is triggered.
6. The method for intelligent voice interaction processing according to claim 5, characterized in that: The step S23 comprises: S51, performing real-time parsing on the RTP data packets inputted in streaming mode in a preset interval time window, and outputting a sequence of text segments; S52, generating a final intention instruction based on the context relevance of the text segment sequence; S53: Generate a feedback voice stream according to the intention instruction, wherein parameters of the feedback voice stream match an audio sampling rate of the real-time communication device.
7. The method for intelligent voice interaction processing according to claim 1, characterized in that: The step S3 comprises: S71 receives the feedback voice stream transmitted from the cloud-based large model through the UDP protocol, and plays it to the user through the local rendering module to achieve low-latency interaction.
8. A real-time communication device, characterized in that: The real-time communication device is communicatively connected to the cloud-based large model, and the real-time communication device is used to receive command voice from the user, and the real-time communication device is used to implement the voice interaction intelligent processing method as described in any one of claims 1 to 7.