This application discloses an intelligent voice dialogue method and
system based on a large
language model, belonging to the field of
speech processing technology. The method includes: acquiring an audio
stream of the user's surrounding environment;
processing the audio
stream by frame and determining whether it contains speech to obtain a valid audio
stream; converting the valid audio stream into a
text string using a
deep learning model; performing deep analysis using a large
language model to dynamically parse the user's immediate state, potential needs, and emotional inclinations, generating
natural language response text, and converting it into an audio stream for playback using
speech synthesis technology. This method leverages the powerful
natural language understanding capabilities of the large
language model to dynamically and deeply understand the user's immediate state, potential needs, and emotional tone, forming a richer, more vivid, and context-dependent user understanding. This results in a more natural, considerate, and accurate
response generation and
service provision, achieving a more human-centered human-computer voice interaction experience.