Intelligent voice interaction method and device based on voice large model, terminal and medium

Through the synergy between voiceprint recognition and speech synthesis model, and the user's personalized data generates interactive content, the problems of single tone, closed data and insufficient expansion in the existing voice interaction system are solved, and a personalized and emotional voice interaction experience is achieved.

CN120472909APending Publication Date: 2025-08-12NANJING KUKAI SMART SCREEN TECH CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510844278.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing voice interaction system cannot achieve personalized tone reproduction, insufficient data autonomy and limited expansion, resulting in a lack of emotional resonance and exclusiveness of the interaction.

Method used

Through the synergy between voiceprint recognition and speech synthesis model, interactive content is generated by combining user personalized data, user customization of prompt words is supported, and user identity is verified through voiceprint recognition, and user's voice characteristics and personal story text are used to generate personalized conversations.

Benefits of technology

It realizes a highly personalized voice interaction experience, enhances the user's emotional resonance and realism, supports the construction of private memory and expands application scenarios, and improves the reliability of the system and user satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472909A_ABST
    Figure CN120472909A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent voice interaction method and device based on a voice large model, a terminal and a medium, and belongs to the technical field of voice recognition and interaction, and the method comprises the steps: recognizing the voiceprint information of a speaker through a voiceprint recognition model when a wake-up instruction is received; when the identity of the current pronunciator is the specified user registered through voiceprint, acquiring question information of the current pronunciator, calling a prompt word of a character event text of the specified user corresponding to the current pronunciator from a cloud end, and combining with a question proposed by the current pronunciator to obtain the question information of the current pronunciator; generating event text information related to a specified user character corresponding to the current pronunciator; and calling the speech synthesis model, synthesizing the generated event text information into speech according to the tone characteristics of the speaker, and outputting and playing the speech. According to the method, the user is supported to upload the character as the cue word, and the cue word is combined with the large model, so that the dialogue content generation based on the personalized information of the user is realized, and convenience is provided for the use of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech recognition and interaction technology, and in particular to an intelligent speech interaction method, device, intelligent terminal and storage medium based on a large speech model. Background Art

[0002] With the development of science and technology and the continuous improvement of people's living standards, the voice interaction function of smart terminals is becoming more and more popular.

[0003] Existing voice interaction systems generally adopt a standardized architecture, which presents three significant drawbacks. First, regarding personalized experience, traditional systems use a fixed voice library for speech synthesis, which cannot accurately reproduce the voice of a specific user. Furthermore, voiceprint recognition is only used for basic identity verification and is completely disconnected from the subsequent conversation content, resulting in a lack of emotional resonance in the interaction process. Second, regarding data autonomy, existing solutions strictly rely on a general knowledge base, preventing users from injecting personalized data (such as family members' life experiences and important anniversaries). The system output is limited to public knowledge, making it difficult to build a conversational experience that reflects personal memories. Finally, regarding system scalability, the closed architecture cannot support the flexible integration of user-defined prompts, limiting the depth of application in scenarios such as family emotional interaction and educational companionship, and hindering the development of innovative applications such as digital reproduction of historical figures. These technical limitations make it difficult for existing voice interaction systems to meet the growing user demand for personalized, emotionally engaging intelligent interactions.

[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention

[0005] In order to solve the above technical problems, the present invention provides an intelligent voice interaction method, device, intelligent terminal and storage medium based on a large voice model, which has the advantages of realizing a highly personalized voice interaction experience, enhancing user emotional resonance, supporting user-defined data injection and improving system scalability; the present invention supports users to upload their own personal deeds as prompt words, and combine them with the large model to realize the generation of dialogue content based on user personalized information, which provides convenience for users.

[0006] The technical solution of this application is as follows: An intelligent voice interaction method based on a large voice model, comprising: The voice information of the designated user is obtained in advance, and a speech synthesis model that simulates the voice of the designated user is synthesized by the sound replication engine; the wake-up word audio of the designated user is obtained in advance, and the voiceprint recognition model is used to perform voiceprint recognition analysis on the wake-up word audio to complete the voiceprint registration of the designated user's voice wake-up; Obtaining in advance the uploaded biographical text of a designated user, and using the unique ID of the designated user's device as an identifier, saving the biographical text of the designated user as a prompt word in the cloud; When a wake-up command is received, the voiceprint information of the speaker is identified by the voiceprint recognition model to determine whether the current speaker is the designated user who has been registered through the voiceprint; When the current speaker is the designated user who has registered through voiceprint, the question information of the current speaker is obtained, and the prompt words of the character deeds text of the designated user corresponding to the current speaker are retrieved from the cloud, and the event text information related to the designated user character corresponding to the current speaker is generated in combination with the question raised by the current speaker; The speech synthesis model is called to synthesize the event text information related to the designated user character corresponding to the current speaker with the speaker's timbre characteristics and output it for playback.

[0007] In the intelligent voice interaction method based on a large voice model, the step of synthesizing a voice synthesis model that simulates the voice of a specified user using a voice replication engine using the pre-acquired voice information of the specified user comprises: Pre-acquire the voice information of the specified user; The sound reproduction engine is used to analyze and extract sound features from the sound information of the specified user; The extracted sound features are synthesized into a speech synthesis model that simulates the voice of the specified user.

[0008] The intelligent voice interaction method based on the large voice model, wherein the step of pre-acquiring the wake-up word audio of the designated user and performing voiceprint recognition analysis on the wake-up word audio through the voiceprint recognition model to complete the voiceprint registration of the designated user's voice wake-up includes: The wake-up word audio of the specified user is obtained in advance, and the voiceprint recognition analysis of the wake-up word audio is performed on the voiceprint recognition model to complete the voiceprint registration of the specified user's voice wake-up, and the voiceprint information of the specified user is associated with the specified user in the system.

[0009] The intelligent voice interaction method based on the large voice model, wherein, when receiving the wake-up instruction, the step of identifying the voiceprint information of the speaker through the voiceprint recognition model and determining whether the current speaker is the designated user who has been registered through the voiceprint includes: Pre-acquire the uploaded biographical text of the specified user; The character deeds text of the designated user is added with the unique ID of the corresponding designated user device, and the unique ID of the designated user device is used as an identifier and stored in the cloud as a prompt word.

[0010] The intelligent voice interaction method based on the large voice model, wherein, when receiving the wake-up instruction, the step of identifying the voiceprint information of the speaker through the voiceprint recognition model and determining whether the current speaker is the designated user who has been registered through the voiceprint includes: When receiving the wake-up command, obtain the voice information of the current speaker; Recognize the voice information of the current speaker and identify the voiceprint information of the speaker through the voiceprint recognition model; Comparing the recognized voice information with the pre-stored voiceprint information of the designated user registered through voiceprint, and determining whether the current speaker is the designated user who has been registered through voiceprint; When the recognized voice information is consistent with the pre-stored voiceprint information of the designated user who has been registered through voiceprint, it is determined that the current speaker is the designated user who has been registered through voiceprint.

[0011] The intelligent voice interaction method based on the voice big model, wherein, when the current speaker is the designated user who has been registered through voiceprint, the question information of the current speaker is obtained, and the prompt words of the character deeds text of the designated user corresponding to the current speaker are retrieved from the cloud, and the event text information related to the designated user character corresponding to the current speaker is generated in combination with the question raised by the current speaker. The steps include: Obtain the information of the current speaker and match it with the identified speaker's identity information to confirm whether it is the designated user who has been registered through voiceprint; When the identity of the current interlocutor is the designated user who has been registered through voiceprint, the question information of the current interlocutor is obtained, and the prompt words of the character deeds text of the designated user corresponding to the current interlocutor are retrieved from the cloud, and combined with the questions raised by the current user, event text information related to the designated user character corresponding to the current speaker is generated.

[0012] The intelligent voice interaction method based on the large voice model, wherein the step of calling the voice synthesis model, synthesizing the generated event text information related to the designated user character corresponding to the current speaker with the speaker's timbre characteristics and outputting the speech for playback includes: Get the generated event text information related to the specified user character corresponding to the current speaker; The speech synthesis model is called, and the event text information generated and related to the designated user character corresponding to the current speaker is synthesized into speech output and played through the speech synthesis model using the speaker's timbre characteristics.

[0013] An intelligent voice interaction device based on a large voice model, wherein the device comprises: A speech synthesis model construction module is used to synthesize a speech synthesis model that simulates the voice of a specified user using a sound replication engine based on the acquired voice information of the specified user. A voiceprint registration module is used to pre-acquire the wake-up word audio of a specified user, and perform voiceprint recognition analysis on the wake-up word audio through a voiceprint recognition model to complete the voiceprint registration of the specified user's voice wake-up; A prompt word saving module is used to pre-acquire the uploaded character deeds text of the designated user, and save the character deeds text of the designated user as a prompt word in the cloud using the unique ID of the designated user device as an identifier; A wake-up module is configured to, upon receiving a wake-up instruction, identify the voiceprint information of the speaker through the voiceprint recognition model and determine whether the current speaker is the designated user who has registered through the voiceprint; The event text information generation module is used to obtain the question information of the current speaker when the current speaker is the designated user who has been registered through voiceprint, and retrieve the prompt words of the character deeds text of the designated user corresponding to the current speaker from the cloud, and generate event text information related to the designated user character corresponding to the current speaker in combination with the question raised by the current speaker; The speaker's timbre output module is used to call the speech synthesis model, synthesize the event text information related to the designated user character corresponding to the current speaker with the speaker's timbre characteristics, and output it for playback.

[0014] An intelligent terminal includes a memory and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by one or more processors, including the method for executing any one of the methods described above.

[0015] A computer-readable storage medium, wherein when instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform any one of the methods described above.

[0016] As can be seen from the above, the intelligent voice interaction method, device, intelligent terminal, and storage medium based on a large voice model provided by this application, through the synergy of voiceprint recognition and speech synthesis model, dynamically generates interactive content with emotional resonance in combination with user personalized data, achieving a highly personalized voice interaction experience. This solves the problems of fixed timbre, closed data, and insufficient scalability of traditional systems, and has the advantages of enhancing user emotional connection, supporting the construction of private memories, and expanding application scenarios. The present invention also has the following effects: 1) Improved personalized voice interaction experience: Through voice replication and personalized prompts, the system can conduct conversations in the voice of a specific person, and the content of the conversation is closely centered around the personal stories provided by the user, making the interaction more realistic and emotionally connected.

[0017] 2) Enables highly customized interactive content: The user-uploaded prompt function breaks the limitations of the shared large model, enabling highly customized conversation content to meet the needs of different users for personalized interaction.

[0018] 3) Accurate identity recognition and interaction: Voiceprint recognition and multi-information confirmation mechanisms ensure the accuracy of interaction, avoid misidentification and incorrect responses, and improve system reliability and user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0020] Figure 1 This is a flow chart of the intelligent voice interaction method based on the large voice model provided in Example 1 of the present invention.

[0021] Figure 2 This is a schematic diagram of the voiceprint registration process of the intelligent voice interaction method based on the large voice model provided in Example 2 of the present invention.

[0022] Figure 3 This is a schematic diagram of the interactive processing flow of the intelligent voice interaction method based on the large voice model provided in Example 2 of the present invention.

[0023] Figure 4 A block diagram showing the principles of an embodiment of an intelligent voice interaction device based on a large voice model provided by the present invention.

[0024] Figure 5 This is a block diagram of the internal structure of the smart terminal provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0025] In order to make the purpose, technical solutions and advantages of the present invention more clear and distinct, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0026] It should be noted that if the embodiments of the present invention involve directional indications (such as up, down, left, right, front, back, etc.), the directional indications are only used to explain the relative position relationship, movement status, etc. between the various components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indications will also change accordingly.

[0027] In existing technologies, voice interaction systems generally use standardized speech synthesis technology, which generates uniform timbres based on a fixed timbre library and cannot reproduce the personalized timbre characteristics of a specific user. Voiceprint recognition technology is only used for account login verification and is not linked to the content of the conversation, resulting in a lack of emotional depth in the interaction process. Users are unable to independently upload personalized data, such as family anniversaries or life experiences, and the system relies on a generalized knowledge base of shared large models, making it difficult to achieve personalized interactions. For example, in family scenarios, smart devices cannot generate conversation content based on the user's personal experiences, and the interactive experience lacks a sense of exclusivity.

[0028] To address these issues, we are considering how to achieve personalized voice generation through sound feature extraction and replication, addressing the monotony of existing speech synthesis techniques. We are also exploring the feasibility of associating voiceprint information with user-specific data to address the disconnect between voiceprint recognition and conversation content. Furthermore, we are exploring the association mechanism between cloud storage structures and device identifiers to enable targeted access to personalized prompts.

[0029] Therefore, this application proposes an intelligent voice interaction method based on a large voice model.

[0030] Example 1 like Figure 1 As shown, an intelligent voice interaction method based on a large voice model in embodiment 1 of the present invention includes the following steps: Step S100: Using the pre-acquired voice information of a designated user, a voice synthesis model simulating the voice of the designated user is synthesized by a voice replication engine; obtaining the wake-up word audio of the designated user in advance, and performing voiceprint recognition analysis on the wake-up word audio by a voiceprint recognition model to complete voiceprint registration for the designated user's voice wake-up; In this step, regarding the collection of voice information and the establishment of a synthesis model, the target user's voice characteristics (such as voice samples) are first obtained through recording or other means. These samples are then fed into the voice replication engine, which generates a speech synthesis model based on these samples to simulate the voice. This model can be used to synthesize a voice that closely resembles the target user, achieving voice imitation or simulation.

[0031] In this invention, voiceprint recognition and registration for wake-up word audio involves capturing the target user's wake-up word audio (e.g., "Hello, assistant") and inputting it into a voiceprint recognition model. This model analyzes the wake-up word's voiceprint characteristics (including speaker characteristics, timbre, intonation, etc.), performs voiceprint recognition, and completes voiceprint registration. This creates a voiceprint profile specific to the target user, enabling the system to identify and confirm the specific user's wake-up.

[0032] Step S200: pre-acquire the uploaded biographical text of the designated user, and use the unique ID of the designated user's device as an identifier to save the biographical text as a prompt word in the cloud; In the embodiment of this step, the text acquisition process collects and uploads the target user's "character story text" in advance, that is, text content that describes the user's life, achievements, background, etc. in detail.

[0033] This is then stored in the cloud using the user's device's unique ID as an identifier. This information is then stored in the cloud as "prompts." Prompts are crucial information that guide the model in generating content or answering questions.

[0034] In this step, storing the user's story text as prompts can help the AI more accurately reflect the user's background information during conversations or content generation, making the generated content more tailored to the user's identity and characteristics. Quickly retrieving relevant prompts based on the user's device's unique ID eliminates the need to re-enter the user's background each time, enabling efficient, automated personalized services and facilitating quick access. Furthermore, based on the specific user's story, the AI's responses and performance will be richer and more authentic, improving the user experience and enhancing authenticity and credibility. This invention is applicable to a variety of application scenarios, including intelligent assistants, virtual characters, personalized Q&A, and content customization.

[0035] Step S300: When a wake-up instruction is received, the voiceprint information of the speaker is identified by the voiceprint recognition model to determine whether the current speaker is the designated user who has been registered through the voiceprint; In this embodiment, when the system hears a wake-up command (e.g., "Hey, assistant"), it automatically initiates the voiceprint recognition process. The voiceprint recognition model analyzes the voice characteristics of the current speaker and extracts voiceprint information (i.e., a unique individual voice fingerprint). The extracted voiceprint is then compared with a previously registered database of target users. If the comparison matches the target user's voiceprint characteristics, the system confirms the current speaker as a "registered target user." Otherwise, specific functions may not be activated or a prompt indicating unrecognized identity may be displayed.

[0036] The benefits of this are: 1) Enhanced security: Voiceprint recognition confirms the speaker's identity, preventing unauthorized users from accidentally triggering the system and effectively protecting user privacy and account security. 2) Personalized experience: After identifying the user's identity, the system can provide customized content or services for each user, such as personalized greetings and preference settings. 3) Automatic recognition eliminates steps: Users no longer need to manually enter passwords or other authentication methods; voice recognition automatically completes identity confirmation, improving convenience. 4) Preventing abuse and misuse: Ensuring that the system only responds to verified user commands, enhancing the accuracy and reliability of operations.

[0037] Step S400: When the current speaker is the designated user who has registered via voiceprint, the question information of the current speaker is obtained, and prompt words of the character deeds text of the designated user corresponding to the current speaker are retrieved from the cloud. In combination with the question raised by the current speaker, event text information related to the designated user character corresponding to the current speaker is generated; In this embodiment, identity verification begins. First, the system uses voiceprint recognition to confirm that the current speaker is the registered target user. After identity verification, the system retrieves the question or request posed by the current speaker. Next, the system retrieves the prompt word for the "character story text" corresponding to the user (i.e., stored background information and keywords) from the cloud. Content generation then proceeds, combining the current question with the user's "character story" prompt word and using a language generation model (such as GPT) to generate event text or answers related to the target user.

[0038] This invention can generate responses based on the user's background information (biographical text) that are more relevant to the user's identity and historical background, making the conversation more natural and heartwarming. Furthermore, by incorporating biographical information, the generated content not only addresses the question but also incorporates the user's stories or achievements, making the information more comprehensive, in-depth, rich, and relevant. Furthermore, during the interaction, users will feel that the system understands and respects their personal circumstances, enhancing their trust and satisfaction, and improving the user experience. Furthermore, the generated content can be intelligently adjusted in terms of perspective and depth based on the user's background and question content, enabling highly personalized services and intelligent customization.

[0039] Step S500: calling the speech synthesis model, synthesizing the generated event text information related to the designated user character corresponding to the current speaker with the speaker's timbre characteristics into speech and outputting it for playback.

[0040] In this embodiment, the system uses previously generated "textual information about user-related events" as input to a text-to-speech (TTS) model. During synthesis, the model leverages the speaker's timbre to make the generated speech sound like the target user's voice. Ultimately, the system synthesizes a piece of audio content in the user's voice, which the user can then hear.

[0041] In this way, the present invention utilizes the user's voice characteristics to generate voices that sound more real and closer to the user, thereby enhancing the sense of immersion and trust and achieving a personalized sound experience; and the present invention combines the content of "character background events", and the system can use the user's voice to tell stories, answer questions or interpret scenes, making the interaction more vivid and achieving rich content expression; and the present invention allows the user who asks the question to hear the synthesized content of his own voice, which will make the user feel that the system is more intimate and more personalized, enhance the usage experience, and improve user satisfaction.

[0042] In more detail, the present application proposes to synthesize the voice information of the designated user into a speech synthesis model through a sound replication engine in advance, pre-acquire the wake-up word audio of the designated user to complete the voiceprint registration, store the user's character deeds text in the cloud as a prompt word with the device's unique ID as the identifier, and after verifying the user's identity through voiceprint recognition in the wake-up stage, generate event text information by combining the user's questions and the cloud prompt word, and finally call the speech synthesis model to output personalized voice.

[0043] Among them, the sound replication engine refers to a system that extracts user voice features through machine learning algorithms. Specifically, it can use deep neural networks to perform spectral analysis and feature encoding on user voice samples to generate an acoustic model containing parameters such as fundamental frequency, resonance peaks, and pronunciation rhythm, which is used to solve the problem of single voice synthesis timbre in existing technologies. The voiceprint recognition model refers to an identity authentication system based on biometrics. Specifically, it can use a Gaussian mixture model or a convolutional neural network to extract Mel-frequency cepstral coefficient features of the wake-up word audio to establish a user voiceprint feature library for linking identity authentication with data retrieval. The device unique ID refers to the identity identification code of the hardware device. Specifically, the IMEI code or MAC address can be used to generate a hash value, which is used to establish a one-to-one correspondence between the user identity and the prompt word when storing data in the cloud, to ensure the accuracy of data retrieval.

[0044] Specifically, a personalized acoustic model is constructed by collecting user voice samples, and the user's voiceprint characteristics are bound and stored with the device ID. When the user issues a wake-up command, the system extracts the current voiceprint characteristics for matching and verification. After confirming the identity, the system retrieves the prompt word data associated with the device ID from the cloud. Based on the user's real-time questions, natural language processing technology is used to generate a response text that incorporates the user's personal experience. Finally, the pre-generated personalized voice model is used for speech synthesis output.

[0045] Compared to existing technologies, existing solutions rely on fixed voice libraries and are unable to generate personalized voices. This solution, however, utilizes a sound replication engine to achieve user-specific voice synthesis. Existing voiceprint recognition is only used for account login, while this solution directly links voiceprint verification results to cloud-based prompt word invocation. Existing systems rely on a shared knowledge base, while this solution allows users to upload personalized story text and enable targeted invocation via device ID, expanding the customization dimension of interactive content.

[0046] Through the above technical solutions, this application achieves user-specific voice interaction output, enabling smart devices to generate customized conversational content based on the user's personal experiences. By binding the device ID to cloud data, the independent storage and precise access of user data are guaranteed. By employing a collaborative working model of voiceprint recognition and speech synthesis, a complete closed loop of identity verification, content generation, and voice output is established, enhancing the exclusiveness and emotional depth of the interaction process.

[0047] The present application further proposes obtaining the voice information of a designated user in advance, analyzing and extracting voice features from the obtained voice information of the designated user through a voice replication engine, and synthesizing the extracted voice features into a speech synthesis model that simulates the voice of the designated user.

[0048] Among them, sound information refers to the voice data collected by the user through the recording device. Specifically, it can be implemented by an audio file containing the user reading a preset text, such as a recording file of more than 10 seconds provided by the user. The sound replication engine refers to a speech feature extraction algorithm based on deep learning. Specifically, a convolutional neural network model can be used to model the sound spectrum to separate and extract acoustic feature parameters such as timbre, pitch, and speaking speed. Sound features refer to a set of acoustic parameters that can characterize the pronunciation habits of an individual. Specifically, they can include fundamental frequency contours, resonance peak distributions, pronunciation duration patterns, etc., which are obtained by quantitatively analyzing the short-time spectrum characteristics of the user's voice. The speech synthesis model refers to a speech generation module based on parameter synthesis technology. Specifically, waveform splicing or statistical parameter synthesis methods can be used to map the extracted sound features into synthesized speech with the user's timbre characteristics.

[0049] Specifically, users record voice samples containing their own pronunciations through terminal devices, such as reading a text containing different phoneme distributions. The sound reproduction engine divides the voice samples into frames and extracts the Mel-frequency cepstral coefficients of each frame of voice as basic features. The timbre features in the voice samples are further analyzed through a pre-trained acoustic model, such as using a generative adversarial network to distinguish the user's unique pronunciation style. After encoding the extracted feature parameters into low-dimensional vectors, they are combined with the text-to-speech synthesis model to build a customized synthesis model that can generate user-specific timbre speech according to the input text. When the model receives text input during the deployment phase, it calls the stored acoustic feature vector to control the speech synthesizer to output a speech waveform with the target timbre.

[0050] Compared to existing technologies, traditional speech synthesis systems rely on a fixed library of voices and are unable to generate customized speech based on individual user voices. This solution, however, uses a voice reproduction engine to extract unique user acoustic features, enabling the synthesized speech to accurately replicate the user's voice characteristics. For example, in family settings, it can generate conversational responses that mimic the voices of relatives, avoiding the mechanical feel of generic voices.

[0051] Through the above technical solution, this application achieves high-fidelity reproduction of a user's personalized voice, resolving the existing problem of speech synthesis being disconnected from the user's identity. By building a user-specific speech synthesis model, the speech output by the interactive system has timbre characteristics consistent with the user's own. For example, when recounting a user's personal experiences, their authentic voice is used, enhancing the authenticity and emotional resonance of the interaction.

[0052] This application further proposes to pre-acquire the wake-up word audio of the specified user, perform voiceprint recognition analysis on the wake-up word audio through a voiceprint recognition model to complete the voiceprint registration of the specified user's voice wake-up, and associate the voiceprint information of the specified user with the specified user in the system.

[0053] The wake-up word audio refers to a pre-recorded voice clip containing specific words or phrases by the user. Specifically, it can be implemented using a voice file of 3-5 seconds in length, for example, the user says "start the conversation" as a trigger command. The voiceprint recognition model refers to a voiceprint feature extraction and matching algorithm built on a deep neural network. Specifically, it can be implemented using Mel-frequency cepstral coefficients combined with a convolutional neural network to extract the speaker's pitch, formant, and other biometric features from the audio. Voiceprint registration refers to the process of establishing a unique binding relationship between voiceprint features and user identity. Specifically, it can be achieved by associating and storing the voiceprint feature vector with the user ID for feature comparison during subsequent identity verification.

[0054] Specifically, when a user uses the system for the first time, they need to record at least three sets of wake-up word audio samples, and the length of each set of audio can be set to 2 seconds. The voiceprint recognition model performs spectral analysis and feature extraction on multiple sets of samples to generate a digital vector with the user's unique voiceprint characteristics. This vector is mapped to the user's registration information in the system, for example, the feature vector is bound to the UUID of the user account and stored in the database. When the user subsequently issues a voice command, the system calculates the similarity between the real-time collected voice data and the registered voiceprint features. If the match reaches the preset threshold, the user is determined to be a legitimate user.

[0055] Compared to existing technologies, traditional voiceprint recognition is only used for authentication during account login and is not linked to interactive content. This solution deeply binds voiceprint registration information to the user's identity, making voiceprint features not only an access credential but also a fundamental data source for personalized interactions. While existing technologies store voiceprint data in isolation, this solution establishes a voiceprint-identity association mechanism that enables the system to dynamically access user-specific data during interactions.

[0056] Through the above technical solution, this application achieves a system-level association between voiceprint characteristics and user identity, effectively preventing unauthorized users from impersonating others' voiceprints to perform illegal operations. In family scenarios, parents can use registered voiceprint characteristics to ensure that only they can wake up and access their children's growth records. At the same time, the system can automatically retrieve the corresponding user's personalized data set based on the voiceprint identity, avoiding data confusion between different users.

[0057] This application further proposes pre-acquiring the uploaded biographical text of the designated user, adding the unique ID of the corresponding designated user device to the biographical text of the designated user, and storing it in the cloud as a prompt word using the unique ID of the designated user device as an identifier.

[0058] Among them, the biographical text of a specified user refers to a text description containing the user's personalized experiences or events. Specifically, it can be implemented by using documents uploaded by the user, text entered in the input box, or text content generated by voice transcription. Its role is to provide a data basis for the subsequent generation of personalized conversation content. The unique ID identifier refers to the identity identification code bound to the user's device. Specifically, it can be implemented by using the device serial number, MAC address, or an encrypted string assigned by the system. Its role is to establish a unique association between the user identity and the cloud storage data to ensure the accuracy of data retrieval. Cloud storage refers to the persistent storage of data through a distributed server cluster. Specifically, it can be implemented by using object storage services or database services. Its role is to enable cross-device access and efficient retrieval of data.

[0059] Specifically, users upload text content containing personal experiences or important events through terminal devices, such as descriptions of family anniversaries, records of professional achievements, or travel experiences. The system automatically extracts the unique identification code of the device, such as a 32-bit hash string, and binds the identification code to the text content as an index identifier. After the binding is completed, the text content is divided into several semantic paragraphs, and keywords are extracted and an inverted index is established through natural language processing technology, and finally stored in the cloud database. When the user initiates a voice interaction request, the system quickly retrieves the corresponding text content from the cloud based on the device's unique ID, inputs it into the language model as context, and generates response content related to the user's personal experience.

[0060] Compared with existing technologies, traditional solutions rely on a shared knowledge base and cannot link to user-defined data, resulting in a lack of targeted interactive content. This solution uses a unique device ID to accurately match user data with the interactive system, allowing users to independently inject personalized content and access it in real time, solving the problems of insufficient data autonomy and limited scalability.

[0061] Through the above technical solution, this application realizes the structured storage and efficient call of user-defined data, so that the voice interaction content can generate responses based on the user's personalized deeds. For example, when answering family-related questions, it automatically quotes the anniversary records uploaded by the user, significantly improving the sense of exclusivity and scene adaptability of the interaction.

[0062] The present application further proposes that when a wake-up command is received, the voice information of the current speaker is obtained; the voice information of the current speaker is identified, and the voiceprint information of the speaker is identified through a voiceprint recognition model; the identified voice information is compared with the pre-stored voiceprint information of the designated user who has registered through the voiceprint, and it is determined whether the identity of the current speaker is the designated user who has registered through the voiceprint; when the identified voice information is consistent with the pre-stored voiceprint information of the designated user who has registered through the voiceprint, it is determined that the identity of the current speaker is the designated user who has registered through the voiceprint.

[0063] Among them, the wake-up command refers to the operation instruction that activates the voiceprint recognition process through specific trigger conditions. Specifically, it can be implemented by voice keyword detection or physical button triggering to start the identity verification process. The voiceprint recognition model refers to the voiceprint feature extraction and matching model constructed by machine learning algorithms. Specifically, it can use deep neural networks to perform spectral analysis and feature encoding on voice signals to extract distinctive voiceprint features from audio. Voiceprint information comparison refers to the similarity calculation between the real-time collected voiceprint feature vector and the pre-stored voiceprint feature vector. Specifically, it can be implemented by cosine similarity or Euclidean distance algorithm to determine whether the identity of the speaker matches that of a registered user.

[0064] Specifically, when the device detects a wake-up command, it first collects a voice clip of the current speaker as input data. The voiceprint recognition model performs spectral analysis and feature encoding on this voice clip, generating a feature vector that represents the speaker's identity. The system then compares the generated feature vector with a database of registered users' voiceprint features stored in the cloud. If the similarity exceeds a preset threshold, the user is deemed legitimate. For example, after the user speaks the wake-up word, the system extracts the voiceprint features within 0.5 seconds and matches them with pre-stored features in real time. A successful match activates the personalized interaction process.

[0065] Compared to existing technologies, voiceprint recognition is only used for static verification during the account login phase and is not deeply integrated into the interaction process. This solution embeds voiceprint recognition into the wake-up phase, completing dynamic identity verification at the beginning of the interaction, preventing unauthorized users from triggering personalized services. It also uses a real-time feature comparison mechanism, which can dynamically adjust the matching accuracy to adapt to different environmental noise conditions, compared to traditional fixed threshold verification methods.

[0066] Through the above technical solution, this application achieves real-time identity verification before interaction, ensuring that subsequent personalized content is only available to authorized users. This solution solves the problem of voiceprint recognition being separated from conversation content in existing technologies. Through the pre-authentication mechanism, it prevents non-registered users from obtaining sensitive information, while also providing a secure foundation for subsequent conversation generation based on user-specific deeds.

[0067] This application further proposes that when the identity of the current interlocutor is a designated user who has been registered through voiceprint, the question information of the current interlocutor is obtained, and the prompt words of the character deeds text of the designated user corresponding to the current interlocutor are retrieved from the cloud, and combined with the questions raised by the current user, event text information related to the designated user character corresponding to the current speaker is generated.

[0068] Among them, the information of the current interlocutor refers to the voiceprint features of the speaker extracted by the voiceprint recognition model. Specifically, it can be represented by a voiceprint feature vector and used to match the voiceprint information of registered users. Identity matching confirmation refers to calculating the similarity between the voiceprint features collected in real time and the registered voiceprint stored in the cloud. Specifically, it can be implemented by using the cosine similarity algorithm or a deep neural network model to verify the legitimacy of the user's identity. The prompt words of the character's deeds text refer to structured text data stored in the cloud with the user's device unique ID as the index. Specifically, it can be stored in JSON format or natural language paragraph form to provide personalized knowledge input for the voice model. Event text information generation refers to combining user questions with cloud prompt words and then generating answer content through large model inference. Specifically, it can be implemented by using the context completion technology of the pre-trained language model to output conversation content related to the user's personal experience.

[0069] Specifically, after voiceprint verification is passed, the system collects the user's voice questions in real time and converts them into text input. The cloud server retrieves the corresponding personalized prompt vocabulary based on the device's unique ID. This prompt vocabulary contains exclusive information such as life events and important dates uploaded by the user. The language model performs semantic correlation analysis between the question text and the prompt words, extracts key information nodes through an attention mechanism, and ultimately generates a natural language response that incorporates the user's personal experience. For example, when a user asks for suggestions for family anniversary activities, the system can generate personalized recommendations based on past celebration records stored in the cloud.

[0070] Compared to existing technologies, traditional voice interaction systems are unable to dynamically link voiceprint authentication information with personalized knowledge bases, resulting in conversations being limited to generic answers from public knowledge bases. This solution precisely binds user identity and private data through a unique device ID, enabling large models to generate personalized interaction content based on user-specific information, resolving the data silos and interaction homogeneity issues inherent in existing technologies.

[0071] Through the above technical solution, this application achieves a deep integration of user private data and voice interaction systems, enabling smart devices to provide customized conversational services based on the user's personal experience. This method effectively improves the relevance and emotional resonance of interactive content, while ensuring the user's independent management rights for data, providing reliable technical support for scenarios such as family memory preservation and personalized education.

[0072] This application further proposes calling a speech synthesis model to output and play the generated event text information related to the designated user character corresponding to the current speaker through the speech synthesis model and synthesized speech using the speaker's timbre characteristics.

[0073] Among them, the event text information generated related to the designated user character corresponding to the current speaker refers to the content dynamically generated by combining the user's questions and the personalized deeds text stored in the cloud. Specifically, a natural language processing model can be used to identify the intention of the user's questions, and the prompt words stored in the cloud can be associated to generate text, thereby achieving a deep connection between the conversation content and the user's personal experience.

[0074] Among them, the speech synthesis model refers to a timbre replication model established based on the user's voiceprint characteristics. Specifically, a deep learning network can be used to extract the acoustic feature parameters of the user's voice, such as the fundamental frequency and resonance peak, and a waveform synthesis algorithm can be used to reconstruct a speech signal with the user's timbre characteristics, so that the output speech has timbre characteristics that are highly similar to those of the real user.

[0075] Specifically, after confirming that the speaker is a registered user, the system uses the voice recognition module to obtain the text of the user's question and then retrieves the character's deeds prompt associated with the user's unique ID from the cloud. The natural language processing engine semantically associates the question text with the deeds prompt, generating a response text containing a personalized event description. This text is input into a pre-generated speech synthesis model, which synthesizes the speech waveform based on the user's voiceprint feature parameters extracted during the registration phase, and ultimately plays a voice response with the user's unique timbre through the audio output device.

[0076] Compared to existing technologies, existing voice interaction systems use a fixed timbre library for speech synthesis, which cannot achieve personalized reproduction of the user's timbre, and the conversation content relies on a general knowledge base. This solution achieves dual customization of timbre characteristics and conversation content by dynamically combining voiceprint feature modeling with personalized prompt words, allowing the interaction process to have a user-specific timbre and integrate their personalized deeds data.

[0077] Through the above technical solution, this application solves the problems of the monotony of speech synthesis timbre and the generalization of dialogue content in the existing technology, realizes personalized voice output based on user voiceprint and customized content generation combined with user deeds, enables the intelligent interactive system to output voice responses with user-exclusive timbre and containing their personal experiences, improves the authenticity and emotional relevance of the interactive experience, and supports users to independently upload personalized data to expand the dimension of interactive content.

[0078] The present invention is further described in detail below through another specific application embodiment.

[0079] Example 2 like Figure 2 and Figure 3 As shown, the second embodiment provides an intelligent voice interaction method based on a large voice model, including: S10: When the audio of a specific person (i.e., the voice signal of a specified user) is obtained, the process proceeds to S11; when the wake-up word audio of the specified user is obtained, the process proceeds to step S12; S11, obtaining the input audio of a specific person and synthesizing a speech synthesis model through a sound replication engine; then proceeding to S14; S12: Register the wake-up word audio of the specified user through the voiceprint recognition model, and then proceed to S14; S13, receiving the character deeds text uploaded by the designated user, adding the character deeds text uploaded by the designated user to the unique ID of the designated user's device, and saving it in the cloud as a prompt word for the corresponding user with the unique ID of the designated user's device; When the system receives the wake-up command, it enters S21; S21, the wake-up command is triggered and enters S22; S22, voiceprint recognition determines the identity of the speaker, and then proceeds to S23; S23, obtain the information of the person to be talked to and confirm it, then enter S24; S24, the large model generates event information in combination with the prompt word, and then enters S25; S25. The speech synthesis model synthesizes speech and plays it.

[0080] That is, the specific application embodiment of the present invention includes a sound replication and voiceprint registration stage, a user uploading prompt word stage and an intelligent interaction stage.

[0081] The voice replication and voiceprint registration phase involves extracting the specific person's audio input from a designated user. The system uses a voice replication engine to analyze the specific person's audio input and extract voice features such as timbre, intonation, and speaking speed, thereby synthesizing a speech synthesis model that can simulate the specific person's voice. Simultaneously, the user inputs the specific person's wake-up word audio, which the voiceprint recognition model identifies and analyzes, completing voiceprint registration and associating the voiceprint information with the system.

[0082] During the user-uploaded prompt phase, designated users can upload textual information, such as personal stories, to the system. The system uses the unique device ID of the designated user as an identifier and stores this information in the cloud as prompts. This process allows users to add personalized content, significantly enriching the system's interactive resources and distinguishing itself from existing technologies that rely solely on shared, large-scale models.

[0083] The intelligent interaction stage of the embodiment of the present application includes: when the system receives a wake-up command, the voiceprint recognition model will identify the voiceprint information of the speaker, thereby determining the identity of the speaker. The system obtains the information of the person to be spoken to and matches it with the identified identity information of the speaker. Based on the confirmed identity of the speaker, the large model retrieves the prompt words uploaded by the corresponding user from the cloud, and generates event information related to the person in combination with the questions raised by the user. Finally, the system calls the previously synthesized speech synthesis model, synthesizes the generated text content into speech with the speaker's timbre characteristics, and plays it.

[0084] As can be seen above, this invention employs personalized voice and voiceprint processing technology. It utilizes a voice replication engine to synthesize a speech synthesis model for a specific person (i.e., a designated user). This is then combined with a voiceprint recognition model to complete voiceprint registration, accurately extracting and utilizing the user's personalized voice characteristics and voiceprint information. Through voice replication and personalized prompts, the system can conduct conversations in the voice of a specific person, closely centered around user-provided personal stories. This creates a more realistic and emotionally connected interaction, achieving a personalized voice interaction experience.

[0085] Furthermore, this invention integrates user-uploaded prompts with a large model: users can upload their own personal stories as prompts, which are then combined with the large model to generate conversation content based on their personalized information. This is a key innovation that distinguishes it from existing shared large-scale interactions. By enabling users to upload their own prompts, this invention overcomes the limitations of the shared large-scale model and enables highly customized conversation content, meeting the needs of different users for personalized interaction and achieving highly customized interactive content.

[0086] The present invention also employs intelligent interaction technology with multi-information confirmation. During the interaction process, the system ensures accuracy by identifying the speaker's identity and confirming the interlocutor's information. This information then generates relevant event information, which is then synthesized and played back using the speaker's timbre. This system utilizes voiceprint recognition and multi-information confirmation to ensure accurate interaction, avoid misidentification and incorrect responses, and enhance system reliability and user experience. It enables precise identity recognition and interaction.

[0087] Exemplary devices like Figure 4 As shown, an embodiment of the present invention provides an intelligent voice interaction device based on a large voice model, the device comprising: The speech synthesis model construction module 310 is used to synthesize a speech synthesis model that simulates the voice of a specified user using a sound reproduction engine based on the acquired voice information of the specified user. The voiceprint registration module 320 is used to pre-acquire the wake-up word audio of the specified user, and perform voiceprint recognition analysis on the wake-up word audio through the voiceprint recognition model to complete the voiceprint registration of the specified user's voice wake-up; The prompt word saving module 330 is used to pre-acquire the uploaded character deeds text of the designated user, and save the character deeds text of the designated user as a prompt word in the cloud using the unique ID of the designated user device as an identifier; The wake-up module 340 is configured to, upon receiving a wake-up instruction, identify the voiceprint information of the speaker through the voiceprint recognition model and determine whether the current speaker is the designated user who has registered through the voiceprint; The event text information generation module 350 is used to obtain the question information of the current speaker when the current speaker is the designated user who has been registered through voiceprint, and retrieve the prompt words of the character deeds text of the designated user corresponding to the current speaker from the cloud, and generate event text information related to the designated user character corresponding to the current speaker in combination with the question raised by the current speaker; The speaker's timbre output module 360 is used to call the speech synthesis model, synthesize the event text information related to the designated user character corresponding to the current speaker with the speaker's timbre characteristics, and output it for playback.

[0088] The speech synthesis model construction module is a component that uses a voice replication engine to analyze user voice characteristics and generate a personalized voice model. Specifically, it uses deep learning algorithms to extract timbre, intonation, and rhythm features from user-provided audio samples to construct a synthesis model that can reproduce the user's voice. The voiceprint registration module is a functional unit that uses voiceprint recognition technology to bind user identities. Specifically, it compares the user-provided wake-up word audio with a preset voiceprint feature library to associate voiceprint information with the user's account. The prompt word storage module is a cloud-based storage unit that binds user-uploaded personalized text data to the device's unique identifier. Specifically, it uses a distributed database to store text content in key-value pairs to ensure unique and secure data retrieval. The wake-up module is a logical control unit that verifies user identity based on voiceprint information. Specifically, it collects voice signals in real time, extracts voiceprint features, and performs similarity matching with registered voiceprint data. The event text information generation module is a processing unit that combines user questions and cloud-based prompt words to generate personalized responses. Specifically, it uses natural language processing technology to semantically associate user questions with character story text to generate contextually coherent responses. The speaker's voice output module refers to a synthesis unit that converts text information into user-sounding speech. Specifically, it can map text input into acoustic parameters that conform to the user's voice characteristics through a speech synthesis model, and drive the vocoder to output audio.

[0089] Specifically, the device generates a user-specific voice model through the speech synthesis model construction module. The voiceprint registration module binds the user's identity to the voiceprint characteristics, and the prompt word storage module stores the personalized text uploaded by the user in the cloud. Once the wake-up module confirms the user's identity through voiceprint matching, the event text information generation module uses the cloud-based prompt word and combines it with the user's question to generate a customized response. Finally, the speaker voice output module plays the response in the user's voice. These modules work together through data flows and control signals to achieve a complete closed loop from identity verification to personalized voice output.

[0090] Compared to existing technologies, existing voice interaction systems use a fixed timbre library and cannot link to user-defined data, resulting in a lack of unique interactive content. This solution, however, uses a modular design to achieve timbre replication, voiceprint binding, and dynamic access to personalized data. This allows speech synthesis to accurately restore the user's timbre and, combined with user-uploaded story text, generates customized conversation content, resolving the disconnect between standardized responses and user needs.

[0091] Through the above technical solution, this application enables the voice interaction system to achieve dual binding of identity verification and content access based on the user's voiceprint. It dynamically generates conversation content incorporating the user's personalized data through cloud-based prompts and outputs responses with the user's unique voice, thereby enhancing the exclusiveness and emotional resonance of the interaction. Furthermore, the modular architecture design supports users to independently inject data such as character deeds, expanding the system's application capabilities in scenarios such as family memory preservation and historical figure restoration.

[0092] Based on the above embodiment, the present invention also provides an intelligent terminal, whose principle block diagram can be shown as follows: Figure 5 The intelligent terminal includes a processor, a memory, a network interface, a display screen, and a database connected via a system bus.

[0093] An intelligent terminal of the present application further includes one or more programs, wherein one or more programs are stored in a memory and configured to be executed by one or more processors, including a method for executing an intelligent voice interaction method based on a large voice model.

[0094] Memory refers to the physical device used to store program code, specifically a solid-state drive or flash memory chip. It is used to store user voiceprint data, speech synthesis models, and biographical text. A processor refers to the computing unit that executes program instructions. It can be implemented as a multi-core central processing unit or embedded chip, and is used to complete voiceprint recognition, text generation, and speech synthesis tasks. A program refers to a code collection containing executable instructions, specifically a compiled binary file or scripting language. It is used to control the terminal to complete user authentication, data retrieval, and speech output processes.

[0095] Specifically, when the terminal boots up, the program stored in the memory is loaded into the processor for execution. During program execution, the voiceprint recognition module first collects the user's voice and extracts voiceprint features, which are then matched against pre-stored voiceprint data to verify the user's identity. Once verified, the processor retrieves the biographical text associated with the user's unique ID from the cloud and, combined with the real-time voice input, generates personalized event text. The speech synthesis model is then invoked to convert the text into a voice signal with the user's timbre, which is then played back through the audio output module.

[0096] In some specific embodiments, the memory can be configured as an encrypted storage area to ensure the security of user voiceprint data. The processor can adopt an edge computing architecture to perform voiceprint recognition and speech synthesis tasks locally, reducing cloud reliance. The program can support a dynamic update mechanism, allowing users to independently upload new character deeds text and update the prompt vocabulary through the terminal interface.

[0097] Compared to existing technologies, existing smart terminals can only run general voice interaction programs and are unable to replicate user voices or integrate personalized data. However, this solution, by combining local terminal storage and processing capabilities, allows users to independently inject personalized data and achieve customized voice interaction, breaking through the limitations of traditional systems that rely on shared knowledge bases.

[0098] Through the above technical solution, this application solves the technical problem that existing voice interaction terminals cannot support user-customized timbre and personalized content integration, realizes voiceprint recognition and data retrieval based on the user's unique ID, ensures that the interactive content is accurately matched with the user's identity, and at the same time reduces the cloud data transmission delay through localized processing, thereby improving the real-time performance and privacy security of speech synthesis.

[0099] The present application further proposes a computer-readable storage medium. When the instructions in the storage medium are executed by the processor of an electronic device, the electronic device is enabled to execute an intelligent voice interaction method based on a large voice model. The method includes pre-building a voice synthesis model and a voiceprint registration module, saving personalized prompt words in the cloud, calling the prompt words after voiceprint verification to generate event text and outputting it through synthesized voice.

[0100] Among them, computer-readable storage medium refers to a physical carrier used to store program instructions, which can be implemented by a solid-state hard drive, a USB flash drive or a cloud storage space. Its function is to provide executable program code for electronic devices to realize voice interaction functions.

[0101] Among them, the execution of instructions by the processor of the electronic device refers to parsing and running the program code in the storage medium through the central processing unit, which can be implemented by multi-threaded processing or distributed computing architecture. Its function is to convert the logical steps of the voice interaction method into an operational process that can be executed by the device.

[0102] Specifically, when the instructions in the storage medium are loaded by the processor, the electronic device will perform the following process: first, the sound reproduction engine analyzes the user's voice characteristics to build a speech synthesis model, and simultaneously completes the voiceprint registration and stores the user's deeds text in the cloud with the device's unique ID; when the wake-up command is detected, the user's identity is verified through voiceprint comparison. If a match is found, the deeds text in the cloud is retrieved to generate the event content, and finally the speech synthesis model outputs the voice in the user's voice. For example, the user can upload a personal experience text, and the device will combine this text to generate a response content containing specific events when answering questions.

[0103] Compared to existing technologies, which rely on fixed voice libraries and are unable to integrate user-defined data, this solution standardizes the deployment of program instructions through storage media, enabling personalized voice interaction processes across different devices. While voiceprints are used solely for account login, this solution integrates voiceprint verification with the invocation of action text, achieving dual matching of identity and content. For example, existing technologies cannot generate customized conversations based on user-uploaded family anniversaries. However, this solution uses cloud-based prompts to make interactive content event-related.

[0104] Through the above technical solution, this application enables electronic devices to replicate the user's voice and combine it with personalized content through executable instructions, solving the problems of voice immobilization and insufficient data autonomy in voice interaction. Users can expand the scope of conversation content by uploading customized travel text. For example, the device can generate voice responses containing locations and times based on the user's travel experiences, realizing a personalized interactive experience.

[0105] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. For those skilled in the art, various modifications and variations of the present application are possible. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. An intelligent voice interaction method based on a large voice model, characterized in that: include: The voice information of the specified user is obtained in advance and synthesized into a speech synthesis model that simulates the voice of the specified user through the voice replication engine; Pre-acquire the wake-up word audio of the specified user, and perform voiceprint recognition analysis on the wake-up word audio through the voiceprint recognition model to complete the voiceprint registration of the specified user's voice wake-up; Obtaining in advance the uploaded biographical text of a designated user, and using the unique ID of the designated user's device as an identifier, saving the biographical text of the designated user as a prompt word in the cloud; When a wake-up command is received, the voiceprint information of the speaker is identified by the voiceprint recognition model to determine whether the current speaker is the designated user who has been registered through the voiceprint; When the current speaker is the designated user who has registered through voiceprint, the question information of the current speaker is obtained, and the prompt words of the character deeds text of the designated user corresponding to the current speaker are retrieved from the cloud, and the event text information related to the designated user character corresponding to the current speaker is generated in combination with the question raised by the current speaker; The speech synthesis model is called to synthesize the event text information related to the designated user character corresponding to the current speaker with the speaker's timbre characteristics and output it for playback.

2. The intelligent voice interaction method based on the voice big model according to claim 1 is characterized in that The step of synthesizing the voice synthesis model simulating the voice of the designated user by using the voice reproduction engine with the voice information of the designated user obtained in advance includes: Pre-acquire the voice information of the specified user; The sound reproduction engine is used to analyze and extract sound features from the sound information of the specified user; The extracted sound features are synthesized into a speech synthesis model that simulates the voice of the specified user.

3. The intelligent voice interaction method based on a large voice model according to claim 1, characterized in that: The step of pre-acquiring a wake-up word audio of a designated user and performing voiceprint recognition analysis on the wake-up word audio through a voiceprint recognition model to complete voiceprint registration for voice wake-up of the designated user includes: The wake-up word audio of the specified user is obtained in advance, and the voiceprint recognition analysis of the wake-up word audio is performed on the voiceprint recognition model to complete the voiceprint registration of the specified user's voice wake-up, and the voiceprint information of the specified user is associated with the specified user in the system.

4. The intelligent voice interaction method based on a large voice model according to claim 1, characterized in that: The step of, upon receiving a wake-up instruction, identifying the voiceprint information of the speaker through the voiceprint recognition model, and determining whether the current speaker is the designated user who has been registered through the voiceprint includes: Pre-acquire the uploaded biographical text of the specified user; The character deeds text of the designated user is added with the unique ID of the corresponding designated user device, and the unique ID of the designated user device is used as an identifier and stored in the cloud as a prompt word.

5. The intelligent voice interaction method based on a large voice model according to claim 1, characterized in that: The step of, upon receiving a wake-up instruction, identifying the voiceprint information of the speaker through the voiceprint recognition model, and determining whether the current speaker is the designated user who has been registered through the voiceprint includes: When receiving the wake-up command, obtain the voice information of the current speaker; Recognize the voice information of the current speaker and identify the voiceprint information of the speaker through the voiceprint recognition model; Comparing the recognized voice information with the pre-stored voiceprint information of the designated user registered through voiceprint, and determining whether the current speaker is the designated user who has been registered through voiceprint; When the recognized voice information is consistent with the pre-stored voiceprint information of the designated user who has been registered through voiceprint, it is determined that the current speaker is the designated user who has been registered through voiceprint.

6. The intelligent voice interaction method based on a large voice model according to claim 1, characterized in that: When the current speaker is the designated user who has registered through voiceprint, the steps of obtaining the question information of the current speaker, retrieving the prompt words of the character deeds text of the designated user corresponding to the current speaker from the cloud, and generating event text information related to the designated user character corresponding to the current speaker in combination with the question raised by the current speaker include: Obtain the information of the current speaker and match it with the identified speaker's identity information to confirm whether it is the designated user who has been registered through voiceprint; When the identity of the current interlocutor is the designated user who has been registered through voiceprint, the question information of the current interlocutor is obtained, and the prompt words of the character deeds text of the designated user corresponding to the current interlocutor are retrieved from the cloud, and combined with the questions raised by the current user, event text information related to the designated user character corresponding to the current speaker is generated.

7. The intelligent voice interaction method based on a large voice model according to claim 1, characterized in that: The step of calling the speech synthesis model, synthesizing the generated event text information related to the designated user character corresponding to the current speaker with the speaker's timbre characteristics and outputting the speech for playback comprises: Get the generated event text information related to the specified user character corresponding to the current speaker; The speech synthesis model is called, and the event text information generated and related to the designated user character corresponding to the current speaker is synthesized into speech output and played through the speech synthesis model using the speaker's timbre characteristics.

8. An intelligent voice interaction device based on a large voice model, characterized in that: The device comprises: A speech synthesis model construction module is used to synthesize a speech synthesis model that simulates the voice of a specified user using a sound replication engine based on the acquired voice information of the specified user. A voiceprint registration module is used to pre-acquire the wake-up word audio of a specified user, and perform voiceprint recognition analysis on the wake-up word audio through a voiceprint recognition model to complete the voiceprint registration of the specified user's voice wake-up; A prompt word saving module is used to pre-acquire the uploaded character deeds text of the designated user, and save the character deeds text of the designated user as a prompt word in the cloud using the unique ID of the designated user device as an identifier; A wake-up module is configured to, upon receiving a wake-up instruction, identify the voiceprint information of the speaker through the voiceprint recognition model and determine whether the current speaker is the designated user who has registered through the voiceprint; The event text information generation module is used to obtain the question information of the current speaker when the current speaker is the designated user who has been registered through voiceprint, and retrieve the prompt words of the character deeds text of the designated user corresponding to the current speaker from the cloud, and generate event text information related to the designated user character corresponding to the current speaker in combination with the question raised by the current speaker; The speaker's timbre output module is used to call the speech synthesis model, synthesize the event text information related to the designated user character corresponding to the current speaker with the speaker's timbre characteristics, and output it for playback.

9. An intelligent terminal, characterized in that: The device comprises a memory and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by one or more processors, and the one or more programs include being used to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Multi-agent matching method and device, storage medium and electronic device

    CN121506152A

  • Voice interaction recognition method and device

    CN121565166A

  • Voice communication method and device based on asynchronous voiceprint prefetching and separated transmission

    CN121687071A

  • AI partner terminal system and method based on cloud edge collaboration

    CN121924140A