Entity token-based large language model context hot switching method and related device
By dynamically switching virtual character configuration information using physical tokens, combined with user input and physical interaction signals, the problem of lack of emotional companionship in human-computer interaction in existing technologies is solved, and immersive character interaction and emotional connection are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUAN GUAN MONOGATARI (BEIJING) TECHNOLOGY CO LTD
- Filing Date
- 2026-03-06
- Publication Date
- 2026-07-03
AI Technical Summary
Existing human-computer interaction technologies cannot achieve immersive character interaction, lack a sense of emotional companionship and realism, and cannot meet users' strong demand for character interaction.
By using an entity token-based method, a unique identifier is received to query a mapping database to obtain virtual role configuration information, dynamically construct a dialogue context, and automatically switch virtual roles when the identifier changes. Voice responses are generated by combining user input and physical interaction signals.
It enables rich interactive scenarios and emotional companionship between users and virtual characters, provides an interactive method that integrates physical entities and AI personalities, and enhances the sense of ritual and emotional connection.
Smart Images

Figure CN122332408A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, specifically to a method for hot-switching context of a large language model based on entity tokens, a cloud server, a smart terminal, a system for hot-switching context of a large language model based on entity tokens, and a computer-readable storage medium. Background Technology
[0002] Currently, people's demand for human-computer interaction has gradually evolved from tool-oriented to companion-oriented. In particular, for some families, the elderly, children, or IP enthusiasts who need companionship, these groups have a strong demand for role-playing interaction to achieve emotional connection.
[0003] However, among related technologies, functional voice assistants, represented by smart speakers, only focus on completing task instructions, resulting in a single personality setting and cumbersome switching processes. Physically triggered hardware, represented by interactive toys, lacks the ability to understand and generate natural language, making it impossible to achieve true "dialogue" and "companionship." While AI character chatbots on pure software platforms perform well in terms of personality setting and dialogue, they are detached from physical entities, lacking the realism of tactile dimension and the physical presence as a carrier of emotions, making it difficult to establish deep emotional attachment and a sense of immersion in the scene.
[0004] Therefore, the mainstream human-computer interaction methods in related technologies cannot achieve immersive role interaction, resulting in a lack of complete technical carriers for users' emotional companionship needs. Summary of the Invention
[0005] This application aims to address at least one of the technical problems existing in the prior art. To this end, this application proposes a method and related device for hot-switching of context in a large language model based on physical tokens, realizing an interactive mode that integrates physical entities and AI personalities, enabling users to obtain a full sense of ritual, companionship, and emotional connection.
[0006] In a first aspect, embodiments of this application provide a method for hot-switching of context in a large language model based on an entity token. The method includes: receiving a unique identifier of an entity token sent from a smart terminal; querying a mapping database based on the unique identifier to obtain virtual role configuration information bound to the unique identifier; wherein the virtual role configuration information can be updated by modifying cloud configuration data; receiving user input information uploaded by the smart terminal; dynamically constructing a dialogue context containing the target role setting based on the virtual role configuration information and the user input information; sending the dialogue context to a large language model for processing and receiving a text response generated by the large language model; converting the text response into a corresponding voice signal based on the virtual role configuration information and sending it to the smart terminal; and automatically interrupting the current dialogue context construction and large language model processing tasks based on the virtual role configuration information when the unique identifier of the entity token changes to a new unique identifier, and loading the corresponding new virtual role configuration information in real time based on the new unique identifier to overwrite the dialogue context.
[0007] In some embodiments, the virtual character configuration information includes system prompt words parameters; dynamically constructing a dialogue context containing the target character settings based on the virtual character configuration information and user input information includes: generating prompt information containing the target character settings and current dialogue content in real time based on the system prompt words parameters and combined with user input information to form a dialogue context.
[0008] In some embodiments, dynamically constructing a dialogue context that includes the target role setting further includes: receiving sensor signals uploaded by a smart terminal that reflect physical interaction with an entity token; converting the sensor signals into corresponding context semantic representations; and integrating the context semantic representations as supplementary context information into the dialogue context.
[0009] In some embodiments, the virtual character configuration information also includes emotion response parameters; the method further includes: dynamically generating emotion indication information that matches the current dialogue content in the dialogue context based on the semantic analysis results of the user input information and the emotion response parameters, so as to guide the large language model to generate a text response with corresponding emotional color.
[0010] In some embodiments, the virtual character configuration information further includes speech synthesis parameters; converting text responses into corresponding speech signals includes: synthesizing text responses into speech signals with specific timbre characteristics set by the target character based on the speech synthesis parameters.
[0011] In some embodiments, the method further includes: after receiving the unique identifier of the new entity token, obtaining the corresponding new virtual role configuration information based on the unique identifier of the new entity token; wherein obtaining the corresponding new virtual role configuration information is automatically triggered in response to the change of the unique identifier.
[0012] Secondly, embodiments of this application provide a cloud server, comprising: a first receiving module configured to receive a unique identifier of an entity token sent from a smart terminal; a query module configured to query a mapping database based on the unique identifier to obtain virtual role configuration information bound to the unique identifier; wherein the virtual role configuration information can be updated by modifying cloud configuration data; a second receiving module configured to receive user input information uploaded by the smart terminal; a construction module configured to dynamically construct a dialogue context containing target role settings based on the virtual role configuration information and the user input information; a processing module configured to send the dialogue context to a large language model for processing and receive a text response generated by the large language model; a feedback module configured to convert the text response into a corresponding voice signal and send it to the smart terminal based on the virtual role configuration information; and a response module configured to automatically interrupt the current dialogue context construction and large language model processing tasks based on the virtual role configuration information when the unique identifier of the entity token is changed to a new unique identifier, and load the corresponding new virtual role configuration information in real time based on the new unique identifier to overwrite the dialogue context.
[0013] Thirdly, embodiments of this application provide a smart terminal, including: a near-field communication module configured to read a unique identifier from a physical token; a collection module configured to collect user input information; a sensing module configured to collect sensor signals reflecting physical interaction with the physical token; a far-field communication module configured to send the unique identifier, user input information, and sensor signals to a cloud server; and an audio output module configured to receive and play voice signals from the cloud server; wherein the voice signals are generated by the cloud server after processing the user input information based on virtual character configuration information bound to the unique identifier.
[0014] Fourthly, embodiments of this application provide a large language model context hot-switching system based on entity tokens, including the cloud server provided in the second aspect and the smart terminal provided in the third aspect.
[0015] Fifthly, embodiments of this application provide a computer-readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the entity token-based large language model context hot-switching method described in the first aspect.
[0016] The technical solution provided in this application first receives a unique identifier of an entity token sent from a smart terminal. Then, it queries a mapping database based on the unique identifier to obtain virtual role configuration information bound to the unique identifier. This virtual role configuration information can be updated by modifying cloud-based configuration data. Next, it receives user input information uploaded by the smart terminal. Based on the virtual role configuration information and the user input information, it dynamically constructs a dialogue context containing the target role's settings. This dialogue context is sent to a large language model for processing, and the text response generated by the large language model is received. Based on the virtual role configuration information, the text response is converted into a corresponding voice signal and sent to the smart terminal. Furthermore, when the identifier changes, the current construction and processing operations are automatically interrupted, and the virtual role configuration information corresponding to the new identifier is loaded, overwriting the current dialogue context. This application receives and identifies the identifier of an entity token, obtains the virtual character configuration information corresponding to the identifier, combines it with the information received from the user input to the terminal, constructs the dialogue environment of the virtual character, and sends the current dialogue context to a language big model to generate corresponding voice signal responses. Through the voice representation of the virtual character, the user can perform corresponding intelligent human-computer interaction at the physical entity level, and can also quickly and freely switch virtual characters, effectively providing users with rich interactive scenarios and emotional companionship.
[0017] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0018] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:
[0019] Figure 1 A schematic diagram of a large language model context hot-switching system architecture based on entity tokens provided in an embodiment of this application; Figure 2 A flowchart of a large language model context hot-switching method based on entity tokens provided in this application embodiment; Figure 3 This is a schematic diagram illustrating the relationship between physical tokens and virtual role configurations provided in an embodiment of this application. Figure 4 This is an interaction timing diagram of the context hot-switching method for a large language model based on entity tokens according to an embodiment of this application; Figure 5 A schematic diagram of a cloud server provided in an embodiment of this application; Figure 6 A schematic diagram of a smart terminal provided in an embodiment of this application; Figure 7 A schematic diagram of a context-switching system for a large language model based on entity tokens provided in an embodiment of this application.
[0020] Figure reference numerals: 110-Client; 120-Cloud server; 130-External server; 111-Entity token; 112-Smart terminal; 121-Cloud gateway; 122-Business layer; 1221-Entity management service module; 123-Algorithm layer; 1231-Dialogue flow orchestration engine; 1232-Dynamic context builder; 130-External server; 1201-First receiving module; 1202-Query module; 1203-Second receiving module; 1204-Construction module; 1205-Processing module; 1206-Feedback module; 1207-Response module; 1121-Near-field communication module; 1122-Acquisition module; 1123-Sensing module; 1124-Far-field communication module; 1125-Audio output module; 700-Entity token-based large language model context hot-switching system. Detailed Implementation
[0021] Embodiments of this application will now be described in more detail with reference to the accompanying drawings. While some embodiments of this application are shown in the drawings, it should be understood that this application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this application. It should be understood that the drawings and embodiments of this application are for illustrative purposes only and are not intended to limit the scope of protection of this application.
[0022] It should be understood that the steps described in the method embodiments of this application may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this application is not limited in this respect.
[0023] refer to Figure 1 This is a schematic diagram of a context-switching system architecture for a large language model based on entity tokens provided in an embodiment of this application.
[0024] Specifically, the large language model context hot-switching system architecture based on physical tokens provided in this application embodiment includes a client 110, a cloud server 120, and an external server 130. The client 110 includes a physical token 111 and a smart terminal 112. When the physical token 111 interacts with the smart terminal 112 (such as touching, connecting, or entering the recognition area), the unique identifier can be read through the corresponding communication protocol and sent to the cloud server 120 through the network communication protocol. The cloud server 120 includes a cloud gateway 121, a business layer 122, and an algorithm layer 123. The cloud server 120 communicates with the client 110 through the cloud gateway 121. The business layer 122 includes an entity management service module 1221, which includes an interface to a cloud database and can manage the mapping relationship between unique identifiers and virtual role configurations. It can also add, delete, modify, and query the virtual role configurations corresponding to unique identifiers through the background. The algorithm layer 123 includes a dialogue flow orchestration engine 1231 and a dynamic context builder 1232. The dialogue flow orchestration engine 1231 adopts a dynamic orchestration mode to load the knowledge base of the corresponding role, the memory module storing the user's and role's historical dialogues, and the preset prompt information templates for the large language model on demand for different virtual roles and dialogue requests. The dynamic context builder 1232 can construct the corresponding dialogue context based on the user's input information and the virtual role configuration, and send the dialogue context to the dialogue flow orchestration engine 1231 for input to the large language model of the external server 130.
[0025] refer to Figure 2 The flowchart below shows a method for hot-switching context of a large language model based on entity tokens, provided in an embodiment of this application.
[0026] Step S201: Receive the unique identifier of the entity token sent from the smart terminal.
[0027] Specifically, the physical token provided by this invention is a physical carrier with an embedded storage module. This storage module stores a unique identifier (UID). The physical token also incorporates a passive or active communication module, such as an NFC tag, RFID chip, or Bluetooth unit; the unique identifier (UID) includes, but is not limited to, the UID of an NFC tag, the TID of an RFID chip, the MAC address of a Bluetooth device, or an encrypted serial number. The physical token of this application is configured such that when a user places the physical token on or near the sensing area of a compatible smart terminal, it can automatically and wirelessly transmit its internally stored unique identifier to the smart terminal via near-field communication (such as NFC, RFID, Bluetooth, etc.).
[0028] Step S202: Query the mapping relationship database based on the unique identifier to obtain the virtual character configuration information bound to the unique identifier; wherein, the virtual character configuration information can be updated by modifying the cloud configuration data.
[0029] Specifically, after receiving the UID, the smart terminal uploads it to the cloud server. The cloud server pre-stores a mapping database between UIDs and virtual character configuration information. This mapping database is a structured data collection stored on the cloud server, including a unique identifier field and a virtual character configuration information field. The unique identifier field and the virtual character configuration information field have a mapping relationship. The mapping database can exist in the form of a relational database table, key-value pair storage, or document database, etc. The virtual character configuration information is a structured data object or configuration file that fully defines the interactive personality and behavioral characteristics of a virtual character.
[0030] By querying the mapping relationship between UID and virtual character configuration, the cloud server can immediately obtain the virtual character settings corresponding to this UID (such as system prompts, voice parameters, personality templates, etc.), and construct the corresponding context for subsequent user dialogue based on the corresponding virtual character settings. Therefore, this physical token essentially acts as a "physical key," and its unique identifier is the core basis for the cloud system to identify user intent, dynamically load, and switch virtual character personalities.
[0031] refer to Figure 3 This is a schematic diagram illustrating the relationship between the entity token and the virtual role configuration provided in this application embodiment.
[0032] Specifically, in Figure 3In this context, "User" represents a registered user of the system, whose attributes include ID and Username, which are the user's unique identifier and username, respectively. In this embodiment, "Figure" can represent a physical, uniquely identified doll, whose attributes include UID, UserID, CharacterID, and Remark. UID is the unique hardware identifier of the figure, UserID is the ID of the user who owns the token, CharacterID is the ID of the virtual character to which the token is currently bound or mapped, and Remark contains other remarks. "Character" represents an activatable virtual character configuration. A complete digital configuration of a virtual character includes attributes such as ID, Name, SystemPrompt, VoiceID, TTSParams, and Capabilities. Among them, ID is the unique identifier of the virtual character, Name is the character name, SystemPrompt is the system prompt word template or content used to define the character's personality, background, and behavioral guidelines, VoiceID is the identifier of the corresponding voice model for the character, TTSParams are detailed parameters of speech synthesis (such as speech rate, tone, etc.), usually stored in a flexible JSON format, and Capabilities is a list of extended capabilities that the character possesses (such as available tools, knowledge base scope, etc.).
[0033] In this embodiment of the application, all logic and resources defining the interactive behavior of the virtual character are stored in the cloud server in the form of data. Therefore, the data information configured for the virtual character can be updated by modifying the configuration data in the cloud server, without the need to modify or upgrade the firmware or local storage of the smart terminal. This can effectively avoid problems such as system update lag and complex update operations.
[0034] Step S203: Receive user input information uploaded by the smart terminal.
[0035] Specifically, after obtaining the virtual character configuration information bound to the unique identifier, the cloud server also needs to receive user input information uploaded to the cloud server by the smart terminal. The user input information includes, but is not limited to, user voice input received by the smart terminal or text information directly input by the user to the smart terminal. For voice input, after the smart terminal uploads the voice input signal to the cloud server, the cloud server will parse it to obtain the corresponding text information.
[0036] Step S204: Dynamically construct a dialogue context containing the target character's settings based on the virtual character configuration information and user input information.
[0037] Specifically, based on the aforementioned role configuration information and user input information, the cloud server first initializes the dialogue workflow through the dialogue flow orchestration engine, i.e., acquiring virtual role configuration information, user input information, and preset prompt information templates. Further, through a dynamic context builder, using the virtual role configuration information as a base and based on the current dialogue state (such as time, location, and past interaction history retrieved from the memory module), it generates an immediate and complete dialogue context as the initial instruction using the dialogue template. This dialogue context is a structured text generated based on the preset prompt information template, containing all the information needed to guide the large language model's response, such as the role, historical dialogue records, and current user input information.
[0038] As an optional embodiment, the virtual character configuration information includes system prompt word parameters; based on the virtual character configuration information and user input information, a dialogue context containing the target character settings is dynamically constructed, including: based on the system prompt word parameters and combined with user input information, generating prompt information containing the target character settings and current dialogue content in real time to constitute the dialogue context.
[0039] Specifically, the virtual character configuration information includes system prompt word parameters. These parameters are used to replace the corresponding placeholders in the dialogue template when user input is received, based on the content of the user input and the virtual character's personality and language habits. This generates prompt information containing the target virtual character and the current dialogue content, which is then used to construct the dialogue context. In addition, the virtual character configuration information also includes the character name, which can serve as a key semantic identifier for the dialogue context. When dynamically constructing the dialogue context, the "character name" can be directly injected into the prompt information.
[0040] In the process of constructing a dialogue context based on system prompt parameters and user input, the system first retrieves the preset prompt template text stored on the cloud server, corresponding to the identifier in the entity token. Then, it combines user input with the current dialogue, replacing placeholders in the preset prompt template text to generate a complete prompt. For example, a preset template might be "You are a {personality tag} {character name}, the current time is {current time}, and the user says to you: {user input information}." The system will replace the placeholders "{personality tag}", "{character name}", and "{current time}" with specific words like "lively", "Sun Wukong", and "3 PM", while simultaneously replacing the actual user input in {user input information}. This replaced prompt is then placed at the beginning of the large language model's dialogue context, forming a complete, structured dialogue context together with existing (or non-existent) historical dialogue records. By parameterizing and templating system prompts and integrating them with user input in real time, the system achieves high flexibility in virtual character settings and improves the relevance of virtual characters to contextual interactions.
[0041] As an optional embodiment, dynamically constructing a dialogue context that includes the target role setting further includes: receiving sensor signals uploaded by the smart terminal that reflect physical interaction with the entity token; converting the sensor signals into corresponding context semantic representations; and integrating the context semantic representations as supplementary context information into the dialogue context.
[0042] Specifically, the smart terminal can also integrate a variety of sensors to sense the user's physical interaction with the physical token. These sensors include, but are not limited to: tactile sensors, which can be sensors distributed inside the smart terminal or the physical token, used to detect the pressure, position, and pattern of contact actions such as "touching," "patting," and "hugging" that the user may perform on the physical token; posture sensors, which can be sensors such as accelerometers and gyroscopes placed inside the physical token, used to detect posture changes such as "picking up," "shaking," and "tilting"; and heartbeat or temperature simulation sensors, used to obtain more realistic bio-feedback signals. When the user engages in the above physical interactions with the physical token, the corresponding sensor signals are collected and preprocessed (such as filtering and feature extraction) in real time by the smart terminal, and then uploaded to the cloud server along with the current session identifier.
[0043] After receiving sensor signals from the smart terminal reflecting the user's physical interaction with the physical token, the cloud server converts the multi-sensor signals into corresponding contextual semantic representations. In this embodiment, the contextual semantic representation is preferably a natural language description. For example, when the user physically interacts with the physical token, after receiving the corresponding sensor signals, the cloud server can analyze the uploaded sensor signal sequence by calling the pre-trained behavior recognition model built into the cloud server. This identifies the specific physical interaction type (such as "continuous gentle stroking" or "quickly patting twice") and its possible intensity or emotional tone (such as "gently" or "excitedly"). Based on the recognition results, the abstract interaction type is converted into a natural language description text containing key verbs and adverbs. For example, it can be converted into: "[The user is gently stroking your head]" or "[You were lifted up happily by the user]". Furthermore, the generated contextual semantic representation is inserted into the constructed dialogue context, providing clear and complete guidance for the large language model's response generation process. By introducing sensor signals of physical interaction, the cloud server can generate responses related to the current physical interaction scenario, thereby improving the realism of the user's interaction with the virtual character.
[0044] As an optional embodiment, the virtual character configuration information also includes emotion response parameters; the method further includes: dynamically generating emotion indication information that matches the current dialogue content in the dialogue context based on the semantic analysis results of the user input information and the emotion response parameters, so as to guide the large language model to generate text responses with corresponding emotional coloring.
[0045] Specifically, the emotion response parameters are structured configuration information stored on a cloud server and bound to a specific virtual character. This information defines the virtual character's response strategies and expression styles to different emotional inputs. Semantic analysis is performed on the user input information to obtain the semantic analysis results, i.e., the user's current emotional information. Based on the semantic analysis results, the emotional state that the virtual character needs to exhibit is mapped. Semantic analysis of user input information can also be performed by analyzing the signal waveform of the user's input speech to obtain the emotional information of the user's current input speech.
[0046] It should be noted that the user's current emotional information can also be inferred based on pattern analysis of the aforementioned physical interaction signals. For example, if the sensor detects that the user is "continuously and gently stroking" the physical token, the cloud server can infer that the user is currently in a relatively relaxed, positive, or affectionate emotional state. Accordingly, the virtual character's emotional response parameters can be selected to tend towards friendly, pleasant, or gentle response parameters. If the sensor detects that the user is "briefly and forcefully slapping" the physical token, the cloud server infers that the user may be in an excited, anxious, or seeking strong feedback emotional state. In this case, the virtual character's emotional response parameters can be adjusted accordingly to a more energetic, confrontational, or emphatic dialogue style.
[0047] To enable virtual characters to better express emotions, emotion response parameters can also include language style templates, common interjections, or sentence structures corresponding to the emotion. For example, when the character's emotion is "excitement," the template will tend to use more exclamation marks and short phrases, thus obtaining emotion prompts corresponding to the user's input. Placing these emotion prompts within the constructed dialogue context is important. It's worth noting that emotion prompts are typically placed within system prompt parameters or adjacent to the user's current input to further ensure clear guidance for the large language model in generating its response. By setting emotion response parameters in the virtual character's configuration information, the system can perceive the user's emotional state and provide human-like emotional responses, greatly enhancing the sense of companionship and realism in the dialogue.
[0048] The following is an example of a specific contextual dialogue: Suppose a user places a physical token representing "Sun Wukong" on a smart terminal. The smart terminal reads its UID (e.g., UID_001) via NFC and obtains the corresponding virtual character configuration information from the mapping database on the cloud server. This configuration information includes a preset prompt message template, the content of which can be as follows (where {} are placeholders to be filled): "You are {character name}, personality {personality tag}. Current user location: {user location}. Current time: {current time}. You possess the following memories: {related memories}. Current context: {context description}. Please use {language style} to communicate with the user. The user's input for this interaction is: {user input}."
[0049] When a user inputs "Wukong, what's the weather like today?" into a smart terminal while simultaneously touching the doll's head, the smart terminal first performs voice recognition, converting the user's speech into text: "Wukong, what's the weather like today?". Further, it analyzes the semantics to determine the user's intent is to query the weather, recording the system prompts as "today" and "weather". Finally, it combines this with the touch gesture detected by the smart terminal's sensors, converting it into text description: "Detected user gently touching your head." Based on the above data, the cloud server replaces the placeholders in the template as follows: It retrieves the {character name} from the virtual character configuration information and replaces it with "Sun Wukong"; it retrieves the {personality tag} from the virtual character configuration information and replaces it accordingly; it retrieves the address information uploaded when the smart terminal is activated from the cloud server system and replaces {user location} accordingly; it retrieves the {current time} from the cloud server system clock and replaces it accordingly; it retrieves dialogue related to today's weather from the cloud server's memory module and replaces {related memory}; it replaces {context description} with the aforementioned physical action-integrated "The user is gently stroking your head"; it retrieves the {language style} of the current virtual character "Sun Wukong" from the virtual character configuration information and replaces it accordingly; and it replaces {user input} with "How's the weather today?".
[0050] After completing the above fusion and replacement, a complete system prompt word containing rich real-time information is generated. For example: "You are Sun Wukong, mischievous, active, and with a strong sense of justice. Your current location is: District B, City A. The current time is 2:30 PM on January 26, 2025. You have the following memories: [Recall: Yesterday, the user made a promise to you that if the weather was nice today, you would go on a 'picnic']. Current situation: [The user is gently stroking your head]. Please speak to the user using a classical novel tone combined with a lively conversational style. The user's input this time is: 'How's the weather today?'"
[0051] Based on the generated prompts, they are further combined with necessary historical dialogues to generate a complete dialogue context.
[0052] Step S205: Send the dialogue context to the large language model for processing and receive the text response generated by the large language model.
[0053] Specifically, after generating the dialogue context according to the above embodiments, the cloud server sends the dialogue context to one or more large language models for processing through a dialogue flow orchestration engine. The large language model can be a general base model or a language model that has been configured or fine-tuned. Upon receiving the dialogue context, the large language model, based on strictly adhering to the role system instructions at the beginning of the dialogue context, understanding user input and contextual information (including physical interaction descriptions), and leveraging its vast parametric knowledge, generates a coherent natural language text response that maintains a high degree of consistency with the target virtual character's settings in terms of content, language style, and emotional tone. This text response is then output to the cloud server.
[0054] Step S206: Based on the virtual character configuration information, convert the text reply into the corresponding voice signal and send it to the smart terminal.
[0055] Specifically, the cloud server's dialogue flow orchestration engine converts text replies into corresponding voice signals based on virtual character configuration information. The converted voice signal stream with virtual character identifiers is then sent to the smart terminal for playback via network communication protocols, thus completing the transformation of the virtual character from a "text character" to an "auditory character".
[0056] It should be noted that this application can also use a physical token to broadcast voice signals sent to a smart terminal. Using a physical token to broadcast voice signals can further enhance the interactive attributes between the virtual character and the user, thereby effectively improving the user experience.
[0057] As an optional embodiment, the virtual character configuration information also includes speech synthesis parameters; converting text responses into corresponding speech signals includes: synthesizing text responses into speech signals with specific timbre characteristics set by the target character based on the speech synthesis parameters.
[0058] Specifically, the virtual character configuration information also includes speech synthesis parameters, which define the control commands and resource identifiers required to convert text content into voice characteristics that conform to a specific target virtual character. These include, but are not limited to, timbre model identifiers and dedicated pronunciation libraries or terminologies. The timbre model identifier can point to the acoustic model or speaker ID of the target virtual character in the speech synthesis service on the cloud server. The dedicated pronunciation library or terminology typically targets the target virtual character's unique catchphrases, spell names, or proper nouns to provide standard or stylized pronunciation guidance, ensuring that the timbre output by the smart terminal is consistent with the virtual character. After the large language model on the cloud server generates a text response that conforms to the character's settings, the speech synthesis parameters and the text response are used as input. The speech synthesis engine performs timbre rendering to synthesize a speech signal with specific timbre characteristics that matches the target virtual character's settings. By binding the virtual character settings and speech synthesis parameters at the configuration level, a high degree of consistency between the virtual character's language content and auditory image is ensured, greatly enhancing the immersiveness and credibility of the interaction.
[0059] In step S207, in response to the change of the unique identifier of the entity token to a new unique identifier, the current task of constructing the dialogue context and processing the large language model based on the virtual role configuration information is automatically interrupted, and the corresponding new virtual role configuration information is loaded in real time based on the new unique identifier to overwrite the dialogue context.
[0060] Specifically, when the smart terminal outputs the voice signal converted from text reply, if the unique identifier of the physical token is updated at this time, the current task of constructing the dialogue context and processing the large language model based on the virtual role configuration information is immediately interrupted. The corresponding virtual role configuration is loaded from the cloud mapping database in real time according to the new unique identifier. Furthermore, the current dialogue context is overwritten according to the newly loaded virtual role configuration information to achieve seamless switching of the personality, language style and interaction capabilities of the new role, thereby realizing instant response and zero-latency hot switching of virtual roles when the physical token is changed.
[0061] As an optional embodiment, the method further includes: after receiving the unique identifier of the new entity token, obtaining the corresponding new virtual role configuration information based on the unique identifier of the new entity token; wherein, obtaining the corresponding new virtual role configuration information is automatically triggered in response to the change of the unique identifier.
[0062] Specifically, when a user switches the physical token that comes into contact with the smart terminal, the smart terminal can re-identify the unique identifier of the new physical token and upload it to the cloud server via the network. After receiving the new unique identifier, the cloud server will interrupt the ongoing preparation of subsequent dialogue logic based on the old identifier, and use the new identifier as the key query key to enable the entity management service module of the cloud server to access the mapping relationship database again to obtain the corresponding new virtual role configuration information, switch from the old role configuration to the newly obtained role configuration, so that the user can switch the target virtual role without waiting, confirming or restarting the device, thus realizing "hot switching" of virtual roles.
[0063] It should be noted that the acquisition of the corresponding new virtual character configuration information is automatically triggered based on the change of the unique identifier. Specifically, when the near-field communication reader of the smart terminal senses that the original physical token has moved out of the recognition area and the new physical token has entered the recognition area, the smart terminal will generate a switching identifier electrical signal. The switching identifier electrical signal and the new unique identifier are transmitted to the cloud server through the communication module. In the cloud server, the execution priority of the switching identifier electrical signal is higher than the dialogue context construction and large language model processing tasks. Therefore, the cloud server will interrupt the dialogue context construction and large language model processing tasks and access the mapping relationship database according to the new unique identifier to obtain the new virtual character configuration information, thereby effectively realizing the user's seamless hot switching experience.
[0064] refer to Figure 4 This is an interaction timing diagram of the context hot-switching method for a large language model based on entity tokens in this application embodiment.
[0065] The user places a physical doll (i.e., a physical token) on the smart terminal and engages in voice conversation with the doll (smart terminal). The smart terminal reads the doll's UID and uploads the UID and the user's voice audio to the cloud gateway. The cloud gateway sends a request to parse the UID signal to the entity management service module. The entity management service module queries the corresponding mapping relationship based on the UID, obtains the virtual role configuration information corresponding to the UID, and sends it to the dialogue flow orchestration engine. After initialization, the dialogue flow orchestration engine sends a context construction request to the dynamic context builder. The dynamic context builder injects virtual role settings and physical state information to construct the dialogue context. The constructed dialogue context is then sent to the large language model for response generation. After generating the response text, the large language model returns the response text to the dynamic context builder. At this point, the dynamic context builder further sends the response text back to the smart terminal, which or the physical doll then broadcasts it via voice to respond to the user's conversation.
[0066] According to the embodiment of this application, the context hot-switching method of a large language model based on entity tokens first receives a unique identifier of an entity token sent from a smart terminal. Further, it queries a mapping database based on the unique identifier to obtain virtual role configuration information bound to the unique identifier. The virtual role configuration information can be updated by modifying cloud configuration data. Then, it receives user input information uploaded by the smart terminal. Based on the virtual role configuration information and the user input information, it dynamically constructs a dialogue context containing the target role setting. The dialogue context is sent to a large language model for processing, and a text response generated by the large language model is received. Based on the virtual role configuration information, the text response is converted into a corresponding voice signal and sent to the smart terminal. Furthermore, when the identifier changes, the current construction and processing business is automatically interrupted, and the virtual role configuration information corresponding to the new identifier is loaded based on the new identifier, overwriting the current dialogue context. This application receives and identifies the identifier of an entity token, obtains the virtual character configuration information corresponding to the identifier, combines it with the information received from the user input to the terminal, constructs the dialogue environment of the virtual character, and sends the current dialogue context to a language big model to generate corresponding voice signal responses. Through the voice representation of the virtual character, the user can perform corresponding intelligent human-computer interaction at the physical entity level, and can also quickly and freely switch virtual characters, effectively providing users with rich interactive scenarios and emotional companionship.
[0067] refer to Figure 5 This is a schematic diagram of a cloud server provided in an embodiment of this application.
[0068] Based on the same concept, corresponding to the entity token-based large language model context hot switching method provided in any of the above embodiments, this application also provides a cloud server 120, including a first receiving module 1201, a query module 1202, a second receiving module 1203, a construction module 1204, a processing module 1205, a feedback module 1206, and a response module 1207.
[0069] The first receiving module 1201 is configured to receive the unique identifier of the entity token sent from the smart terminal; the query module 1202 is configured to query the mapping relationship database based on the unique identifier to obtain the virtual role configuration information bound to the unique identifier; wherein, the virtual role configuration information can be updated by modifying cloud configuration data; the second receiving module 1203 is configured to receive user input information uploaded by the smart terminal; the construction module 1204 is configured to dynamically construct a dialogue context containing the target role setting based on the virtual role configuration information and the user input information; the processing module 1205 is configured to send the dialogue context to the large language model for processing and receive the text reply generated by the large language model; the feedback module 1206 is configured to convert the text reply into the corresponding voice signal based on the virtual role configuration information and send it to the smart terminal; the response module 1207 is configured to automatically interrupt the current dialogue context construction and large language model processing task based on the virtual role configuration information when the unique identifier of the entity token is changed to a new unique identifier, and load the corresponding new virtual role configuration information in real time based on the new unique identifier to overwrite the dialogue context.
[0070] In some embodiments, the virtual character configuration information includes system prompt words parameters, and the construction module 1204 is further configured to: based on the system prompt words parameters and combined with user input information, generate prompt information in real time that includes the target character settings and the current dialogue content to form a dialogue context.
[0071] In some embodiments, a dialogue context containing the target role setting is dynamically constructed. The construction module 1204 is further configured to: receive sensor signals uploaded by the smart terminal that reflect physical interaction with the entity token; convert the sensor signals into corresponding context semantic representations; and integrate the context semantic representations as supplementary context information into the dialogue context.
[0072] In some embodiments, the virtual character configuration information also includes emotion response parameters, and the construction module 1204 is further configured to: dynamically generate emotion indication information that matches the current dialogue content in the dialogue context based on the semantic analysis results of the user input information and the emotion response parameters, so as to guide the large language model to generate text responses with corresponding emotional coloring.
[0073] In some embodiments, the virtual character configuration information also includes speech synthesis parameters, and the feedback module 1206 is further configured to synthesize the text reply into a speech signal with specific timbre characteristics set by the target character based on the speech synthesis parameters.
[0074] In some embodiments, the response module 1207 is further configured to: upon receiving the unique identifier of the new entity token, obtain the corresponding new virtual role configuration information based on the unique identifier of the new entity token; wherein, obtaining the corresponding new virtual role configuration information is automatically triggered in response to the change of the unique identifier.
[0075] According to the cloud server in this application embodiment, it first receives a unique identifier of an entity token sent from a smart terminal. Further, it queries a mapping database based on the unique identifier to obtain virtual role configuration information bound to the unique identifier. The virtual role configuration information can be updated by modifying cloud configuration data. Then, it receives user input information uploaded by the smart terminal. Based on the virtual role configuration information and the user input information, it dynamically constructs a dialogue context containing the target role settings. The dialogue context is sent to a large language model for processing, and a text response generated by the large language model is received. Based on the virtual role configuration information, the text response is converted into a corresponding voice signal and sent to the smart terminal. Furthermore, when the identifier changes, the current construction and processing business is automatically interrupted, and the virtual role configuration information corresponding to the new identifier is loaded based on the new identifier, overwriting the current dialogue context. In this embodiment, the cloud server receives and identifies the identifier of the entity token, obtains the virtual character configuration information corresponding to the identifier, combines the information received from the user input to the terminal, constructs the dialogue environment of the virtual character, and sends the current dialogue context to the language big model to generate the corresponding voice signal response. Through the voice representation of the virtual character, the user can perform corresponding intelligent human-computer interaction at the physical entity level, and can also quickly and freely switch virtual characters, effectively providing users with rich interactive scenarios and emotional companionship.
[0076] refer to Figure 6 This is a schematic diagram of a smart terminal provided in an embodiment of this application.
[0077] Based on the same concept, corresponding to the entity token-based large language model context hot-switching method provided in any of the above embodiments, this application also provides a smart terminal 112, including a near-field communication module 1121, a data acquisition module 1122, a sensing module 1123, a far-field communication module 1124, and an audio output module 1125.
[0078] The near-field communication module 1121 is configured to read a unique identifier from a physical token. In this embodiment, the near-field communication module is preferably an NFC / RFID reader, which can quickly read the identifier information in an NFC tag or RFID chip.
[0079] The acquisition module 1122 is configured to acquire user input information. In this embodiment, the acquisition module is preferably a microphone array for acquiring the user's voice input. In this embodiment, the user's input information can also be text information. The acquisition module can also be a touch screen, in which case the user can input text information through the touch screen to achieve interaction with the virtual character.
[0080] The sensing module 1123 is configured to collect sensor signals reflecting physical interaction with the physical token. In this embodiment, the sensing module may preferably be a tactile sensor, posture sensor, heartbeat or temperature simulation sensor, etc., and can transmit the collected physical interaction sensor signals through communication.
[0081] The far-field communication module 1124 is configured to send a unique identifier, user input information, and sensor signals to a cloud server; that is, the far-field communication module is a network module.
[0082] The audio output module 1125 is configured to receive and play voice signals from a cloud server; wherein the voice signals are generated by the cloud server after processing user input information based on virtual character configuration information bound to a unique identifier.
[0083] According to the embodiments of this application, the smart terminal reads the identifier of the physical token and collects user input information, and transmits the identifier and user input information to the cloud server, thereby realizing the real-time conversion of physical interaction actions to requests from the cloud server. This provides the necessary data foundation for the cloud server to load the corresponding virtual character configuration based on the identifier and construct the dialogue context by combining the user input information.
[0084] refer to Figure 7 This is a schematic diagram of a large language model context hot-switching system based on entity tokens provided in an embodiment of this application.
[0085] Based on the same concept, corresponding to the entity token-based large language model context hot-switching method provided in any of the above embodiments, this application also provides an entity token-based large language model context hot-switching system 700, including the cloud server 120 and the smart terminal 112. The entity token-based large language model context hot-switching system 700 is used to execute the entity token-based large language model context hot-switching method described above, and has the beneficial effects of the corresponding entity token-based large language model context hot-switching method embodiments, which will not be repeated here.
[0086] Based on the same concept, corresponding to the entity token-based large language model context hot-switching method provided in any of the above embodiments, this application also provides a computer-readable storage medium storing a program or instructions, which, when executed by a processor, implements the entity token-based large language model context hot-switching method.
[0087] The aforementioned computer-readable storage medium can be any available medium or data storage device that a computer can access, including but not limited to magnetic storage (e.g., floppy disks, hard disks, magnetic tapes, magneto-optical disks (MOs), etc.), optical storage (e.g., CDs, DVDs, BDs, HVDs, etc.), and semiconductor storage (e.g., ROMs, EPROMs, EEPROMs, non-volatile memory (NAND flash), solid-state drives (SSDs)).
[0088] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the corresponding entity token-based large language model context hot-switching method in any of the foregoing embodiments, and have the beneficial effects of the corresponding entity token-based large language model context hot-switching method embodiments, which will not be repeated here.
[0089] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0090] From the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of computer software products plus necessary general-purpose hardware platforms, and of course, they can also be implemented by hardware. The computer software product is stored in a storage medium (such as ROM, RAM, magnetic disk, optical disk, etc.) and includes several instructions to cause the terminal or network-side device to execute the methods described in the various embodiments of this application.
[0091] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other implementations under the guidance of this application without departing from the spirit and scope of the claims. All of these implementations are within the protection scope of this application.
Claims
1. An entity token-based large language model context hot switching method, characterized in that, The method includes: A unique identifier for receiving entity tokens sent from smart terminals; The mapping database is queried based on the unique identifier to obtain the virtual character configuration information bound to the unique identifier; wherein, the virtual character configuration information can be updated by modifying cloud configuration data; Receive user input information uploaded by the smart terminal; Based on the virtual character configuration information and the user input information, a dialogue context containing the target character settings is dynamically constructed; The dialogue context is sent to a large language model for processing, and the text response generated by the large language model is received. Based on the virtual character configuration information, the text reply is converted into a corresponding voice signal and sent to the smart terminal; When the unique identifier of the entity token is changed to a new unique identifier, the current dialogue context construction and large language model processing tasks based on the virtual character configuration information are automatically interrupted, and the corresponding new virtual character configuration information is loaded in real time based on the new unique identifier to overwrite the dialogue context.
2. The method of claim 1, wherein, The virtual character configuration information includes system prompt word parameters; The step of dynamically constructing a dialogue context containing the target character's settings based on the virtual character configuration information and the user input information includes: Based on the system prompt words parameters and combined with the user input information, prompt information containing the target role setting and the current dialogue content is generated in real time to form the dialogue context.
3. The method for hot-switching context of a large language model based on entity tokens according to claim 1, characterized in that, The dynamic construction includes the dialogue context set by the target role, and also includes: Receive sensor signals uploaded by the smart terminal that reflect physical interaction with the physical token; Convert the sensor signals into corresponding contextual semantic representations; The contextual semantic representation is integrated into the dialogue context as supplementary contextual information.
4. The method for hot-switching context of a large language model based on entity tokens according to claim 2, characterized in that, The virtual character configuration information also includes emotion response parameters; The method further includes: Based on the semantic analysis results of the user input information and the emotion response parameters, emotion indication information matching the current dialogue content is dynamically generated in the dialogue context to guide the large language model to generate text responses with corresponding emotional coloring.
5. The method for hot-switching context of a large language model based on entity tokens according to claim 2, characterized in that, The virtual character configuration information also includes voice synthesis parameters; The step of converting the text reply into a corresponding voice signal includes: Based on the speech synthesis parameters, the text response is synthesized into a speech signal with the specific timbre characteristics set for the target character.
6. The method for hot-switching context of a large language model based on entity tokens according to claim 1, characterized in that, The method further includes: Upon receiving the unique identifier of the new entity token, the corresponding new virtual role configuration information is obtained based on the unique identifier of the new entity token; wherein, the obtaining of the corresponding new virtual role configuration information is automatically triggered in response to the change of the unique identifier.
7. A cloud server, characterized in that, include: The first receiving module is configured to receive a unique identifier for an entity token sent from a smart terminal; The query module is configured to query the mapping relationship database based on the unique identifier to obtain the virtual character configuration information bound to the unique identifier; wherein, the virtual character configuration information can be updated by modifying cloud configuration data; The second receiving module is configured to receive user input information uploaded by the smart terminal; The construction module is configured to dynamically construct a dialogue context containing the target character's settings based on the virtual character configuration information and the user input information; The processing module is configured to send the dialogue context to a large language model for processing and to receive the text response generated by the large language model. The feedback module is configured to convert the text reply into a corresponding voice signal and send it to the smart terminal based on the virtual character configuration information; The response module is configured to automatically interrupt the current dialogue context construction and large language model processing tasks based on the virtual character configuration information when the unique identifier of the entity token is changed to a new unique identifier, and to load the corresponding new virtual character configuration information in real time based on the new unique identifier to overwrite the dialogue context.
8. A smart terminal, characterized in that, include: The near-field communication module is configured to read a unique identifier from an entity token; The data acquisition module is configured to collect user input information; The sensing module is configured to acquire sensor signals reflecting physical interactions with the physical token; The far-field communication module is configured to send the unique identifier, the user input information, and the sensor signals to a cloud server. An audio output module is configured to receive and play voice signals from the cloud server; wherein the voice signals are generated by the cloud server after processing the user input information based on the virtual character configuration information bound to the unique identifier.
9. A context-switching system for a large language model based on entity tokens, characterized in that, It includes the cloud server as described in claim 7 and the smart terminal as described in claim 8.
10. A computer-readable storage medium, characterized in that, The program or instructions are stored on the readable storage medium, and when the program or instructions are executed by a processor, they implement the steps of the entity token-based large language model context hot-switching method as described in any one of claims 1 to 6.