AI role interaction method and system of intelligent lollipop

By building a role identification module and a personalized interaction model in the smart lollipop, the interaction between emotional perception and character evolution is achieved, the problem of insufficient interaction capabilities of the existing smart lollipop is solved, and the multi-character companionship and emotional resonance of the user experience is improved.

CN120580997AActive Publication Date: 2025-09-02AMES (GUANGDONG) FOOD TECH CO LTD +1
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510673920.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-09-02
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

Existing smart lollipops lack dynamic semantic understanding and personalized interaction capabilities, and cannot meet users' deep needs for multi-character companionship and emotional communication.

Method used

By building a character identification module in the sugar head of the smart lollipop, combining the smart handle for voice data acquisition and emotional intention analysis, a pre-trained personalized exclusive interaction model is used to generate response voices that conform to the AI ​​character style, realizing the interaction between emotional perception and character evolution.

Benefits of technology

It improves the context adaptability and emotional resonance of interactive content, enhances the multi-character companionship and immersion of user experience, and fills the technical gap between multimodal interaction and emotional companionship.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120580997A_ABST
    Figure CN120580997A_ABST
Patent Text Reader

Abstract

The invention discloses an AI role interaction method and system of an intelligent lollipop, the intelligent lollipop comprises an intelligent handle and a plurality of lollipop heads with different physical shapes, the intelligent handle is in communication connection with a user side, the lollipop heads are detachably installed at the front end of the intelligent handle, and a role identification module is arranged in the lollipop head with each physical shape; the method comprises the steps that after a current candy head is inserted into an intelligent handle, a role identification module in the current candy head is read to load a current AI role; voice data of the user are collected, and emotion categories and intention information are analyzed; and inputting an analysis result into a personalized exclusive interaction model corresponding to the current AI role, and outputting response voice conforming to the style of the current AI role. According to the intelligent lollipop, a hardware form technology, an emotion analysis technology and a dynamic response technology are coupled, an existing intelligent lollipop is upgraded from a one-way playing tool to an AI interaction terminal with emotion perception, role evolution and personalized dialogue capabilities, and multi-modal interaction and emotional accompanying are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an AI character interaction method and system for a smart lollipop. Background Art

[0002] In recent years, smart lollipops, as interactive electronic consumer products, have gradually evolved from simple candy forms to integrated functions. Existing products on the market typically focus on basic voice playback and recording capabilities, such as pre-stored songs and story audio via built-in memory, or allowing users to record and playback short voice messages. While these products offer a certain level of entertainment, their interaction model remains limited to one-way output or mechanical repetition, lacking dynamic semantic understanding and personalized interaction capabilities. They fail to meet users' deeper needs for multi-role companionship and emotional communication. Summary of the Invention

[0003] The purpose of the present invention is to provide an AI character interaction method and system for a smart lollipop, aiming to solve the problem that existing smart lollipops lack dynamic semantic understanding and personalized interaction capabilities.

[0004] In a first aspect, an embodiment of the present invention provides an AI character interaction method for a smart lollipop, wherein the smart lollipop includes a smart handle and multiple lollipop heads with different physical shapes. The smart handle is communicatively connected to a user terminal, and each lollipop head is detachably mounted on the front end of the smart handle. Each lollipop head with different physical shapes has a built-in character identification module. The AI ​​character interaction method includes:

[0005] In response to the operation of plugging the current sugar head into the front end of the smart handle, the character identification module in the current sugar head is read by the smart handle to load the current AI character;

[0006] The smart lollipop collects voice data sent by the user and performs emotion analysis and intention analysis to obtain corresponding emotion categories and intention information;

[0007] Input the voice data, emotion category, and intent information into a pre-trained personalized interaction model corresponding to the current AI character for question-answering processing, and generate a response voice that matches the style of the current AI character;

[0008] The response voice is fed back to the user via the smart lollipop.

[0009] In a second aspect, an embodiment of the present invention provides an AI character interaction system for a smart lollipop, wherein the smart lollipop includes a smart handle and multiple lollipop heads with different physical shapes. The smart handle is communicatively connected to a user terminal, and each lollipop head is detachably mounted on the front end of the smart handle. Each lollipop head with different physical shapes has a built-in character identification module. The AI ​​character interaction system includes:

[0010] An identification unit, configured to respond to an operation of plugging a current sugar head into the front end of the smart handle and read the character identification module in the current sugar head through the smart handle to load the current AI character;

[0011] An analysis unit, configured to collect voice data emitted by the user through the smart lollipop and perform emotion analysis and intention analysis to obtain corresponding emotion categories and intention information;

[0012] A questioning unit, configured to input the voice data, emotion category, and intention information into a pre-trained personalized interaction model corresponding to the current AI character for question-answering processing, and generate a response voice that conforms to the style of the current AI character;

[0013] The answering unit is used to feed back the answering voice to the user through the smart lollipop.

[0014] The beneficial effects of the embodiments of the present invention are:

[0015] By coupling the three technologies of hardware form, emotion analysis and dynamic response, the existing smart lollipop is upgraded from a one-way playback tool to an AI interactive terminal with the ability of emotion perception, role evolution and personalized dialogue, filling the technical gap in multimodal interaction and emotional companionship in existing technologies. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0017] Figure 1 A flowchart of the AI ​​character interaction method provided in an embodiment of the present invention.

[0018] Figure 2 This is a schematic diagram of a sub-flow of step S101 provided in an embodiment of the present invention.

[0019] Figure 3 This is a schematic diagram of a sub-flow of step S102 provided in an embodiment of the present invention.

[0020] Figure 4This is a schematic diagram of a sub-flow of step S305 provided in an embodiment of the present invention.

[0021] Figure 5 This is a schematic diagram of a sub-flow chart of step S306 provided in an embodiment of the present invention.

[0022] Figure 6 This is a schematic diagram of a sub-flow of step S103 provided in an embodiment of the present invention.

[0023] Figure 7 Another flowchart of the AI ​​character interaction method provided by an embodiment of the present invention.

[0024] Figure 8 Another flowchart of the AI ​​character interaction method provided by an embodiment of the present invention.

[0025] Figure 9 A schematic block diagram of an AI character interaction system provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0026] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort shall fall within the scope of protection of the present invention.

[0027] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0028] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0029] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0030] See also Figure 1 , Figure 1 A flowchart of an AI character interaction method for a smart lollipop provided by an embodiment of the present invention.

[0031] To facilitate understanding of the present application, the smart lollipop of the present invention is first introduced.

[0032] An embodiment of the present invention provides a smart lollipop, comprising a smart handle and multiple lollipop heads with different physical shapes. The smart handle is communicatively connected to a user terminal, and each lollipop head is detachably mounted on the front end of the smart handle. Each lollipop head with a physical shape has a built-in character identification module. The smart handle is used to trigger the reading and loading of the character identification module in the lollipop head when the lollipop head is mounted on the front end of the smart handle.

[0033] In this embodiment, the smart handle can provide users with an interactive experience related to the character by intelligently identifying the character identification module in the sugar head.

[0034] Furthermore, an NFC chip is built into the sugar head, and the NFC chip is used to store the role identification information corresponding to the current sugar head to form a role identification module.

[0035] Furthermore, the sugar head and the smart handle are connected to each other through a USB connector, NFC, Bluetooth, etc.; preferably, the sugar head and the smart handle are connected to each other through a USB connector.

[0036] Furthermore, a bone conduction vibration component is provided in the candy head, which transmits the user's voice data to the smart handle through the bone conduction vibration component. The smart handle has a built-in communication module for communicating with the user end to send the voice data to the user end for processing.

[0037] Furthermore, the smart handle is provided with a sound receiving module for receiving voice data sent by the user, and a speaker module for answering the user.

[0038] Based on the above-mentioned smart lollipop solution, the following specifically introduces the AI ​​character interaction method of the smart lollipop.

[0039] like Figure 1 As shown, the AI ​​character interaction method provided by the embodiment of the present invention includes steps S101-S104.

[0040] S101: In response to the operation of plugging the current sugar head into the front end of the smart handle, the smart handle reads the character identification module in the current sugar head to load the current AI character;

[0041] In this step, the built-in character identification module of the detachable candy head is linked to the intelligent handle, so that each physical candy head corresponds to a unique AI character identity. When the user changes to a candy head made of different materials, the loading of a different AI character is automatically triggered without manually switching programs. This effectively solves the problem of disconnection between physical modeling and digital content, directly mapping physical form changes to AI personality feature switching, creating an immersive interactive scene where "model is character", significantly enhancing the user's perception of the realism of multi-character companionship.

[0042] In this step, the AI ​​character is loaded from a pre-created character portrait library. This library creates a three-dimensional portrait of each AI character, encompassing their basic attributes, language style, and knowledge domains. This 3D portrait includes: character metadata (character type, historical context, personality matrix [rigor, liveliness, curiosity]); style corpus (high-frequency vocabulary, sentence template library, rhetorical device distribution); and domain knowledge topology (main domain knowledge graph, cross-domain association rules, taboo topic list). Specific AI character types can be customized based on specific demographics.

[0043] S102, collecting voice data sent by the user through the smart lollipop and performing emotion analysis and intention analysis to obtain corresponding emotion categories and intention information;

[0044] S103: Input the voice data, emotion category, and intent information into a pre-trained personalized interaction model corresponding to the current AI character for question-answering processing, and generate a response voice that matches the style of the current AI character;

[0045] In steps S102-S103, real-time emotion category recognition and intent analysis based on user voice data, combined with dynamic response generation using a pre-trained personalized interaction model, enables the system to adjust feedback strategies based on the user's emotional state (e.g., excitement, depression) and interaction intent (e.g., asking a question, sharing a conversation). Compared to existing interaction methods that rely on fixed command matching or static audio playback, this significantly improves the contextual adaptability and emotional resonance of conversational content, resolving the core issue of a monotonous user experience caused by mechanical, repetitive responses.

[0046] S104, feeding back the answer voice to the user via the smart lollipop;

[0047] In this step, the smart controller is equipped with an external speaker module. It can convert the sound signal of the response voice into a mechanical vibration signal, and use bone conduction technology to feed the mechanical vibration signal back to the candy head. That is, the user can directly interact with the AI ​​character while holding the candy head; and / or, the response voice can be played through the external speaker module on the smart controller, so that the user can still hear and interact even if the candy head is taken out of the mouth.

[0048] In this embodiment, through the coupling of hardware form, emotion analysis and dynamic response triple technologies, the existing smart lollipop is upgraded from a one-way playback tool to an AI interactive terminal with emotion perception, role evolution and personalized dialogue capabilities, filling the technical gap in multimodal interaction and emotional companionship in the existing technology.

[0049] In some embodiments, the various models for analyzing voice data (i.e., personalized exclusive interaction models and models such as the dual-stream deep pulse neural network described below) can be deployed online in the user terminal (in the APP application), and the smart lollipop is connected to the user terminal using wired or wireless communication to realize two-way data transmission and reception. Obviously, a data processing module can also be built into the smart handle. The trained models can be deployed offline in the data processing module of the smart handle to realize offline application in a network-free environment.

[0050] In one embodiment, after the sugar head and the smart handle are plugged in and out, communication connection can be achieved through various methods, such as USB connector, NFC, Bluetooth, etc.

[0051] In one embodiment, the candy head and the smart handle are connected via a USB connector, and an NFC chip is built into the candy head. The NFC chip is used to store the character identification information corresponding to the current candy head to form a character identification module. Figure 2 As shown, in one embodiment, step S101 includes steps S201-S203.

[0052] S201: In response to the operation of plugging the current candy head into the front end of the smart handle, the character identification information in the current candy head is read through the smart handle;

[0053] S202, obtaining a unique identification code in the role identification information;

[0054] S203: Match the corresponding AI character on the user side according to the unique identification code, load the character, and then access the personalized exclusive interaction model corresponding to the current AI character.

[0055] In this embodiment, based on the current material and shape of the candy head selected by the user, after the current candy head is plugged into the front end of the smart controller, the smart controller can be triggered to read the NFC chip in the current candy head to obtain the character identification information, and then automatically load the AI ​​character corresponding to the current candy head according to the unique identification code in the character identification information. Throughout this process, through the intelligent recognition and matching process, the user and the AI ​​character are quickly linked, which not only improves the user's operational convenience, but also enhances the fun and immersiveness of entertainment interaction, bringing users a richer and more diverse user experience.

[0056] In one embodiment, a bone conduction vibration component is provided in the candy head, and a communication module for communicating with the user end is built into the smart handle. Figure 3 As shown, step S102 includes steps S301-S306.

[0057] S301. Collecting voice data from the user using bone conduction technology through the candy head, and / or receiving the voice data from the user through a sound receiving module provided on the smart controller;

[0058] S302, extracting speech waveform features of the speech data and inputting them into the main branch of the dual-stream deep pulse neural network, extracting time-domain emotional features in the speech data through a multi-layer pulse convolution kernel, and generating a time-domain pulse sequence;

[0059] S303, extracting the MFCC coefficients of the speech data and converting them into sparse pulse representations through a frequency domain pulse encoder, and then inputting them into the side branch of the dual-stream deep pulse neural network, using a formant-sensitive pulse pooling layer to capture the frequency domain features related to the intent, and generating a frequency domain pulse stream;

[0060] S304, using a pulse cross-correlation algorithm to calculate the mutual information entropy value of the time-domain pulse sequence and the frequency-domain pulse stream, generating a spatiotemporal coherence matrix, and extracting cross-modal joint features from the spatiotemporal coherence matrix through pulse emission synchronization detection;

[0061] S305: Using a pulse sequence dynamic time warping algorithm, the emotion classifier performs dynamic feature comparison between the cross-modal joint feature and a preset emotion feature template library, and outputs a corresponding emotion category.

[0062] In step S305, emotions are not only reflected by the time domain features of speech, but the frequency domain features also contain important information. Therefore, emotion analysis based on cross-modal joint features is more accurate.

[0063] S306, performing pulse sequence symbolization inference on the frequency domain pulse stream through the intent analyzer to locate the corresponding intent information in the preset intent map library;

[0064] In step S306, the intent analysis focuses more on the vocabulary and grammatical structure in the speech. The frequency domain features can be converted into symbol sequences after processing, which is more suitable for state machine reasoning.

[0065] The dual-stream deep spiking neural network (DSNN) of this embodiment is a computational model that combines a spiking neural network (SNN) and a dual-stream architecture. The dual-stream deep spiking neural network includes a main branch and a side branch. The specific construction method of the dual-stream deep spiking neural network is as follows: the main branch is composed of a stack of multiple layers of pulse convolution kernels, which is used to extract time-domain emotional features from the original speech waveform. After each layer of pulse convolution, a pulse activation function is embedded, and a time-domain pulse sequence is generated through a membrane potential accumulation-discharge mechanism; the side branch includes a frequency-domain pulse encoder and a formant-sensitive pulse pooling layer. The frequency-domain pulse encoder maps MFCC coefficients (Mel-Frequency Cepstral Coefficients) to a sparse pulse representation. The formant-sensitive pulse pooling layer selects frequency band features related to speech intent through a tunable pulse discharge threshold; and the spatiotemporal feature fusion module connects the main branch and the side branch. The spike synchronization detection algorithm is used to align the time-domain and frequency-domain pulse streams, and a cross-modal joint feature vector is generated through pulse cross-correlation calculation.

[0066] In this embodiment, based on the process of steps S301-S306, the user's voice data is jointly analyzed for emotion and intention through a dual-stream deep pulse neural network, which can accurately decode the user's emotional state and intention information, providing the user with a more intimate and personalized interactive service.

[0067] like Figure 4 As shown, in one embodiment, step S305 includes:

[0068] S401, pre-building a preset emotion feature template library containing multiple types of emotions;

[0069] S402: Slice the cross-modal joint features in time series to generate a plurality of continuous real-time pulse trajectory segments of a preset duration (e.g., 200-500 ms);

[0070] S403: Using a pulse sequence dynamic time warping algorithm, elastically align the multiple real-time pulse trajectory segments with the multiple standard pulse trajectory segments of each emotion type, and then perform similarity comparison to obtain the similarity between the cross-modal joint feature and the trajectory corresponding to each emotion type;

[0071] S404 , selecting multiple candidate emotions from each trajectory similarity in descending order, and performing multi-scale emotion probability estimation to output the candidate emotion with the highest probability as the target emotion.

[0072] In this embodiment, the emotional feature template library can include multiple emotions such as joy, sadness, anger, surprise, fear, and calmness, and can be specifically designed according to the target audience. Each emotion type stores a three-dimensional feature template: Dimension 1: Pulse density gradient histogram (which counts the rate of change of the number of pulses according to a preset time window), Dimension 2: Cross-layer pulse synchronization matrix (which records the frequency of coordinated triggering of the main branch in the time domain and the side branch in the frequency domain within a preset time window), and Dimension 3: Membrane potential oscillation spectrum (which extracts the energy distribution of the preset characteristic frequency band).

[0073] In this embodiment, the Pulse Sequence Dynamic Time Warping Algorithm (PS-DTW) is used to solve the alignment problem between time series of different lengths. That is, the pulse trajectory segments to be compared need to be "elastically aligned" to solve problems such as misalignment of pulse segments caused by different speaking rates. Only after "elastic alignment" can their similarities be compared more objectively and accurately.

[0074] In this embodiment, multi-scale emotion probability estimation is performed on multiple candidate emotions, and the characteristics of different scales (such as pulse density change rate, oscillation frequency, etc.) need to be comprehensively considered to improve the accuracy of emotion recognition. Finally, the probability results of each scale are integrated to determine the target emotion.

[0075] like Figure 5 As shown, in one embodiment, step S306 includes:

[0076] S501, using a pulse timing logic rule engine to discretize the emission time of the frequency domain pulse stream and map the pulse emission cluster into a symbol sequence, where each symbol corresponds to a semantic primitive in a preset intent map library;

[0077] S502: In a hierarchical topology of a preset intent map, perform spatiotemporal similarity matching between the symbol sequence and the intent nodes in the preset intent map library, wherein the preset intent map includes an intent-action association layer constructed by a knowledge graph and a dynamic context parameter layer driven by user historical behavior;

[0078] In this step, the intent-action association layer represents static logical relationships based on the knowledge graph (e.g., "User says 'play music' → triggers the music app"). The dynamic context parameter layer extracts historical user behavior features to generate contextual weight coefficients and modify the spatiotemporal similarity score. Spatiotemporal similarity matching involves matching the similarities between symbol sequences and intent nodes in both the temporal dimension (sequential similarity between the symbol sequence and the intent node) and the spatial dimension (feature similarity between the symbol sequence and the intent node), and then performing a weighted fusion to obtain spatiotemporal similarity.

[0079] S503: When the spatiotemporal similarity exceeds a first threshold, outputting the intention information of successful positioning and the associated confidence level;

[0080] S504: When the spatiotemporal similarity does not exceed the first threshold, triggering the feedback fine-tuning mechanism of the dual-stream deep spike neural network to readjust the frequency band weight coefficient of the resonance peak sensitive spike pooling layer.

[0081] In this embodiment, based on steps S501-S505, pulse sequence symbolization and a hierarchical intent matching mechanism are used to transform complex frequency-domain pulse streams into interpretable semantic primitives. This is combined with knowledge graphs and dynamic user behavior data to achieve high-precision intent recognition. Furthermore, a feedback fine-tuning mechanism is introduced to automatically optimize the neural network band weights when matching fails, enhancing the system's adaptability to non-stationary signals.

[0082] like Figure 6 As shown, in one embodiment, step S103 includes:

[0083] S601: Input the voice data, emotion category, and intent information into the pre-trained personalized interaction model corresponding to the current AI character;

[0084] S602: Extracting a style parameter set of the current AI character, where the style parameter set includes a corpus preference coefficient, a tone feature vector, and a domain knowledge weight;

[0085] S603, mapping emotion categories to emotion reinforcement factors;

[0086] S604: Convert the intent information into a domain knowledge index;

[0087] S605: A style adaptation layer in a personalized interaction model is used to combine and calculate the corpus preference coefficient and the sentiment enhancement factor to generate dynamic style constraints.

[0088] S606. Retrieve matching domain corpus from the pre-built multi-dimensional knowledge graph based on domain knowledge weight and domain knowledge index;

[0089] S607: Input the voice data, dynamic style constraints and domain corpus into the dialogue generation network, and generate a response voice that conforms to the language characteristics of the current AI character through adversarial style transfer.

[0090] In this embodiment, based on the process of steps S601-S607, deep customization of the AI ​​character's responses is achieved. The core of this process lies in the use of style parameters to dynamically integrate user emotions and corpus preferences, combined with adversarial transfer generation technology. This ensures that the response voice not only maintains the unique language characteristics of the current AI character (such as specific intonation and domain expression habits), but also adapts to the emotional intensity and professional knowledge requirements of the conversation scenario in real time, thereby significantly improving the realism, emotional resonance, and response accuracy of the smart lollipop when interacting with the user.

[0091] like Figure 7 As shown, in one embodiment, the AI ​​character interaction method of the smart lollipop further includes:

[0092] S701, generating an emotional memory node with a time dimension mark based on the voice data, emotion category, intention information and corresponding response voice, and establishing a dynamic emotional map consisting of multiple emotional memory nodes;

[0093] S702, extracting text information from the voice data and performing semantic analysis to obtain at least three levels of interest entities including basic entities, composite entities, and abstract concept entities;

[0094] In this step, basic entities refer to atomic semantic units directly extracted through syntactic parsing, corresponding to observable objects in the physical world; composite entities refer to complex semantic units composed of basic entities through dependency relationships, representing specific events or behavior patterns; abstract concept entities refer to surreal mapping relationships generated through semantic reasoning, representing users' potential values ​​or cognitive tendencies;

[0095] S703, cross-modally associating the interest entity with the corresponding emotional memory node, establishing a multi-dimensional association matrix of entity-emotion-time, and establishing a user interest knowledge graph consisting of multiple multi-dimensional association matrices;

[0096] S704: When it is detected that the emotion intensity change rate of a specific interest entity within a continuous time window exceeds a threshold, the corresponding emotion memory node is activated;

[0097] S705. Adjust the response parameters of the current AI character based on the activated emotional memory node. The response parameters include voice intonation parameters and response content generation strategy.

[0098] In this embodiment, through the process of steps S701-S705, a dynamic emotion graph and a user interest knowledge graph are constructed to achieve a deep understanding and memory tracking of user emotions and interests by the AI ​​character. This core feature is the use of time dimension markers and a multidimensional correlation matrix to capture the changes in user emotion intensity and the evolution of interest entities. Cross-modal data fusion is then used to adjust the AI ​​response strategy in real time, ensuring emotional coherence and interest relevance in interactive responses, significantly improving the naturalness and adaptability of personalized interactions.

[0099] In this embodiment, in order to better construct the user interest knowledge graph, an open-ended questioning strategy can be used to detect the user's interests when the user uses it for the first time, or the information entered by the user can be collected through the user terminal to build a preliminary user interest knowledge graph based on this.

[0100] For example, suppose the emotional memory node marked with the time dimension is the user's regular outdoor activity once a week. After it is activated, the AI ​​character can adjust its response parameters to respond with more positive and encouraging voice and tone parameters. At the same time, the response content generation strategy may revolve around the theme of outdoor activities, providing relevant suggestions, activity arrangements, or words of encouragement. Such adjustments not only reflect the AI ​​character's sensitive capture of user emotional changes, but also demonstrate its continued attention to and personalized response to user interests. In this way, the AI ​​character can more naturally integrate into the user's daily life, establish a deep emotional connection and interest resonance, thereby improving user experience and satisfaction.

[0101] like Figure 8 As shown, in one embodiment, the AI ​​character interaction method of the smart lollipop further includes:

[0102] S801, extracting text information of the voice data and performing word segmentation processing to obtain word groups;

[0103] S802: Detect dangerous keywords on the word group. If the word group triggers the crisis warning mechanism, evaluate the danger level of the word group according to the dangerous keyword level table.

[0104] S803. Adjust the current AI interaction strategy according to the triggered dangerous keywords and the corresponding danger level, including pre-defined different interaction strategies such as comforting, guiding or emergency assistance.

[0105] In this embodiment, based on the process of steps S801-S803, dynamic safety regulation of AI interaction strategy is realized through voice text analysis and dangerous keyword classification detection mechanism. It can identify potential dangerous contexts (such as violence, self-harm tendencies, etc.) in real time when the user uses the smart lollipop, and intelligently switch to soothing guidance or start emergency assistance procedures according to the risk level, effectively improving the level of emotional emergency response.

[0106] In this embodiment, when the dangerous keyword and the corresponding danger level reach the level requiring intervention, the emergency contact function is activated through the hardware interface of the smart lollipop, and the preset emergency contact number is automatically dialed or an emergency help message is sent; in addition, the entire process of the crisis event is recorded, including user voice, emotional response, interactive content, and response measures taken, to provide data support for subsequent crisis analysis and processing; finally, based on the crisis handling results and user feedback, the interaction logic of the AI ​​character and the crisis warning mechanism are continuously optimized and improved to enhance user experience and safety.

[0107] An embodiment of the present invention also provides an AI character interaction system for a smart lollipop. The smart lollipop includes a smart handle and multiple lollipop heads with different physical shapes. The smart handle is communicated with the user end. Each lollipop head can be detachably installed on the front end of the smart handle. Each lollipop head with different physical shapes has a built-in character identification module.

[0108] The AI ​​character interaction system of the embodiment of the present invention is used to execute any embodiment of the aforementioned AI character interaction method. Figure 9 , Figure 9 This is a schematic block diagram of an AI character interaction system provided by an embodiment of the present invention.

[0109] like Figure 9 As shown, the AI ​​character interaction system 900 includes: an identification unit 901, an analysis unit 902, a questioning unit 903 and a response unit 904.

[0110] The identification unit 901 is configured to read the character identification module in the current candy head through the smart handle to load the current AI character in response to the operation of plugging the current candy head into the front end of the smart handle;

[0111] An analysis unit 902 is configured to collect voice data sent by the user through the smart lollipop and perform emotion analysis and intention analysis to obtain corresponding emotion categories and intention information;

[0112] The questioning unit 903 is used to input the voice data, emotion category and intention information into the pre-trained personalized exclusive interaction model corresponding to the current AI character for question and answer processing, and generate a response voice that conforms to the style of the current AI character;

[0113] The response unit 904 is configured to feed back the response voice to the user via the smart lollipop.

[0114] By coupling hardware form, emotion analysis, and dynamic response technologies, this system upgrades the existing smart lollipop from a one-way playback tool to an AI interactive terminal with emotion perception, role evolution, and personalized dialogue capabilities, filling the technical gap in multimodal interaction and emotional companionship in existing technologies.

[0115] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working process of the system described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0116] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.

Claims

1. A method for AI character interaction of smart lollipop, characterized in that: The smart lollipop includes a smart handle and multiple lollipop heads with different physical shapes. The smart handle is in communication with a user terminal. Each lollipop head is detachably mounted on the front end of the smart handle. Each lollipop head with different physical shapes has a built-in character identification module. The AI ​​character interaction method includes: In response to the operation of plugging the current sugar head into the front end of the smart handle, the character identification module in the current sugar head is read by the smart handle to load the current AI character; The smart lollipop collects voice data sent by the user and performs emotion analysis and intention analysis to obtain corresponding emotion categories and intention information; Input the voice data, emotion category, and intent information into a pre-trained personalized interaction model corresponding to the current AI character for question-answering processing, and generate a response voice that matches the style of the current AI character; The response voice is fed back to the user via the smart lollipop.

2. The AI ​​character interaction method of the smart lollipop according to claim 1, characterized in that: The sugar head is connected to the smart handle via a USB connector. An NFC chip is built into the sugar head. The NFC chip is used to store the character identification information corresponding to the current sugar head to form the character identification module. In response to the operation of plugging the current sugar head into the front end of the smart handle, reading the character identification module in the current sugar head through the smart handle to load the current AI character, including: In response to the operation of plugging the current sugar head into the front end of the smart handle, reading the character identification information in the current sugar head through the smart handle; Obtaining a unique identification code from the role identification information; According to the unique identification code, the corresponding AI character is matched in the user terminal and the character is loaded and then accessed to the personalized exclusive interaction model corresponding to the current AI character.

3. The AI ​​character interaction method of the smart lollipop according to claim 1, characterized in that: The candy head is provided with a bone conduction vibration component, and the smart handle is built-in with a communication module for communicating with the user end; The smart lollipop collects voice data sent by the user and performs emotion analysis and intention analysis to obtain corresponding emotion categories and intention information, including: The user's voice data is collected by the candy head using bone conduction technology, and / or the user's voice data is received by the sound receiving module provided on the smart handle; Extracting speech waveform features of the speech data and inputting them into the main branch of a dual-stream deep pulse neural network, extracting time-domain emotional features in the speech data through a multi-layer pulse convolution kernel, and generating a time-domain pulse sequence; Extracting the MFCC coefficients of the speech data and converting them into sparse pulse representations through a frequency domain pulse encoder, and then inputting them into the side branch of a dual-stream deep pulse neural network. A formant-sensitive pulse pooling layer is used to capture the frequency domain features related to the intent, and a frequency domain pulse stream is generated. A pulse cross-correlation algorithm is used to calculate the mutual information entropy value of the time-domain pulse sequence and the frequency-domain pulse stream to generate a spatiotemporal coherence matrix, and a cross-modal joint feature is extracted from the spatiotemporal coherence matrix through pulse emission synchronization detection; The emotion classifier uses a pulse sequence dynamic time warping algorithm to perform dynamic feature comparison between the cross-modal joint feature and a preset emotion feature template library, and outputs a corresponding emotion category; The intention analyzer performs pulse sequence symbolization inference on the frequency domain pulse stream to locate the corresponding intention information in the preset intention map library.

4. The AI ​​character interaction method of the smart lollipop according to claim 3, characterized in that: The emotion classifier uses a pulse sequence dynamic time warping algorithm to perform dynamic feature comparison on the cross-modal joint feature and a preset emotion feature template library, and outputs the corresponding emotion category, including: A pre-built library of preset emotional feature templates containing multiple categories of emotions; Performing time series slicing on the cross-modal joint features to generate a plurality of continuous real-time pulse trajectory segments of a preset duration; Using a pulse sequence dynamic time warping algorithm, multiple real-time pulse trajectory segments are elastically aligned with multiple standard pulse trajectory segments for each type of emotion, and then similarity comparison is performed to obtain the trajectory similarity between the cross-modal joint feature and each type of emotion; Multiple candidate emotions are selected from each trajectory similarity in descending order, and after multi-scale emotion probability estimation, the candidate emotion with the highest probability is output as the target emotion.

5. The AI ​​character interaction method of the smart lollipop according to claim 3, characterized in that: The performing pulse sequence symbolization inference on the frequency domain pulse stream by the intent analyzer to locate corresponding intent information in a preset intent map library includes: A pulse timing logic rule engine is used to discretize the emission time of the frequency domain pulse stream and map the pulse emission cluster into a symbol sequence, wherein each symbol corresponds to a semantic primitive in the preset intention map library; In the hierarchical topology of the preset intent map, the symbol sequence is matched with the intent nodes in the preset intent map library for spatiotemporal similarity, wherein the preset intent map includes an intent-action association layer constructed by a knowledge graph and a dynamic context parameter layer driven by user historical behavior; When the spatiotemporal similarity exceeds a first threshold, outputting the intention information of successful positioning and the associated confidence; When the spatiotemporal similarity does not exceed the first threshold, the feedback fine-tuning mechanism of the dual-stream deep spike neural network is triggered to readjust the frequency band weight coefficient of the resonance peak sensitive spike pooling layer.

6. The AI ​​character interaction method of the smart lollipop according to claim 1, characterized in that: Input the voice data, emotion category, and intent information into the pre-trained personalized interaction model corresponding to the current AI character for question-answering processing, and generate a response voice that conforms to the style of the current AI character, including: Input the voice data, emotion category, and intent information into a pre-trained personalized interaction model corresponding to the current AI character; Extracting a style parameter set of the current AI character, wherein the style parameter set includes a corpus preference coefficient, a tone feature vector, and a domain knowledge weight; mapping the emotion categories into emotion reinforcement factors; Converting the intent information into a domain knowledge index; The corpus preference coefficient and the sentiment enhancement factor are integrated and calculated by the style adaptation layer in the personalized exclusive interaction model to generate dynamic style constraints; Based on the domain knowledge weight and domain knowledge index, searching for matching domain corpus from a pre-built multi-dimensional knowledge graph; The speech data, dynamic style constraints and domain corpus are input into the dialogue generation network, and a response speech that conforms to the language characteristics of the current AI character is generated through adversarial style transfer.

7. The AI ​​character interaction method of the smart lollipop according to claim 1, characterized in that: The smart handle is provided with an external speaker module; Feeding back the response voice to the user through the smart lollipop includes: Converting the sound signal of the response voice into a mechanical vibration signal, and feeding the mechanical vibration signal back to the canine head using bone conduction technology; and / or The answering voice is played through the external speaker module on the smart handle.

8. The AI ​​character interaction method of the smart lollipop according to claim 1, characterized in that: Also includes: Generate an emotional memory node with a time dimension mark based on the voice data, emotion category, intention information and corresponding response voice, and establish a dynamic emotional map consisting of multiple emotional memory nodes; Extracting text information from the speech data and performing semantic analysis to obtain interest entities of at least three levels, including basic entities, composite entities, and abstract concept entities; Cross-modally associate the interest entities with corresponding emotional memory nodes, establish a multi-dimensional association matrix of entity-emotion-time, and establish a user interest knowledge graph consisting of multiple multi-dimensional association matrices; When it is detected that the change rate of the emotion intensity of a specific entity of interest in a continuous time window exceeds a threshold, the corresponding emotion memory node is activated; Adjust the response parameters of the current AI character based on the activated emotional memory node, wherein the response parameters include voice intonation parameters and response content generation strategy.

9. The AI ​​character interaction method of the smart lollipop according to claim 1, characterized in that: Also includes: Extracting text information from the speech data and performing word segmentation processing to obtain word groups; Performing a dangerous keyword detection on the segmented word group, and if the segmented word group triggers a crisis warning mechanism, performing a danger level assessment on the segmented word group according to a dangerous keyword level table; Adjust the current AI interaction strategy based on the triggered dangerous keywords and the corresponding danger level, including pre-defined different interaction strategies such as comforting, guiding or emergency assistance.

10. An AI character interaction system for smart lollipop, characterized in that: The smart lollipop includes a smart handle and multiple lollipop heads with different physical shapes. The smart handle is in communication with the user terminal. Each lollipop head is detachably mounted on the front end of the smart handle. Each lollipop head with different physical shapes has a built-in character identification module. The AI ​​character interaction system includes: An identification unit, configured to respond to an operation of plugging a current sugar head into the front end of the smart handle and read the character identification module in the current sugar head through the smart handle to load the current AI character; An analysis unit, configured to collect voice data emitted by the user through the smart lollipop and perform emotion analysis and intention analysis to obtain corresponding emotion categories and intention information; A questioning unit, configured to input the voice data, emotion category, and intention information into a pre-trained personalized interaction model corresponding to the current AI character for question-answering processing, and generate a response voice that conforms to the style of the current AI character; The answering unit is used to feed back the answering voice to the user through the smart lollipop.

Citation Information

Patent Citations

  • Intelligent cross-dressing robot and scene intelligentizing method

    CN106621354A

  • Voice emotion interaction method, computer equipment and computer readable storage medium

    CN110085221A

  • Bone conduction sound production device, sound production processing method, equipment, medium and program product

    CN116528683A

  • Intelligent toy control method and system based on multi-round dialogues

    CN119207397A

  • Virtual character interaction control method, apparatus and device, and computer program product

    CN119488706A