A method and system for AI character interaction in a smart lollipop
By combining a smart handle with a detachable candy head, the smart lollipop solves the problem of the lack of dynamic semantic understanding in existing smart lollipops by utilizing emotion analysis and personalized interaction models, and achieves the effects of multi-role interaction and emotional companionship.
Patent Information
- Application Number
- CN202510673920.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2026-03-06
- Estimated Expiration
- 2045-05-23
Smart Images

Figure CN120580997B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to an AI character interaction method and system for a smart lollipop. Background Technology
[0002] In recent years, smart lollipops, as an interactive electronic consumer product, have gradually evolved from simple candy-like forms towards more integrated functions. Existing products on the market typically focus on basic voice playback and recording functions, such as pre-storing songs and story audio via built-in memory, or allowing users to record short voice messages for playback. While these products offer a degree of entertainment, their interaction modes remain at a stage of one-way output or mechanical repetition, lacking dynamic semantic understanding and personalized interaction capabilities; they fail to meet users' deeper needs for multi-role companionship and emotional communication. Summary of the Invention
[0003] The purpose of this invention is to provide an AI character interaction method and system for smart lollipops, aiming to solve the problem that existing smart lollipops lack dynamic semantic understanding and personalized interaction capabilities.
[0004] In a first aspect, embodiments of the present invention provide an AI character interaction method for a smart lollipop. The smart lollipop includes a smart handle and multiple candy heads with different physical shapes. The smart handle is communicatively connected to a user terminal. Each candy head is detachably installed at the front end of the smart handle, and each type of candy head has a built-in character identification module. The AI character interaction method includes:
[0005] In response to the current candy head being plugged into the front end of the smart handle, the smart handle reads the character identification module in the current candy head to load the current AI character;
[0006] The smart lollipop collects voice data from users and performs emotion and intent analysis to obtain corresponding emotion categories and intent information.
[0007] The voice data, emotion category, and intent information are input into a pre-trained personalized interaction model corresponding to the current AI character for question-and-answer processing, generating a response voice that matches the style of the current AI character.
[0008] The smart lollipop transmits the response voice back to the user.
[0009] Secondly, embodiments of the present invention provide an AI character interaction system for a smart lollipop. The smart lollipop includes a smart handle and multiple candy heads with different physical shapes. The smart handle is communicatively connected to a user terminal. Each candy head is detachably installed at the front end of the smart handle. Each type of candy head has a built-in character identification module. The AI character interaction system includes:
[0010] The identification unit is used to respond to the operation of the current candy head being plugged into the front end of the smart handle, and to read the role identification module in the current candy head through the smart handle to load the current AI role;
[0011] The analysis unit is used to collect voice data emitted by the user through the smart lollipop and perform emotion analysis and intent analysis to obtain the corresponding emotion category and intent information;
[0012] The questioning unit is used to input the voice data, emotion category and intent information into the pre-trained personalized interaction model corresponding to the current AI role for question and answer processing, and generate response voice that conforms to the style of the current AI role;
[0013] A response unit is used to transmit the response voice back to the user via the smart lollipop.
[0014] The beneficial effects of the embodiments of the present invention are as follows:
[0015] By coupling hardware form, emotion analysis and dynamic response technologies, the existing smart lollipop is upgraded from a one-way playback tool to an AI interactive terminal with emotion perception, role evolution and personalized dialogue capabilities, filling the technological gap in multimodal interaction and emotional companionship in the existing technology. Attached Figure Description
[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a flowchart illustrating the AI character interaction method provided in an embodiment of the present invention.
[0018] Figure 2 This is a schematic diagram of a sub-process of step S101 provided in an embodiment of the present invention.
[0019] Figure 3 This is a schematic diagram of a sub-process of step S102 provided in an embodiment of the present invention.
[0020] Figure 4This is a schematic diagram of the sub-process of step S305 provided in an embodiment of the present invention.
[0021] Figure 5 This is a schematic diagram of a sub-process of step S306 provided in an embodiment of the present invention.
[0022] Figure 6 This is a schematic diagram of a sub-process of step S103 provided in an embodiment of the present invention.
[0023] Figure 7 This is another flowchart illustrating the AI character interaction method provided in an embodiment of the present invention.
[0024] Figure 8 This is another flowchart illustrating the AI character interaction method provided in an embodiment of the present invention.
[0025] Figure 9 This is a schematic block diagram of an AI character interaction system provided in an embodiment of the present invention. Detailed Implementation
[0026] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0027] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0028] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0029] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0030] Please see Figure 1 , Figure 1 This is a flowchart illustrating the AI character interaction method for a smart lollipop provided in an embodiment of the present invention.
[0031] To facilitate understanding of this application, the smart lollipop of the present invention will first be described.
[0032] This invention provides a smart lollipop, comprising a smart handle and multiple candy heads with different physical shapes. The smart handle is communicatively connected to a user terminal. Each candy head is detachably installed at the front end of the smart handle, and each type of candy head has a built-in character identification module. The smart handle is used to trigger the reading and loading of the character identification module in the candy head when the candy head is installed at the front end of the smart handle.
[0033] In this embodiment, by intelligently recognizing the character identifier module in the candy head, the smart controller can provide users with an interactive experience related to that character.
[0034] Furthermore, the candy head has a built-in NFC chip, which is used to store the character identification information corresponding to the current candy head to form a character identification module.
[0035] Furthermore, the candy head and the smart controller communicate via methods such as a USB connector, NFC, or Bluetooth; preferably, the candy head and the smart controller are connected via a plug-in USB connector.
[0036] Furthermore, the candy head is equipped with a bone conduction vibration component, which transmits the user's voice data to the smart handle through the bone conduction vibration component. The smart handle has a built-in communication module that communicates with the user terminal to send the voice data to the user terminal for processing.
[0037] Furthermore, the smart handle is equipped with a radio module for receiving voice data emitted by the user and a speaker module for responding to the user.
[0038] Based on the above-mentioned smart lollipop solution, the AI character interaction method of the smart lollipop will be introduced in detail below.
[0039] like Figure 1 As shown, the AI character interaction method provided in this embodiment of the invention includes steps S101-S104.
[0040] S101. In response to the current operation of plugging the candy head into the front end of the smart controller, the smart controller reads the character identification module in the current candy head to load the current AI character;
[0041] In this step, the detachable candy head's built-in character identification module and the smart handle's linkage design ensure that each physical candy head corresponds to a unique AI character identity. When users change candy heads with different material shapes, different AI characters are automatically triggered without manually switching programs. This effectively solves the problem of the disconnect between physical shapes and digital content, directly mapping changes in physical form to changes in AI personality traits, constructing an immersive interactive scene where "shape is character," and significantly enhancing the user's perception of the realism of multiple characters accompanying them.
[0042] In this step, AI characters are loaded from a pre-created character profile library. This library contains a 3D profile for each AI character, including basic attributes, language style characteristics, and knowledge domain. The 3D profile can include: character metadata: {character type, historical context, personality matrix [rigor, liveliness, curiosity]}; style corpus: {high-frequency word list, sentence template library, rhetorical device distribution}; and domain knowledge topology: {main domain knowledge graph, cross-domain association rules, list of taboo topics}. Specific AI character types can be set according to different user groups.
[0043] S102. Collect voice data emitted by the user through the smart lollipop and perform emotion analysis and intent analysis to obtain the corresponding emotion category and intent information;
[0044] S103. Input the voice data, emotion category and intent information into the pre-trained personalized interaction model corresponding to the current AI character for question and answer processing, and generate response voice that matches the style of the current AI character;
[0045] In steps S102-S103, real-time emotion category recognition and intent analysis based on user-generated voice data, combined with dynamic response generation from a pre-trained personalized interaction model, enable the system to adjust its feedback strategy according to the user's emotional state (e.g., excitement, depression) and interaction intent (e.g., asking questions, confiding). Compared to existing technologies that rely on fixed command matching or static audio playback, this significantly improves the contextual adaptability and emotional resonance of the dialogue content, solving the core problem of monotonous user experience caused by mechanical repetitive responses.
[0046] S104. The smart lollipop will provide voice feedback to the user.
[0047] In this step, the smart handle is equipped with an external speaker module. It can convert the audio signal of the response voice into a mechanical vibration signal, and use bone conduction technology to feed the mechanical vibration signal back to the candy head. This means that the user can interact directly with the AI character while holding the candy head in their mouth; and / or, the response voice can be played through the external speaker module on the smart handle, so that the user can hear and interact even after removing the candy head from their mouth.
[0048] In this embodiment, by coupling three technologies—hardware form, emotion analysis, and dynamic response—the existing smart lollipop is upgraded from a one-way playback tool into an AI interactive terminal with emotion perception, role evolution, and personalized dialogue capabilities, filling the technological gap in multimodal interaction and emotional companionship in the existing technology.
[0049] In some embodiments, the various models analyzing the voice data (i.e., personalized interactive models and models such as the dual-stream deep spiking neural network described below) can be deployed online in the user terminal (APP application). The smart lollipop connects to the user terminal via wired or wireless communication to achieve bidirectional data transmission and reception. Obviously, the smart controller can also have a built-in data processing module. The trained models can be deployed offline in the smart controller's data processing module to achieve offline application in environments without a network.
[0050] In one embodiment, the candy head and the smart handle can communicate with each other through various means after being plugged in or unplugged, such as USB connector, NFC, Bluetooth, etc.
[0051] In one embodiment, the candy head and the smart controller are connected via a USB connector. The candy head has a built-in NFC chip, which stores the character identification information corresponding to the current candy head to form a character identification module. Based on this, as... Figure 2 As shown, in one embodiment, step S101 includes steps S201-S203.
[0052] S201. In response to the current operation of plugging the candy head into the front end of the smart controller, read the role identification information in the current candy head through the smart controller;
[0053] S202. Obtain the unique identifier from the role identification information;
[0054] S203. Based on the unique identification code, match the corresponding AI character in the user terminal, load the character, and then connect to the personalized exclusive interaction model corresponding to the current AI character.
[0055] In this embodiment, based on the candy head with the current material shape selected by the user, after the candy head is plugged into the front end of the smart handle, the smart handle can read the NFC chip in the candy head to obtain the character identification information. Then, based on the unique identification code in the character identification information, the AI character corresponding to the current candy head is automatically loaded. Throughout the process, through an intelligent identification and matching process, rapid interaction between the user and the AI character is achieved; this not only improves the user's ease of operation but also enhances the fun and immersion of entertainment interaction, bringing users a richer and more diverse user experience.
[0056] In one embodiment, the candy head contains a bone conduction vibration component, and the smart handle has a built-in communication module for communicating with the user terminal. Based on this, as... Figure 3 As shown, step S102 includes steps S301-S306.
[0057] S301. Collects voice data emitted by the user through bone conduction technology via the candy head, and / or receives voice data emitted by the user through the sound receiving module set on the smart handle.
[0058] S302. Extract the speech waveform features of the speech data and input them into the main branch of the dual-stream deep pulse neural network. Extract the temporal emotion features in the speech data through multi-layer pulse convolution kernels to generate a temporal pulse sequence.
[0059] S303. Extract the MFCC coefficients of the speech data and convert them into sparse pulse representations by a frequency domain pulse encoder. Then, input the MFCC coefficients into the side branch of the dual-stream deep pulse neural network and use formant-sensitive pulse pooling layers to capture intent-related frequency domain features to generate a frequency domain pulse stream.
[0060] S304. The mutual information entropy value of the time-domain pulse sequence and the frequency-domain pulse stream is calculated using the pulse cross-correlation algorithm to generate a spatiotemporal coherence matrix, and cross-modal joint features are extracted from the spatiotemporal coherence matrix through pulse firing synchronization detection.
[0061] S305. Using the pulse sequence dynamic time warping algorithm, the cross-modal joint features are dynamically compared with the preset emotion feature template library through the emotion classifier, and the corresponding emotion category is output.
[0062] In step S305, emotion is not only reflected by the temporal features of speech, but also contains important information in the frequency domain features. Therefore, emotion analysis based on cross-modal joint features is more accurate.
[0063] S306. Perform pulse sequence symbolization reasoning on the frequency domain pulse stream through the intent analyzer to locate the corresponding intent information in the preset intent map library;
[0064] In step S306, intent analysis focuses more on the vocabulary and grammatical structure in speech. After processing, frequency domain features can be transformed into symbol sequences, which are more suitable for state machine reasoning.
[0065] The Dual-Stream Deep Spiking Neural Network (DSNN) in this embodiment is a computational model that combines a Spiking Neural Network (SNN) and a dual-stream architecture. The DSNN includes a main branch and side branches. Specifically, the main branch consists of multiple stacked spike convolution kernels used to extract temporal emotional features from the original speech waveform. Each spike convolution kernel embeds a spike activation function, generating a temporal spike sequence through a membrane potential accumulation-firing mechanism. The side branches contain a frequency-domain spike encoder and a formant-sensitive spike pooling layer. The frequency-domain spike encoder maps MFCC coefficients (Mel-Frequency Cepstral Coefficients) to a sparse spike representation. The formant-sensitive spike pooling layer selects frequency band features related to speech intent through a tunable spike firing threshold. A spatiotemporal feature fusion module connects the main branch and side branches, using a spike synchronization detection algorithm to align the temporal and frequency-domain spike streams, and generating a cross-modal joint feature vector through spike cross-correlation calculation.
[0066] In this embodiment, based on the process of steps S301-S306, the user's voice data is analyzed jointly by a dual-stream deep spiking neural network to determine the emotion and intent. This can accurately decode the user's emotional state and intent information, providing the user with more considerate and personalized interactive services.
[0067] like Figure 4 As shown, in one embodiment, step S305 includes:
[0068] S401. Pre-build a library of preset emotion feature templates containing multiple types of emotions;
[0069] S402. Perform time series slicing on the cross-modal joint features to generate multiple continuous real-time pulse trajectory segments with a preset duration (e.g., 200-500ms).
[0070] S403. The pulse sequence dynamic time warping algorithm is used to elastically align multiple real-time pulse trajectory segments with multiple standard pulse trajectory segments of each emotion before performing similarity comparison to obtain the cross-modal joint features and the trajectory similarity corresponding to each emotion.
[0071] S404. Select multiple candidate emotions from each trajectory similarity in descending order, and after performing multi-scale emotion probability estimation, output the candidate emotion with the highest probability as the target emotion.
[0072] In this embodiment, the emotion feature template library can include multiple emotions such as joy, sadness, anger, surprise, fear, and calmness, and can be specifically designed according to the target audience. Each emotion category stores a three-dimensional feature template: Dimension 1: Pulse density gradient histogram (statistically counting the rate of change of pulse quantity according to a preset time window); Dimension 2: Cross-layer pulse synchronization matrix (recording the frequency of coordinated triggering of pulses in the main time-domain branch and the side-branch in the frequency domain within a preset time window); Dimension 3: Membrane potential oscillation spectrum (extracting the energy distribution of preset feature frequency bands).
[0073] In this embodiment, the Pulse Sequence Dynamic Time Warping Algorithm (PS-DTW) is used to solve the alignment problem between time series of different lengths. That is, the pulse trajectory segments to be compared need to be "flexibly aligned" to solve problems such as misalignment of pulse segments caused by different speech rates. Only after "flexible alignment" can their similarity be compared more objectively and accurately.
[0074] In this embodiment, multi-scale emotion probability estimation is performed on multiple candidate emotions. It is necessary to comprehensively consider the features of different scales (such as pulse density change rate, oscillation frequency, etc.) to improve the accuracy of emotion recognition. Finally, the probability results of each scale are fused to determine the target emotion.
[0075] like Figure 5 As shown, in one embodiment, step S306 includes:
[0076] S501. The pulse timing logic rule engine is used to discretize the emission time of the frequency domain pulse stream and map the pulse emission cluster into a symbol sequence, where each symbol corresponds to a semantic primitive in the preset intent map library.
[0077] S502. In the hierarchical topology of the preset intent map, the symbol sequence is matched with the intent nodes in the preset intent map library for spatiotemporal similarity. The preset intent map includes an intent-action association layer constructed by a knowledge graph and a dynamic context parameter layer driven by user history behavior.
[0078] In this step, the intent-action association layer refers to the static logical relationship based on the knowledge graph (such as "user says 'play music' → triggers music app"). The dynamic context parameter layer is used to extract user's historical behavior features, generate context weight coefficients, and correct the spatiotemporal similarity score. Spatiotemporal similarity matching refers to matching the two similarities between the symbol sequence and the intent node from the time dimension (sequence temporal similarity between the symbol sequence and the intent node) and the spatial dimension (feature similarity between the symbol sequence and the intent node), and then performing a weighted fusion to obtain the spatiotemporal similarity.
[0079] S503. When the spatiotemporal similarity exceeds the first threshold, output the intent information and association confidence of successful localization.
[0080] S504 When the spatiotemporal similarity does not exceed the first threshold, the feedback fine-tuning mechanism of the dual-stream deep spiking neural network is triggered to readjust the frequency band weight coefficients of the formant-sensitive pulse pooling layer.
[0081] In this embodiment, based on steps S501-S505, the complex frequency domain pulse signal flow is transformed into interpretable semantic primitives through pulse sequence symbolization and a hierarchical intent matching mechanism. Combined with knowledge graphs and dynamic user behavior data, high-precision intent recognition is achieved. Simultaneously, a feedback fine-tuning mechanism is introduced to automatically optimize the neural network frequency band weights when matching fails, enhancing the system's adaptability to non-steady-state signals.
[0082] like Figure 6 As shown, in one embodiment, step S103 includes:
[0083] S601. Input the voice data, emotion category and intent information into the pre-trained personalized interaction model corresponding to the current AI character;
[0084] S602. Extract the style parameter set of the current AI character. The style parameter set includes corpus preference coefficients, intonation feature vectors, and domain knowledge weights.
[0085] S603. Map emotion categories to emotion reinforcement factors;
[0086] S604. Convert intent information into a domain knowledge index;
[0087] S605. By fusing and calculating the corpus preference coefficient and sentiment enhancement factor through the style adaptation layer in the personalized exclusive interaction model, dynamic style constraints are generated.
[0088] S606. Based on domain knowledge weights and domain knowledge indexes, retrieve matching domain corpora from a pre-constructed multi-dimensional knowledge graph;
[0089] S607. Input speech data, dynamic style constraints and domain corpus into the dialogue generation network, and generate response speech that conforms to the language characteristics of the current AI role through adversarial style transfer.
[0090] In this embodiment, based on steps S601-S607, deep customization of the AI character's responses is achieved. The core lies in dynamically fusing user emotions and corpus preferences using style parameters, combined with adversarial transfer learning technology. This ensures that the response speech retains the unique linguistic characteristics of the current AI character (such as specific intonation and domain-specific expression habits) while adapting in real-time to the emotional intensity and professional knowledge requirements of the dialogue scenario. This significantly improves the realism, emotional resonance, and response accuracy of the smart lollipop's interaction with the user.
[0091] like Figure 7 As shown, in one embodiment, the AI character interaction method for the smart lollipop further includes:
[0092] S701. Based on voice data, emotion category, intent information and corresponding response voice, generate emotional memory nodes with time dimension markers, and establish a dynamic emotion map composed of multiple emotional memory nodes.
[0093] S702. Extract the text information from the speech data and perform semantic parsing to obtain at least three levels of interest entities, including basic entities, combined entities, and abstract concept entities;
[0094] In this step, basic entities refer to atomic-level semantic units directly extracted through syntactic parsing, corresponding to observable objects in the physical world; composite entities refer to complex semantic units composed of basic entities through dependency relationships, representing specific events or behavioral patterns; and abstract concept entities refer to surreal mapping relationships generated through semantic reasoning, representing users' potential values or cognitive tendencies.
[0095] S703. Cross-modal association between interest entities and corresponding emotional memory nodes is established to create a multi-dimensional association matrix of entity-emotion-time, and a user interest knowledge graph composed of multiple multi-dimensional association matrices is established.
[0096] S704. When the rate of change of the emotional intensity of a specific interest entity within a continuous time window exceeds the threshold, the corresponding emotional memory node is activated.
[0097] S705. Adjust the response parameters of the current AI character based on the activated emotional memory nodes. The response parameters include voice tone parameters and response content generation strategy.
[0098] In this embodiment, based on steps S701-S705, a deep understanding and memory tracking of user emotions and interests by the AI character is achieved by constructing a dynamic sentiment graph and a user interest knowledge graph. The core of this is: using time-dimensional labeling and a multi-dimensional association matrix to capture changes in user emotion intensity and the evolution of interest entities; and adjusting the AI response strategy in real time through cross-modal data fusion, enabling interactive responses to possess emotional coherence and interest relevance, significantly improving the naturalness and adaptability of personalized interaction.
[0099] In this embodiment, in order to better construct a user interest knowledge graph, an open-ended question strategy can be used to explore the user's interests when the user uses the service for the first time. Alternatively, information entered by the user can be collected through the user terminal, and a preliminary user interest knowledge graph can be constructed based on this information.
[0100] For example, suppose a time-stamped emotional memory node represents a user's weekly outdoor activity. Once activated, the AI character can adjust its response parameters to use a more positive and encouraging tone of voice. Simultaneously, the response content generation strategy might revolve around the outdoor activity theme, offering relevant suggestions, activity arrangements, or encouraging words. This adjustment not only demonstrates the AI character's sensitive capture of user emotional changes but also showcases its continuous attention to user interests and personalized responses. In this way, the AI character can more naturally integrate into the user's daily life, establishing deep emotional connections and shared interests, thereby enhancing user experience and satisfaction.
[0101] like Figure 8 As shown, in one embodiment, the AI character interaction method for the smart lollipop further includes:
[0102] S801. Extract the text information from the speech data and perform word segmentation to obtain word groups;
[0103] S802. Perform dangerous keyword detection on word segments. If a word segment triggers a crisis warning mechanism, assess the danger level of the word segment according to the dangerous keyword level table.
[0104] S803. Adjust the current AI's interaction strategy based on the triggered dangerous keywords and the corresponding danger level, including different pre-defined interaction strategies such as soothing, guiding, or emergency assistance.
[0105] In this embodiment, based on the process of steps S801-S803, dynamic safety control of AI interaction strategy is realized through voice text analysis and dangerous keyword classification detection mechanism. It can identify potential dangerous contexts (such as violence, self-harm tendencies, etc.) in real time during the user's use of smart lollipop, and intelligently switch to soothing guidance or start emergency help procedure according to risk level, effectively improving the level of emotional emergency response.
[0106] In this embodiment, when a dangerous keyword and its corresponding danger level reach the level requiring intervention, the emergency contact function is activated through the hardware interface of the smart lollipop, automatically dialing a preset emergency contact number or sending an emergency help message. In addition, the entire process of the crisis event is recorded, including user voice, emotional reactions, interaction content, and response measures taken, providing data support for subsequent crisis analysis and handling. Finally, based on the crisis handling results and user feedback, the interaction logic of the AI role and the crisis early warning mechanism are continuously optimized and improved to enhance user experience and safety.
[0107] This invention also provides an AI character interaction system for a smart lollipop. The smart lollipop includes a smart handle and multiple candy heads with different physical shapes. The smart handle is connected to a user terminal for communication. Each candy head is detachably installed at the front end of the smart handle. Each type of candy head has a built-in character identification module.
[0108] The AI character interaction system of this invention is used to execute any of the aforementioned AI character interaction methods. Specifically, please refer to... Figure 9 , Figure 9 This is a schematic block diagram of the AI role interaction system provided in the embodiments of the present invention.
[0109] like Figure 9 As shown, the AI role interaction system 900 includes: a recognition unit 901, an analysis unit 902, an inquiry unit 903, and a response unit 904.
[0110] The identification unit 901 is used to respond to the operation of the current candy head being plugged into the front end of the smart handle, and to read the role identification module in the current candy head through the smart handle to load the current AI role;
[0111] The analysis unit 902 is used to collect voice data emitted by the user through the smart lollipop and perform emotion analysis and intent analysis to obtain the corresponding emotion category and intent information;
[0112] The questioning unit 903 is used to input voice data, emotion category and intent information into the pre-trained personalized interaction model corresponding to the current AI role for question and answer processing, and generate response voice that matches the style of the current AI role;
[0113] The response unit 904 is used to provide voice feedback to the user via a smart lollipop.
[0114] This system upgrades the existing smart lollipop from a one-way playback tool to an AI interactive terminal with emotional perception, role evolution, and personalized dialogue capabilities by coupling three technologies: hardware form, emotion analysis, and dynamic response. This fills the technological gap in multimodal interaction and emotional companionship in the existing technology.
[0115] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the system described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0116] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. An AI character interaction method of a smart lollipop, characterized in that, The smart lollipop includes a smart handle and a plurality of physically different lollipop heads, the smart handle is in communication connection with a user terminal, each of the lollipop heads is detachably installed at a front end of the smart handle, and each of the physically different lollipop heads is internally provided with a role identification module; the AI role interaction method includes: In response to an operation of plugging the current lollipop head into the front end of the smart handle, reading the role identification module in the current lollipop head through the smart handle to load the current AI role; Collecting voice data emitted by a user through the smart lollipop and performing emotion analysis and intent analysis to obtain corresponding emotion categories and intent information; Inputting the voice data, emotion categories and intent information into a pre-trained personalized exclusive interaction model corresponding to the current AI role to perform question and answer processing and generate response voice conforming to the style of the current AI role; specifically including: inputting the voice data, emotion categories and intent information into the pre-trained personalized exclusive interaction model corresponding to the current AI role; extracting a style parameter set of the current AI role, the style parameter set including a corpus preference coefficient, a tone feature vector and a domain knowledge weight; mapping the emotion categories to an emotion reinforcement factor; converting the intent information to a domain knowledge index; performing fusion calculation on the corpus preference coefficient and the emotion reinforcement factor through a style adaptation layer in the personalized exclusive interaction model to generate a dynamic style constraint condition; based on the domain knowledge weight and the domain knowledge index, retrieving matching domain corpus from a pre-constructed multi-dimensional knowledge graph; inputting the voice data, dynamic style constraint condition and domain corpus into a dialogue generation network to generate response voice conforming to the language characteristics of the current AI role through adversarial style transfer; Feeding back the response voice to the user through the smart lollipop. 2.The AI character interaction method of the smart lollipop according to claim 1, wherein, The lollipop head and the smart handle are in plug-in connection through a USB connector, the lollipop head is internally provided with an NFC chip, and the NFC chip is used to store role identification information corresponding to the current lollipop head to constitute the role identification module; The operation of plugging the current lollipop head into the front end of the smart handle, reading the role identification module in the current lollipop head through the smart handle to load the current AI role, includes: In response to an operation of plugging the current lollipop head into the front end of the smart handle, reading the role identification information in the current lollipop head through the smart handle; Obtaining a unique identification code in the role identification information; According to the unique identification code, matching the corresponding AI role in the user terminal and loading the role to access the personalized exclusive interaction model corresponding to the current AI role. 3.The AI character interaction method of the smart lollipop according to claim 1, wherein, The lollipop head is internally provided with a bone conduction vibration component, and the smart handle is internally provided with a communication module in communication with the user terminal; The operation of plugging the current lollipop head into the front end of the smart handle, reading the role identification module in the current lollipop head through the smart handle to load the current AI role, includes: Collecting voice data emitted by a user through the lollipop head using bone conduction technology, and / or receiving voice data emitted by a user through a sound receiving module arranged on the smart handle; extracting a speech waveform feature of the speech data and inputting the speech waveform feature into a main branch of a double-flow deep pulse neural network, extracting a time-domain emotion feature in the speech data through a plurality of pulse convolution kernels, and generating a time-domain pulse sequence; extracting an MFCC coefficient of the speech data, converting the MFCC coefficient into a sparse pulse representation through a frequency-domain pulse encoder, and inputting the sparse pulse representation into a side branch of the double-flow deep pulse neural network, capturing an intention-related frequency-domain feature through a formant-sensitive pulse pooling layer, and generating a frequency-domain pulse flow; calculating mutual information entropy values of the time-domain pulse sequence and the frequency-domain pulse flow through a pulse cross-correlation algorithm, generating a space-time coherence matrix, and extracting a cross-modal joint feature from the space-time coherence matrix through pulse firing synchronization detection; comparing the cross-modal joint feature with a preset emotion feature template library through a pulse sequence dynamic time warping algorithm of an emotion classifier, and outputting a corresponding emotion category; performing pulse sequence symbolic reasoning on the frequency-domain pulse flow through an intention analyzer to locate corresponding intention information in a preset intention map library. 4.The AI character interaction method of the smart lollipop according to claim 3, characterized in that, The method of comparing the cross-modal joint feature with the preset emotion feature template library through the pulse sequence dynamic time warping algorithm of the emotion classifier and outputting the corresponding emotion category comprises: pre-constructing a preset emotion feature template library containing multiple emotion categories; time series slicing the cross-modal joint feature to generate a plurality of continuous real-time pulse trajectory segments of a preset time length; performing elastic alignment of a plurality of the real-time pulse trajectory segments with a plurality of standard pulse trajectory segments of each emotion category through a pulse sequence dynamic time warping algorithm, and then performing similarity comparison to obtain a trajectory similarity of the cross-modal joint feature corresponding to each emotion category; selecting a plurality of candidate emotions from the trajectory similarities in order from high to low, performing multi-scale emotion probability estimation, and outputting a candidate emotion with the highest probability as a target emotion. 5.The AI character interaction method of the smart lollipop according to claim 3, characterized in that, The method of performing pulse sequence symbolic reasoning on the frequency-domain pulse flow through the intention analyzer to locate corresponding intention information in the preset intention map library comprises: performing firing time discretization coding on the frequency-domain pulse flow through a pulse timing logic rule engine to map pulse firing clusters into a symbol sequence, wherein each symbol corresponds to a semantic primitive in the preset intention map library; performing space-time similarity matching of the symbol sequence with an intention node in the preset intention map library in a hierarchical topological structure of the preset intention map, wherein the preset intention map includes an intention-action association layer constructed by a knowledge graph and a dynamic context parameter layer driven by user historical behavior; when the space-time similarity exceeds a first threshold, outputting the located intention information and the associated confidence; when the space-time similarity does not exceed the first threshold, triggering a feedback fine-tuning mechanism of the double-flow deep pulse neural network to re-adjust a frequency band weight coefficient of the formant-sensitive pulse pooling layer. 6.The AI character interaction method of the smart lollipop according to claim 1, wherein, The smart handle is provided with a external playing module; The method of feeding back the response speech to the user through the smart lollipop comprises: The sound signal of the response voice is converted into a mechanical vibration signal, and the mechanical vibration signal is fed back to the sugar head by using bone conduction technology; And / or The response voice is played by a loudspeaker module on the smart handle. 7.The AI character interaction method of the smart lollipop according to claim 1, wherein, Also includes: According to the voice data, emotion category, intention information and corresponding response voice, a sentimental memory node with time dimension mark is generated, and a dynamic sentimental map composed of multiple sentimental memory nodes is established; Extracting the text information of the voice data and performing semantic analysis to obtain at least three levels of interest entities including basic entities, combined entities and abstract concept entities; Cross-modal association is performed between the interest entities and the corresponding sentimental memory nodes to establish a multi-dimensional association matrix of entities-sentiment-time, and a user interest knowledge graph composed of multiple multi-dimensional association matrices is established; When it is detected that the sentimental intensity change rate of a specific interest entity within a continuous time window exceeds a threshold, the corresponding sentimental memory node is activated; Based on the activated sentimental memory node, the response parameters of the current AI role are adjusted, including voice tone parameters and response content generation strategies. 8.The AI character interaction method of the smart lollipop according to claim 1, wherein, Also includes: Extracting the text information of the voice data and performing word segmentation processing to obtain a word segmentation group; If the word segmentation group triggers a crisis warning mechanism, the word segmentation group is evaluated for danger level according to a danger keyword level table; According to the triggered danger keyword and the corresponding danger level, the interaction strategy of the current AI is adjusted, including different interaction strategies such as pre-prepared pacification, guidance or emergency help.
9. An AI character interaction system of a smart lollipop, characterized in that, The smart lollipop includes a smart handle and multiple sugar heads with different physical shapes, the smart handle is in communication connection with a user terminal, each sugar head is detachably installed at the front end of the smart handle, and a role identification module is built-in in each sugar head with different physical shapes; the AI role interaction system includes: An identification unit is configured to read the role identification module in the current sugar head through the smart handle to load the current AI role in response to the operation of plugging the current sugar head into the front end of the smart handle; An analysis unit is configured to collect voice data emitted by a user through the smart lollipop and perform emotion analysis and intention analysis to obtain corresponding emotion categories and intention information. The inquiring unit is configured to input the voice data, the emotion category, and the intention information into a pre-trained personalized exclusive interaction model corresponding to the current AI role to perform question and answer processing, and generate response voice conforming to the style of the current AI role. Specifically, the voice data, the emotion category, and the intention information are input into the pre-trained personalized exclusive interaction model corresponding to the current AI role. The style parameter set of the current AI role is extracted, and the style parameter set includes a corpus preference coefficient, a tone feature vector, and a domain knowledge weight. The emotion category is mapped to an emotion reinforcement factor. The intention information is converted into a domain knowledge index. The corpus preference coefficient and the emotion reinforcement factor are fused and calculated by a style adaptation layer in the personalized exclusive interaction model to generate a dynamic style constraint condition. Based on the domain knowledge weight and the domain knowledge index, matching domain corpus is retrieved from a pre-constructed multi-dimensional knowledge graph. The voice data, the dynamic style constraint condition, and the domain corpus are input into a dialogue generation network to generate response voice conforming to the language characteristics of the current AI role through adversarial style transfer. The response unit is configured to feed back the response voice to the user through the smart lollipop.
Citation Information
Patent Citations
Intelligent cross-dressing robot and scene intelligentizing method
CN106621354A
Voice emotion interaction method, computer equipment and computer readable storage medium
CN110085221A
Bone conduction sound production device, sound production processing method, equipment, medium and program product
CN116528683A