Children's AI dialogue generation method, system and storage medium based on IP character setting
By constructing an IP character personality constraint matrix and performing triplet decoupling, combined with a multi-objective personality constraint loss function and a real-time correction mechanism, the problems of character consistency and authenticity in the children's AI dialogue generation system are solved, and the deep integration and precise matching of dialogue content and IP character settings are achieved, thereby improving the acceptance of educational content and learning effects.
Patent Information
- Application Number
- CN202510978660.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-07-16
AI Technical Summary
Existing AI dialogue generation systems for children have shortcomings in personalized character creation and deep emotional connection, making it difficult to maintain character consistency and authenticity, resulting in the generated dialogue content deviating from the character settings.
By constructing the IP character personality constraint matrix and performing triple decoupling, personality feature sub-vectors, style feature sub-vectors, and knowledge feature sub-vectors are generated. Combined with a multi-objective personality constraint loss function and a real-time correction mechanism, the dialogue content is ensured to conform to the IP character settings.
It achieves deep integration and precise matching between dialogue content and IP character personality, avoids character identity drift, and improves the acceptance of educational content and learning effects.
Smart Images

Figure CN120492596B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of AI dialogue generation technology, and in particular to a method, system and storage medium for generating children's AI dialogue based on IP character settings. Background Art
[0002] With the widespread application of artificial intelligence technology in children's education, children's AI dialogue systems based on IP characters have become an important means of enhancing children's learning interest and interactive experience. Existing children's interactive applications mostly use pre-recorded audio or general AI voice technology. While these technologies meet basic human-computer interaction needs to a certain extent, they still have significant shortcomings in personalized character creation and deep emotional connection. Traditional dialogue generation systems often lack precise modeling and constraint mechanisms for specific IP characters, making it difficult to maintain character consistency and authenticity during conversations, thereby affecting children's immersion and learning outcomes.
[0003] Current AI dialogue generation technology for IP characters suffers from several key technical deficiencies in terms of character constraints. Existing methods oversimplify the processing of character data, using only simple label descriptions or single-dimensional feature representations, which cannot fully capture the complex personality traits, unique language style, and multi-dimensional constraints between areas of expertise of IP characters. Secondly, the coupling between character characteristics and the semantic space of dialogue is insufficient, resulting in the generated dialogue content often deviating from the character setting, resulting in problems such as confusing personality expression and inconsistent language style. In addition, there is a lack of an effective fusion mechanism between the IP character knowledge base and personalized expression, making it easy for characters to exceed their set knowledge scope or express themselves in a way that is inconsistent with their character setting when answering questions. Summary of the Invention
[0004] The main purpose of the present invention is to provide a method, system and storage medium for generating children's AI dialogue based on IP character settings. The present invention can continuously monitor and correct character setting deviations during the dialogue generation process, ensuring that the generated content always conforms to the IP character settings, and effectively solves the technical problem of character identity drift in multiple rounds of dialogue.
[0005] To achieve the above objectives, the present invention provides a method for generating children's AI dialogue based on IP character settings, comprising the following steps:
[0006] Obtain IP character personality data and construct IP character personality constraint matrix;
[0007] Perform triple decoupling on the IP character constraint matrix to obtain a personality feature sub-vector, a style feature sub-vector, and a knowledge feature sub-vector;
[0008] Performing personality weighted attention calculation on the user input text based on the personality feature sub-vector, the style feature sub-vector, and the knowledge feature sub-vector to obtain an attention weight distribution matrix;
[0009] The attention weight distribution matrix is input into the multi-objective character constraint loss function for joint optimization to obtain the intermediate dialogue content that meets the IP character character constraint conditions;
[0010] Based on the intermediate dialogue content, educational content adaptation and role identity consistency verification are performed to obtain a pre-generated dialogue text, and the pre-generated dialogue text is corrected for real-time character deviation to obtain a target dialogue generation result.
[0011] Optionally, in a first implementation of the first aspect of the present invention, obtaining IP character personality data and constructing an IP character personality constraint matrix includes:
[0012] Parse and standardize the input IP character data to obtain a structured character data set containing character personality descriptions, language style characteristics, and professional knowledge areas;
[0013] Performing feature extraction and numerical quantification mapping on the character attribute information according to the structured character data set to obtain a character feature parameter set including personality constraint parameters, style constraint parameters, and knowledge constraint parameters;
[0014] Performing multi-dimensional matrix structure organization and dimension allocation based on the personality feature parameter set to obtain a multi-dimensional constraint matrix framework having a personality constraint layer, a style constraint layer, and a knowledge constraint layer;
[0015] Based on the multi-dimensional constraint matrix framework, matrix element filling and index construction are performed to obtain the IP character personality constraint matrix.
[0016] Optionally, in a second implementation of the first aspect of the present invention, performing triple decoupling on the IP character personality constraint matrix to obtain a personality feature sub-vector, a style feature sub-vector, and a knowledge feature sub-vector includes:
[0017] Perform dimensional analysis and subspace target determination based on the IP character personality constraint matrix to obtain a decoupled target parameter set;
[0018] Performing non-negative matrix decomposition on the IP character constraint matrix according to the decoupling target parameter set to obtain a personality subspace matrix, a style subspace matrix, and a knowledge subspace matrix;
[0019] Perform orthogonal constraint optimization and linear independence check based on the personality subspace matrix, the style subspace matrix, and the knowledge subspace matrix to obtain a decoupled subspace matrix set;
[0020] Feature vector extraction and vectorized dimension reduction processing are performed on the decoupled subspace matrix set to obtain personality feature subvectors, style feature subvectors, and knowledge feature subvectors.
[0021] Optionally, in a third implementation of the first aspect of the present invention, performing non-negative matrix decomposition on the IP character personality constraint matrix according to the decoupling target parameter set to obtain a personality subspace matrix, a style subspace matrix, and a knowledge subspace matrix includes:
[0022] Initializing the decomposition parameters based on the decoupling target parameter set to obtain a decomposition initialization parameter set including a gradient descent learning rate, an iteration number threshold, and a convergence error threshold;
[0023] Performing gradient descent iterative decomposition on the IP character personality constraint matrix according to the decomposition initialization parameter set to obtain a decomposition intermediate result matrix set;
[0024] Dimension allocation and subspace extraction are performed based on the decomposition intermediate result matrix set to obtain a personality subspace matrix, a style subspace matrix and a knowledge subspace matrix.
[0025] Optionally, in a fourth implementation of the first aspect of the present invention, the performing of personality weighted attention calculation on the user input text based on the personality feature sub-vector, the style feature sub-vector, and the knowledge feature sub-vector to obtain an attention weight distribution matrix includes:
[0026] Perform word embedding encoding and sequence vectorization on the user input text to obtain the input text vector sequence;
[0027] Constructing and initializing a character query matrix and a character key matrix according to the personality feature sub-vector, the style feature sub-vector, and the knowledge feature sub-vector to obtain a multi-dimensional character attention calculation matrix group;
[0028] Perform multi-head attention weight calculation and semantic attention weight calculation based on the multi-dimensional human attention calculation matrix group and the input text vector sequence to obtain a human attention weight matrix and a semantic attention weight matrix;
[0029] The human-setting attention weight matrix and the semantic attention weight matrix are gated and fused to obtain an attention weight distribution matrix.
[0030] Optionally, in a fifth implementation of the first aspect of the present invention, inputting the attention weight distribution matrix into a multi-objective character constraint loss function for joint optimization to obtain intermediate dialogue content that satisfies the IP character character constraint conditions includes:
[0031] Constructing a multi-objective personality constraint loss function based on the attention weight distribution matrix, which includes a personality constraint loss component, a style constraint loss component, and a knowledge constraint loss component;
[0032] Calculating the constraint deviation of the generated text's sentiment vector, language style features, and knowledge domain information based on the multi-objective character constraint loss function to obtain a character constraint loss value, a style constraint loss value, and a knowledge constraint loss value;
[0033] Performing weighted linear combination and AdamW optimizer iterative processing based on the personality constraint loss value, the style constraint loss value, and the knowledge constraint loss value to obtain a comprehensive optimization loss value;
[0034] The comprehensive optimization loss value is input into the dialogue generation decoder for gradient backpropagation and constraint guidance generation to obtain the intermediate dialogue content that meets the IP character setting constraints.
[0035] Optionally, in a sixth implementation of the first aspect of the present invention, performing educational content adaptation and role identity consistency verification based on the intermediate dialogue content to obtain a pre-generated dialogue text, and performing real-time character setting deviation correction on the pre-generated dialogue text to obtain a target dialogue generation result, includes:
[0036] Performing educational knowledge point retrieval and content matching calculation on the intermediate conversation content to obtain a candidate educational content library containing fire safety knowledge, emotional cognition knowledge, and social etiquette knowledge;
[0037] Performing consistency check and time decay processing on the role identity preservation factor and the dialogue history memory matrix according to the candidate education content library to obtain a screening education content set that meets the role identity consistency condition;
[0038] Based on the screened educational content set, layered fusion processing is performed on the content selection layer, the language style is adjusted on the expression conversion layer, and the character emotion is added on the emotion coloring layer to obtain role-based educational fusion content;
[0039] Performing content integration and semantic coherence checking on the role-based education fusion content and the intermediate dialogue content to obtain a pre-generated dialogue text;
[0040] The pre-generated dialogue text is corrected for deviations in the human setting in real time to obtain a target dialogue generation result.
[0041] Optionally, in a seventh implementation of the first aspect of the present invention, performing real-time character setting deviation correction on the pre-generated dialogue text to obtain a target dialogue generation result includes:
[0042] Inputting the pre-generated dialogue text into the candidate generation layer for beam search and diversified expansion to obtain a candidate dialogue text set containing multiple candidate responses;
[0043] Calculating the personality consistency score, style matching score, and knowledge accuracy score based on the candidate dialogue text set to obtain a personality conformity assessment result and a deviation detection value corresponding to each candidate response;
[0044] Comparing and judging the personality deviation threshold and adaptively adjusting the correction strength according to the deviation detection value to obtain a screening candidate response set that meets the personality constraint conditions or regenerate a trigger instruction;
[0045] The screened candidate response set is input into the persona fidelity assessment algorithm for quality assessment and optimal response selection to obtain the target dialogue generation result.
[0046] The present invention also provides a children's AI dialogue generation system based on IP character settings, including:
[0047] The acquisition module is used to obtain IP character personality data and build the IP character personality constraint matrix;
[0048] A decoupling module, configured to perform triple decoupling on the IP character constraint matrix to obtain a personality feature sub-vector, a style feature sub-vector, and a knowledge feature sub-vector;
[0049] a calculation module, configured to perform personality weight attention calculation on the user input text based on the personality feature sub-vector, the style feature sub-vector, and the knowledge feature sub-vector to obtain an attention weight distribution matrix;
[0050] A joint optimization module is used to input the attention weight distribution matrix into a multi-objective character constraint loss function for joint optimization to obtain intermediate dialogue content that meets the IP character character constraint conditions;
[0051] A real-time correction module is used to perform educational content adaptation and role identity consistency verification based on the intermediate dialogue content to obtain a pre-generated dialogue text, and to perform real-time character deviation correction on the pre-generated dialogue text to obtain a target dialogue generation result.
[0052] The present invention also provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of any of the above methods are implemented.
[0053] In summary, the technical solution provided by the present invention constructs a multi-dimensional IP character personality constraint matrix including a personality constraint layer, a style constraint layer and a knowledge constraint layer. The present invention can realize all-round and accurate quantitative modeling of IP character personality, and significantly improves the expression integrity and constraint accuracy of personality characteristics compared to the simple label description method of the prior art. The triple decoupling mechanism is used to decompose the personality constraint matrix into independent personality feature sub-vectors, style feature sub-vectors and knowledge feature sub-vectors, which avoids the problem of mutual interference between different personality dimensions in the prior art and ensures the linear independence and independent controllability of the features of each dimension. Based on the personality weight attention modulation mechanism, the present invention can simultaneously consider semantic relevance and personality constraint requirements in the dialogue generation process. Compared with the traditional attention mechanism that only focuses on semantics, it realizes the deep integration and precise matching of dialogue content and IP character personality. Through the joint optimization mechanism of the multi-objective personality constraint loss function, the present invention can simultaneously meet the personality constraint conditions of the IP character in the three dimensions of personality expression, language style and knowledge field, avoiding the personality deviation problem caused by a single optimization target. By adopting a layered fusion strategy and a role identity preservation mechanism, the present invention can naturally integrate educational guidance content into the conversation in a manner that is consistent with the IP character setting. Compared with the blunt insertion method of the prior art, it significantly improves the acceptance of educational content and the learning effect. Through a real-time character deviation correction mechanism and multi-level quality assessment, the present invention can continuously monitor and correct character deviations during the dialogue generation process, ensuring that the generated content always complies with the IP character setting, and effectively solves the technical problem of character identity drift in multiple rounds of dialogue. Through optimized matrix operations and attention calculation mechanisms, the present invention achieves low computational complexity and good real-time performance while ensuring the accuracy of character constraints, meeting the immediate response requirements of children's AI dialogue systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 This is a schematic diagram of the steps of a method for generating a children's AI dialogue based on an IP character setting in one embodiment of the present invention;
[0055] Figure 2 This is a structural block diagram of a children's AI dialogue generation system based on IP character settings in one embodiment of the present invention.
[0056] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0057] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0058] Reference Figure 1 This embodiment provides a method for generating children's AI dialogue based on IP character settings, including the following steps:
[0059] S1, obtain IP character personality data and construct IP character personality constraint matrix;
[0060] The input IP character personality data is parsed and standardized beforehand. This step processes the original IP character information, typically described in natural language, extracting personality descriptions, language style characteristics, and areas of expertise from the free text and converting it into a structured dataset that facilitates subsequent quantitative modeling. Natural language processing techniques are used to segment the text, perform part-of-speech tagging, and identify entities, extracting keywords describing the personality, and using semantic analysis to determine the specific representational strength of each personality trait. For language style characteristics, the uniqueness of the character's expression is extracted by analyzing sentence length, emotional word frequency, and sentence complexity indicators. For areas of expertise, keyword matching and knowledge graph linking techniques are used to identify the character's field labels, such as science, history, and art, to form a structured personality dataset containing the character's personality descriptions, language style characteristics, and areas of expertise. Based on the personality feature parameter set, the character attribute information is subjected to feature extraction and numerical quantification mapping. To achieve a quantifiable representation of personality traits, a word embedding model is used to map textual personality descriptions into quantitative features. Incorporating the five dimensions of personality in psychology, character characterization is quantified into a set of multidimensional personality parameters, including bravery, gentleness, sense of humor, responsibility, and extroversion. Each value is controlled within the [0, 1] range with an accuracy of 0.01. Similarly, for language style, statistical indicators such as sentence complexity, proportion of sentiment words, and technical term density are normalized to map them into a set of style feature parameters. For professional knowledge, multi-label binary encoding or dense vector embedding methods are used to ensure that knowledge constraint parameters accurately cover the character's knowledge boundaries and domain characteristics. Through feature extraction and numerical quantification, a character feature parameter set is generated, consisting of personality constraint parameters, style constraint parameters, and knowledge constraint parameters. Based on this character feature parameter set, a multidimensional matrix structure is organized and dimensionally allocated, constructing a multidimensional constraint matrix framework with personality constraint layers, style constraint layers, and knowledge constraint layers. The personality constraint layer is divided into 64 dimensions, covering a rich array of psychological personality indicators; the style constraint layer is divided into 128 dimensions to capture complex differences in linguistic expression; and the knowledge constraint layer is divided into 256 dimensions to depict the character's knowledge structure across multiple domains. These three layers maintain orthogonality, ensuring the independence of personality, style, and knowledge in mathematical space, preventing information coupling from blurring character definitions. The matrix framework utilizes a dense matrix storage structure, supporting parallel processing and rapid indexing of large-scale IP character data. Matrix element filling and indexing are performed based on the multi-dimensional constraint matrix framework to generate the IP character constraint matrix. Matrix element filling involves applying previously extracted and quantified personality characteristic parameters to pre-defined matrix locations, corresponding to the three categories of personality, style, and knowledge. Each element is represented as a floating-point number with a precision of 0.01, ensuring the accuracy and stability of the matrix data.At the same time, an index system is constructed based on the structural characteristics of matrix elements, and a hash table combined with an inverted index method is used to achieve rapid retrieval and dynamic update support for specific role features, thereby improving the retrieval and matching efficiency in the subsequent dialogue generation process.
[0061] S2, performing triple decoupling on the IP character constraint matrix to obtain a personality feature sub-vector, a style feature sub-vector, and a knowledge feature sub-vector;
[0062] Specifically, dimensional analysis and subspace target determination are performed based on the IP character personality constraint matrix. Through eigenvalue decomposition or principal component analysis technology, the contribution of each dimensional feature to the overall data structure is evaluated, and combined with the personality modeling requirements, the decoupling target parameter set of the three subspaces is determined. The target dimension of the personality subspace is set to 64 dimensions, aiming to capture the diversity of the character's psychological characteristics; the target dimension of the style subspace is set to 128 dimensions, used to cover language style features such as sentence structure, intonation, and emotional color; the target dimension of the knowledge subspace is set to 256 dimensions to depict the various knowledge areas mastered by the character. According to the decoupling target parameter set, the IP character personality constraint matrix is subjected to non-negative matrix decomposition. As a feature learning method, non-negative matrix decomposition can effectively achieve feature decoupling and noise reduction by decomposing the original matrix into the product of two low-rank non-negative matrices. In the specific operation, regularization terms were introduced to limit the sparsity and interpretability of the decomposition results. Iterative optimization was performed using gradient descent with a learning rate of 0.001 and a maximum number of iterations of 10,000. Termination was determined when the decomposition error was less than 5%. Three sub-matrices were obtained: the personality subspace matrix, the style subspace matrix, and the knowledge subspace matrix. Each sub-matrix corresponds to the feature mapping of the IP character in terms of personality, language style, and knowledge domain, and the elements are kept non-negative to facilitate subsequent physical interpretation and modeling. After obtaining the initial decomposed subspace matrices, orthogonality-constrained optimization and linear independence verification were performed on the personality, style, and knowledge subspace matrices to enhance the independence and expressiveness of the features in each subspace. An orthogonality loss term was introduced during the decomposition process, and Lagrange multipliers were used to constrain the sub-matrices to be as orthogonal as possible in the vector space, reducing overlap and interference between the subspaces. The covariance matrix between the sub-matrices was calculated, and linear independence criteria (such as rank test and singular value decomposition) were used to verify that the three subspaces were mathematically independent. The Adam optimizer was used during optimization with a learning rate of 2e-5 and a batch size of 32. Early stopping was also used to prevent overfitting. This yielded a set of decoupled subspace matrices that satisfied orthogonality constraints and exhibited good linear independence. Furthermore, the decoupled subspace matrices were subjected to feature vector extraction and dimensionality reduction to obtain personality, style, and knowledge subvectors. Principal component analysis or t-SNE dimensionality reduction was applied to each subspace matrix to extract the most representative low-dimensional vectors. The personality subvectors were maintained at 64 dimensions to capture the character's personality traits; the style subvectors were compressed to 128 dimensions to reflect the language's stylistic characteristics; and the knowledge subvectors were maintained at 256 dimensions to ensure both breadth and depth of knowledge coverage.
[0063] The decomposition parameters are initialized based on the decoupling target parameter set, providing appropriate hyperparameter configuration for the subsequent decomposition algorithm and ensuring the convergence speed and stability of the matrix decomposition process. Based on the dimensionality of the personality, style, and knowledge subspaces preset in the decoupling target parameter set, a gradient descent learning rate of 0.001 is set to ensure that each parameter update is moderate, neither causing excessive oscillation nor falling into local minima. The iteration threshold is set to 10,000 to ensure a sufficient search step length within the high-dimensional matrix space to achieve the optimal decomposition. A convergence error threshold is also set to 5%. The decomposition process is considered converged when the loss function decreases below this threshold, thus avoiding wasted computational resources caused by invalid iterations. Based on the decomposition initialization parameter set, the IP character constraint matrix is decomposed iteratively using gradient descent. Non-negative matrix factorization is a constrained optimization problem that requires minimizing the reconstruction error between the product of the original matrix and the decomposition matrix, assuming that all matrix elements are non-negative. A gradient descent method based on a multiplication update rule is used to iteratively update the coefficient matrix and the basis matrix. After each update, elements less than 0 are truncated to 0 to ensure matrix non-negativity. The Frobenius norm error between the current decomposition result and the original matrix is calculated at each iteration, and this error is used to guide the next update. To prevent gradient explosion or vanishing, a gradient clipping strategy is used during the update process, limiting the maximum gradient amplitude to less than 10. Furthermore, a momentum term is introduced to improve the convergence speed of the decomposition process and reduce oscillation. Through continuous gradient descent iterations, the optimal decomposition state is gradually approached, resulting in a set of intermediate decomposition result matrices. Dimension allocation and subspace extraction are performed based on this set of intermediate decomposition result matrices, resulting in personality subspace matrices, style subspace matrices, and knowledge subspace matrices. Based on the subspace dimensionality criteria preset in the decoupling target parameter set, the matrix is partitioned according to its column vectors: the first 64-dimensional vectors are assigned to the personality subspace, the 128-dimensional vectors are assigned to the style subspace, and the 256-dimensional vectors are assigned to the knowledge subspace. This dimensionality-based partitioning ensures that the eigenvectors contained in each subspace matrix are non-overlapping while covering all information dimensions of the original personality matrix. To enhance the separability and independence of the subspaces, a one-time orthogonalization process is introduced after subspace extraction, employing Gram-Schmidt orthogonalization or QR decomposition. This ensures that the subspace matrices remain internally orthogonal and linearly independent of each other, minimizing feature redundancy and overlap. The resulting personality subspace matrix captures fine-grained differences in the character's psychological traits, the style subspace matrix effectively captures the character's unique style of expression, and the knowledge subspace matrix characterizes the distribution of the character's knowledge system and domain characteristics.
[0064] S3, performing personality weighted attention calculation on the user input text based on the personality feature sub-vector, the style feature sub-vector, and the knowledge feature sub-vector to obtain an attention weight distribution matrix;
[0065] It should be noted that the user input text is subjected to word embedding encoding and sequence vectorization. For natural language text input, word vector embedding technology is used to map each word in the input text to a high-dimensional continuous vector space. Encoding methods include pre-trained Word2Vec, GloVe models, or the more context-aware BERT encoder. This embedding method can capture the semantic features of words and contextual relationships. After word embedding is completed, the word vectors in the text are arranged in sequence to form an input text vector sequence through sequence modeling. This allows each word in the sequence to not only contain its own semantic representation, but also implicitly retain its position and order information in the context. Based on the personality feature sub-vector, style feature sub-vector, and knowledge feature sub-vector, a persona query matrix and a persona key matrix are constructed and initialized to obtain a multi-dimensional persona attention calculation matrix group. Based on the personality trait subvectors, a personality query matrix and a personality key matrix are generated. These subvectors are linearly transformed to the same dimensional space as the input text vector. Similarly, based on the style and knowledge trait subvectors, a style personality matrix and a knowledge personality matrix are generated, respectively, to independently capture different aspects of personality characteristics. To maintain numerical stability and enhance computational efficiency, all matrices are initialized using Xavier initialization or LayerNorm normalization to ensure a balanced and unbiased matrix distribution, thus avoiding exploding or vanishing gradients in the early stages of training. This matrix grouping enables encoding personality features as structured inputs that can participate in attention computation. Multi-head attention weights and semantic attention weights are calculated based on the personality attention matrix group and the input text vector sequence. The multi-head attention mechanism concurrently computes attention weights across different subspaces. Specifically, 12 attention heads are configured, each focusing on capturing the matching relationship between the personality, style, or knowledge subvector and the input text along different feature dimensions. Using a scaled dot-product attention mechanism, the correlation scores between the input text vector and the persona query matrix and key matrix are calculated, and then softmax normalized to obtain the persona attention weight matrix. Simultaneously, a traditional semantic attention calculation method is used to calculate the semantic relevance within the text based on the input text vector's own query, key, and value matrices, resulting in a semantic attention weight matrix. The persona attention weight matrix and the semantic attention weight matrix are gated and fused to obtain an attention weight distribution matrix. The gating mechanism introduces a learnable fusion gating parameter to perform a weighted sum of the persona attention weight and the semantic attention weight. The gating weights are dynamically adjusted, initially set to 0.6 for persona and 0.4 for semantics, and are continuously optimized through backpropagation during training to adapt to the changing demands of different conversation turns and input content.Gated fusion can maintain the necessary semantic coherence while keeping the dominant characteristics of the character, and can adaptively adjust according to the degree of deviation from the character, thereby enhancing the balance between personalization and naturalness in the generated dialogue.
[0066] S4, inputting the attention weight distribution matrix into a multi-objective character constraint loss function for joint optimization to obtain the intermediate dialogue content that meets the IP character character constraint conditions;
[0067] Specifically, based on the attention weight distribution matrix, a multi-objective character constraint loss function is constructed, which includes personality constraint loss components, style constraint loss components, and knowledge constraint loss components. For personality constraint, the loss function uses the cosine similarity of sentiment vectors as a metric, calculating the angle between the generated text in the sentiment space and the IP character's personality vector to measure the degree of emotional alignment. For style constraint, the generated text's linguistic style features are extracted based on an n-gram language model, and the distribution differences in grammatical structure and emotional tone between the generated text and the target style samples are quantified using KL divergence. For knowledge constraint, knowledge graph embedding and entity recognition techniques are used to extract knowledge entities in the text and calculate the accuracy of entity links between them and the target knowledge domain, ensuring that the generated content does not deviate from the character's pre-set knowledge boundaries. In this way, a multi-objective character constraint loss function is comprehensively constructed, covering the three dimensions of sentiment, linguistic style, and knowledge domain. Based on this multi-objective character constraint loss function, the constraint deviation degree is calculated for the generated text's sentiment vector, linguistic style features, and knowledge domain information. For personality constraints, the cosine similarity between the generated text's sentiment vector and the IP character's personality vector is calculated to generate a personality constraint loss. A smaller value indicates a closer match between the sentiment and the pre-defined character. For style constraints, the KL divergence between the generated text's n-gram distribution and the target style sample distribution is calculated to generate a style constraint loss, reflecting the degree of deviation in linguistic expression. For knowledge domain constraints, the accuracy of entity linking is calculated by extracting text entities and comparing them with the knowledge graph, thereby generating a knowledge constraint loss. This value reflects the accuracy and consistency of the generated text within the scope of professional knowledge. A weighted linear combination of the personality, style, and knowledge constraint losses is performed to generate a comprehensive optimization loss. The specific weighting coefficients are set to 0.4 for personality loss, 0.3 for style loss, and 0.3 for knowledge loss. To optimize the generation process, the AdamW optimizer is used for iteration. The AdamW optimizer features adaptive learning rate adjustment and introduces a weight decay mechanism to effectively prevent overfitting and improve the generalization of the generated text. During the optimization process, a learning rate of 2e-5, a batch size of 32, and a gradient clipping threshold of 1.0 were set to ensure that each parameter update was both efficient and stable, enabling rapid convergence to the optimal solution that satisfied multiple objective constraints. The comprehensive optimization loss value was input into the dialogue generation decoder for gradient backpropagation and constraint-guided generation. Based on the Transformer architecture, the decoder incorporates a character-based attention mechanism during each decoding layer. By using the gradient signal fed back by the comprehensive optimization loss value, the generation strategy is adjusted so that each generated token is simultaneously guided by the triple character constraints of personality, style, and knowledge, thereby gradually generating intermediate dialogue content that conforms to the IP character's personality settings.
[0068] S5, based on the intermediate dialogue content, educational content adaptation and role identity consistency verification are performed to obtain a pre-generated dialogue text, and real-time character deviation correction is performed on the pre-generated dialogue text to obtain a target dialogue generation result.
[0069] The intermediate dialogue content is retrieved for educational knowledge points and content matching is calculated. For the currently generated intermediate dialogue text, a combination of keyword extraction and semantic matching is used to retrieve relevant knowledge points from a pre-set educational knowledge base, which covers three categories: fire safety knowledge, emotional cognition knowledge, and social etiquette knowledge. Matching is calculated based on the semantic similarity between the content and the character's feature vectors. A similarity threshold of 0.7 is set to ensure that the selected knowledge points sufficiently match the character's personality, language style, and knowledge domain. This results in the construction of a preliminary candidate educational content library. Based on the candidate educational content library, the character identity preservation factor and the conversation history memory matrix are subjected to consistency verification and time decay. The character identity preservation factor constructs a historical conversation state matrix that records the emotion vectors and semantic feature vectors of the last ten conversations and dynamically adjusts the weight of historical content using an exponential decay function. By comparing the candidate educational content against the historical memory matrix, the consistency with the character's existing performance is assessed. Educational content that deviates significantly from the character's emotional state or language style is eliminated, and a set of educational content that meets the character identity coherence criteria is selected. Based on the selected educational content, a layered fusion process is performed: content selection filtering, expression conversion language style adjustment, and emotional coloring character emotion addition. The content selection layer further filters based on character matching, prioritizing educational knowledge points that best align with the character's personality. The expression conversion layer rewrites the selected content by sentence structure, enhances emotional tone, and replaces terminology to align it with the IP character's established language style and maintain the character's consistent expression habits. The emotional coloring layer adapts emotional vocabulary and expression tone based on the character's emotional characteristics, imbuing the educational content with the character's unique emotional coloring. This ensures that the educational information is naturally integrated while enhancing the authenticity and relatability of the character's personalized expression. Through multi-layered fusion processing, character-based educational content is generated. The character-based educational content is then integrated with the intermediate dialogue content and semantic coherence is checked. Based on the Transformer encoder-decoder architecture, the two components are embedded in a unified semantic space. Using semantic similarity and contextual coherence scoring, transitional sentences are optimized to avoid awkward content splicing or logical jumps, ensuring the overall dialogue's natural flow and semantic consistency, thus forming a preliminary pre-generated dialogue text. To ensure that the generated content meets the IP character design requirements, pre-generated dialogue text undergoes real-time character deviation correction. A character consistency detection module scores the text for deviation based on personality, style, and knowledge. If the deviation threshold exceeds 0.2, a reconstruction mechanism is automatically triggered, adjusting the generation strategy and regenerating or correcting the relevant paragraphs. Furthermore, a real-time feedback mechanism dynamically adjusts the deviation tolerance threshold based on user interaction data to enhance the personalized adaptability of dialogue generation. Through this real-time character deviation correction strategy, the target dialogue generation result is ultimately output.
[0070] Pre-generated dialogue text is fed into the candidate generation layer, where it is generated using a beam search algorithm combined with a diversified expansion strategy. This yields a candidate dialogue text set containing multiple candidate responses. Beam search retains multiple sequences with the highest probability at each generation step. By setting a beam width (such as 5 or 10) and introducing randomness or temperature adjustment strategies during the search process to enhance diversity, the generated results ensure not only coverage of potential high-quality responses but also avoid pattern collapse, resulting in a rich and diverse selection of candidate dialogue texts during the initial generation phase. Based on the candidate dialogue text set, personality consistency scores, style matching scores, and knowledge accuracy scores are calculated for each candidate response. The personality consistency score is calculated by performing a cosine similarity calculation between the candidate text's sentiment vector and the target IP character's personality trait vector. A higher score indicates a closer fit between the emotional expression and the character's setting. The style match score extracts sentence structure features, sentiment intensity, and language style indicators from the candidate text and evaluates the distribution fit with the character's language style vector, using KL divergence or Wasserstein distance as evaluation metrics. The knowledge accuracy score, based on knowledge graph entity recognition, checks whether the professional terms and knowledge points mentioned in the text fall within the character's pre-defined knowledge domain, using entity linking accuracy as a metric. This series of evaluations determines the degree of conformity of each candidate response to the multi-dimensional characteristics of the character's setting, and calculates the corresponding deviation detection value to quantify the degree of deviation from the target character's setting. The deviation detection value is then compared to the character's deviation threshold, and the detection results are used to adaptively adjust the correction strength. The default deviation threshold is set to 0.2. When the candidate response's deviation is below this threshold, it is considered to meet the character's setting requirements. If the deviation exceeds the threshold, the correction strength is dynamically adjusted based on the degree of deviation, controlling the extent to which the response is regenerated or slightly modified. If a majority of candidate responses deviate significantly, a regeneration mechanism is triggered, restarting the generator to construct the response. If some candidate responses meet the requirements, a set of candidate responses that meet the character constraints is selected. The correction strength adjustment not only relies on the current detected deviation value but also takes into account conversational turn length, contextual coherence, and user feedback history, forming a dynamic adjustment mechanism to ensure an adaptive balance between generation quality and character consistency in different interaction environments. The selected candidate response set is input into the character fidelity assessment algorithm for quality evaluation and optimal response selection. The character fidelity assessment algorithm combines three indicators: personality consistency, style match, and knowledge accuracy, calculating an overall fidelity score by weighted average (e.g., each weighing 1 / 3). The final selected response is required to achieve a fidelity score of 0.85 or above, ensuring that the final output content meets the IP character's character settings across multiple dimensions. For multiple candidate responses that meet the requirements, a fine-grained comparison is performed based on contextual coherence, language fluency, and emotional naturalness, and the optimal response is selected as the final target dialogue generation result.
[0071] In one example, obtaining IP character personality data and constructing an IP character personality constraint matrix includes:
[0072] Parse and standardize the input IP character data to obtain a structured character data set containing character personality descriptions, language style characteristics, and professional knowledge areas;
[0073] Performing feature extraction and numerical quantification mapping on the character attribute information according to the structured character data set to obtain a character feature parameter set including personality constraint parameters, style constraint parameters, and knowledge constraint parameters;
[0074] Performing multi-dimensional matrix structure organization and dimension allocation based on the personality feature parameter set to obtain a multi-dimensional constraint matrix framework having a personality constraint layer, a style constraint layer, and a knowledge constraint layer;
[0075] Based on the multi-dimensional constraint matrix framework, matrix element filling and index construction are performed to obtain the IP character personality constraint matrix.
[0076] In this example, the input IP character personality data is parsed and pre-processed for standardization. The input personality data exists in the form of free text or semi-structured documents, and contains the character's basic settings, personality descriptions, language characteristics, and professional background information. Natural language processing technology is applied to the input text for word segmentation. Context-based pre-trained models (such as BERT or RoBERTa) are used for semantic segmentation and keyword extraction. Key personality descriptors, typical language style words, and domain knowledge entities in natural language are identified. Based on this, part-of-speech tagging and entity naming recognition are performed to ensure the semantic accuracy and completeness of the extracted information. Through this processing step, the raw text data is standardized into a structured personality dataset containing personality description fragments, language style fragments, and knowledge domain labels. Based on this structured personality dataset, feature extraction and numerical quantification mapping of the character attribute information are performed. For personality descriptions, we used the five personality trait models from psychology: openness, conscientiousness, extroversion, agreeableness, and emotional stability. Personality keywords extracted from the text were matched to a mapping table of the five traits. A weighted score corresponding to each trait was calculated based on a sentiment lexicon, ensuring that personality traits were mathematically quantifiable. The scores were normalized to the range [0, 1] with an accuracy of 0.01. Regarding language style, we used a language style analysis tool to statistically analyze the text for features such as sentence complexity, emotionally charged word frequency, and technical term density. The results were normalized and mapped to form a style feature vector, covering key dimensions such as sentence structure, emotional tendencies, and the use of professional terminology. For knowledge domain information, we used knowledge graph-based entity linking to identify the character's domain (e.g., science, art, literature, etc.). The identified results were encoded into a multi-label binary vector, or a dense knowledge vector was generated using knowledge graph embedding to ensure accurate modeling of the character's knowledge boundaries. These steps resulted in a set of personality feature parameters, including personality constraint parameters, style constraint parameters, and knowledge constraint parameters. Based on the character feature parameter set, a multi-dimensional matrix structure is organized and dimensionally allocated to construct a multi-dimensional constraint matrix framework with personality constraint layers, style constraint layers, and knowledge constraint layers. For personality feature parameters, 64 dimensions are set to describe the character's various psychological trait indicators in a fine-grained manner, covering specific sub-items such as bravery, curiosity, responsibility, and sense of humor; for style feature parameters, 128 dimensions are set to describe language style characteristics such as sentence complexity, emotional intensity, and the frequency of professional terms; for knowledge feature parameters, 256 dimensions are set to ensure that the knowledge field is widely and meticulously covered. The entire matrix framework uses dense matrix storage, storing each parameter value as floating-point numbers, with numerical precision maintained at two decimal places to ensure the delicacy and stability of parameter representation.The three matrices are organized in a vertically stacked fashion to ensure spatial independence for personality, style, and knowledge, preventing cross-contamination and providing mathematical support for subsequent decoupling and decomposition operations. Column-major storage is used within the matrix to improve retrieval and computational efficiency in high-dimensional spaces, ensuring compatibility with the performance requirements of parallel loading and processing of large-scale IP character libraries. After establishing a multi-dimensional constraint matrix framework, matrix element filling and indexing are performed within this framework to form the IP character personality constraint matrix. The matrix element filling process is based on a set of feature parameters, with each feature parameter being filled into the corresponding matrix dimension in a predetermined order. Personality parameters are filled into the first 64 dimensions, style parameters into the next 128 dimensions, and knowledge parameters into the remaining 256 dimensions. To ensure accuracy and consistency in the filling process, an automated mapping table mechanism is implemented to dynamically map parameter values to target matrix locations based on the correspondence between parameter names and matrix dimension indices. Outlier detection and value verification are also performed during the filling process to eliminate missing or abnormal data and improve matrix data quality. In terms of index construction, a hash index combined with an inverted index structure is used to bind the matrix position of each role to the role ID, supporting fast retrieval and reverse query based on the role ID; at the same time, in order to improve concurrent processing performance, a distributed index management module is introduced to perform matrix index partitioning and synchronous update in a multi-node environment, ensuring that the system can still maintain efficient response in large-scale concurrent request scenarios.
[0077] In one example, the triple decoupling of the IP character constraint matrix to obtain a personality feature sub-vector, a style feature sub-vector, and a knowledge feature sub-vector includes:
[0078] Perform dimensional analysis and subspace target determination based on the IP character personality constraint matrix to obtain a decoupled target parameter set;
[0079] Performing non-negative matrix decomposition on the IP character constraint matrix according to the decoupling target parameter set to obtain a personality subspace matrix, a style subspace matrix, and a knowledge subspace matrix;
[0080] Perform orthogonal constraint optimization and linear independence check based on the personality subspace matrix, the style subspace matrix, and the knowledge subspace matrix to obtain a decoupled subspace matrix set;
[0081] Feature vector extraction and vectorized dimension reduction processing are performed on the decoupled subspace matrix set to obtain personality feature subvectors, style feature subvectors, and knowledge feature subvectors.
[0082] In this example, dimensional analysis and subspace target determination are performed based on the IP character constraint matrix. As a multi-dimensional, high-order feature carrier, the IP character constraint matrix comprises three primary feature clusters: personality, language style, and knowledge domain. Each feature cluster contains several sub-feature dimensions. Feature distribution analysis and singular value decomposition (SVD) of the matrix are performed using principal component analysis (PCA) or eigenvalue decomposition (Eigenvalue Decomposition) to extract the feature contribution of each dimension. The cumulative contribution rate curve determines the appropriate retention ratio for feature dimensions, set at above 90%, ensuring sufficient feature representation while compressing redundant information. Based on experience and actual application requirements, the target dimensionality for the personality subspace is set at 64, the target dimensionality for the style subspace at 128, and the target dimensionality for the knowledge subspace at 256. These three together form the decoupled target parameter set, defining the dimensionality partitioning criteria and subspace structure layout for subsequent non-negative matrix factorization. Non-negative matrix factorization is performed on the IP character constraint matrix based on the decoupled target parameter set, achieving low-rank decomposition of the original matrix and independent extraction of subspace information. Non-negative matrix factorization represents the original matrix as the product of two low-rank non-negative matrices, one corresponding to the subspace basis matrix and the other to the coefficient matrix. The basis matrix encodes the eigenvectors of each subspace, while the coefficient matrix represents the weight distribution of each eigenvector. The decomposition process utilizes a multiplication update rule based on gradient descent, with an initial learning rate of 0.001 and a maximum number of iterations of 10,000. Regularization constraints are introduced during the update process to avoid abnormalities such as excessively sparse or dense matrix elements. The reconstruction error after decomposition is calculated at each iteration, and the objective function is minimized using the Frobenius norm. The iteration is terminated when the error falls below 5%. This results in preliminary character subspace matrices, style subspace matrices, and knowledge subspace matrices. These three subspace matrices represent the low-dimensional features of the IP character across the three dimensions of personality. Orthogonality-constrained optimization and linear independence checks are performed on the character, style, and knowledge subspace matrices to eliminate potential redundancy and correlation between the subspaces and enhance the independent expressive power of each subspace. Orthogonality constraints are introduced to minimize the correlation index between subspace matrices, forcing the basis vectors of each subspace to be nearly orthogonal in Euclidean space. Orthogonal projection matrices and Lagrange multipliers are used to introduce orthogonality conditions into the matrix update process, ensuring the orthogonality of the submatrix column vector group after each gradient descent update. After orthogonalization is completed, a linear independence check is performed, using methods such as rank test and singular value decomposition to verify whether the submatrix column vector group has full rank, thereby ensuring that there is no linear dependency between the eigenvectors within each subspace. Through orthogonal optimization and linear independence verification, a highly independent, orthogonal, and non-redundant set of decoupled subspace matrices is obtained. Eigenvector extraction and vectorized dimensionality reduction are performed on the decoupled subspace matrix set.For each subspace matrix, principal component analysis (PCA) is used for further dimensionality reduction. The top eigenvectors with the highest principal component contributions are selected. The character subvector is compressed to 64 dimensions, the style subvector to 128 dimensions, and the knowledge subvector to 256 dimensions, ensuring that the essential information of the subspace features is preserved to the greatest extent possible while reducing the dimensionality. The reduced eigenvectors are normalized to the interval [0, 1] to ensure scale consistency between vectors and avoid numerical instability caused by scale discrepancies in subsequent attention weight calculation and generative model training. To enhance the semantic expressiveness of the vectors, embedding enhancement is performed on the eigenvectors. Deep autoencoders are then used to reconstruct the vectors, ensuring that they retain a low-dimensional and compact representation while enhancing their ability to capture nonlinear features and generalize. Through these steps, the character, style, and knowledge subvectors are ultimately obtained.
[0083] In one example, the non-negative matrix decomposition is performed on the IP character personality constraint matrix according to the decoupling target parameter set to obtain a personality subspace matrix, a style subspace matrix, and a knowledge subspace matrix, including:
[0084] Initializing the decomposition parameters based on the decoupling target parameter set to obtain a decomposition initialization parameter set including a gradient descent learning rate, an iteration number threshold, and a convergence error threshold;
[0085] Performing gradient descent iterative decomposition on the IP character personality constraint matrix according to the decomposition initialization parameter set to obtain a decomposition intermediate result matrix set;
[0086] Dimension allocation and subspace extraction are performed based on the decomposition intermediate result matrix set to obtain a personality subspace matrix, a style subspace matrix and a knowledge subspace matrix.
[0087] In this example, key control parameters in the decomposition process are first initialized based on the decoupling objective parameter set to construct the decomposition initialization parameter set. Considering that non-negative matrix factorization is essentially an optimization problem under non-negativity constraints, it requires continuous iteration to minimize the matrix reconstruction error. Therefore, appropriate gradient descent learning rates, maximum iteration thresholds, and convergence error thresholds are set in advance. The learning rate, as a key hyperparameter controlling the step size of each parameter update, directly impacts the convergence speed and stability of the decomposition. To avoid oscillation caused by excessively large step sizes or slow convergence caused by excessively small step sizes, the initial learning rate is set within a relatively robust range of 0.001. An adaptive learning rate adjustment mechanism is introduced in subsequent iterations to dynamically fine-tune the learning rate based on the downward trend in the loss. Furthermore, the iteration threshold is set to 10,000 to ensure ample optimization space for the decomposition algorithm while effectively preventing wasted iterations due to vanishing gradients or local minima. The convergence error threshold is set to 5%. When the matrix reconstruction error in the current iteration falls within 5% of the original matrix norm, the decomposition process is considered converged and the iteration process is automatically terminated. The above parameter settings initially establish a decomposition initialization parameter set, including the gradient descent learning rate, iteration threshold, and convergence error threshold. Based on this decomposition initialization parameter set, the IP character constraint matrix is subjected to iterative gradient descent decomposition. The basic idea of the non-negative matrix factorization algorithm is to decompose the original matrix into the product of two low-rank non-negative matrices: a basis matrix and a coefficient matrix. Through continuous iterative optimization, the product of these two matrices is optimized to be as close as possible to the original matrix, with all elements non-negative. In the actual decomposition process, a multiplication update rule is employed based on the principle of gradient descent. In each iteration, the coefficient matrix is updated with the basis matrix fixed, followed by the basis matrix fixed with the coefficient matrix fixed, alternating between these two steps to ensure the correct overall convergence direction. To ensure non-negativity, all negative elements are truncated and set to zero after the update to prevent mathematical anomalies caused by numerical fluctuations. After each iteration, the Frobenius norm error between the current reconstructed matrix and the original matrix is calculated and compared with the set convergence error threshold. If the error falls below the threshold or the maximum number of iterations is reached, the iteration is terminated and the current basis matrix and coefficient matrix are output as the intermediate decomposition result matrix set. To improve decomposition stability and convergence speed, a momentum factor and gradient clipping mechanism are introduced during the optimization process to limit the gradient update amplitude to a reasonable range (for example, setting the gradient clipping upper limit to 10) to prevent extreme gradient fluctuations from interfering with the convergence process. Through this iterative process, the resulting decomposition intermediate result matrix not only retains the key feature information of the original IP character constraint matrix, but also effectively reduces data redundancy. After obtaining the decomposition intermediate result matrix set, dimension allocation and subspace extraction are performed based on the established decoupling target parameter set to form the personality subspace matrix, style subspace matrix, and knowledge subspace matrix.Specifically, the decoupling target parameter set clearly divides the personality dimension into 64 dimensions, the style dimension into 128 dimensions, and the knowledge dimension into 256 dimensions. Therefore, during the subspace extraction process, according to this preset standard, column vectors of corresponding dimensions are sequentially intercepted from the decomposed basis matrix. The first 64 column vectors are used as the personality subspace matrix, the next 128 column vectors are used as the style subspace matrix, and the remaining 256 column vectors are used as the knowledge subspace matrix. To ensure the numerical stability and structural rationality of the subspace matrix, a one-time normalization process is performed during the extraction process to normalize the L2 norm of each vector to within 1 to prevent vector scale differences from affecting subsequent feature decoupling and attention mechanism modeling. At the same time, to improve the expressive power and mutual differences of each subspace, an orthogonal projection process is introduced after the subspace division is completed to ensure that the personality, style, and knowledge subspaces maintain good orthogonality and linear independence in the vector space, avoiding feature cross-contamination and confusion in character expression.
[0088] In one example, the user input text is subjected to personality weight attention calculation based on the personality feature sub-vector, the style feature sub-vector, and the knowledge feature sub-vector to obtain an attention weight distribution matrix, including:
[0089] Perform word embedding encoding and sequence vectorization on the user input text to obtain the input text vector sequence;
[0090] Constructing and initializing a character query matrix and a character key matrix according to the personality feature sub-vector, the style feature sub-vector, and the knowledge feature sub-vector to obtain a multi-dimensional character attention calculation matrix group;
[0091] Perform multi-head attention weight calculation and semantic attention weight calculation based on the multi-dimensional human attention calculation matrix group and the input text vector sequence to obtain a human attention weight matrix and a semantic attention weight matrix;
[0092] The human-setting attention weight matrix and the semantic attention weight matrix are gated and fused to obtain an attention weight distribution matrix.
[0093] In this example, word embedding encoding and sequence vectorization are performed on user input text. The input sentence is segmented using a tokenizer such as the BERT Tokenizer, WordPiece, or SentencePiece, breaking the text into words or subword units to balance vocabulary size with fine-grained linguistic expression. After segmentation, each word is mapped into a high-dimensional semantic space using a pre-trained word embedding model. Embedding methods include Word2Vec, GloVe, or the Transformer-based BERT and RoBERTa models. This embedding approach ensures that each word not only retains semantic information but also includes contextual relationships with other words. The word embedding dimension is set to 768 to ensure rich and detailed semantic expression. The word vectors are concatenated according to the natural order of the text to form a sequence of input text vectors. To improve the model's ability to perceive text structure, positional encoding is introduced. The position of each word is added to the word vector using a specific encoding method. This allows the model to capture the word's position within the sentence while recognizing its meaning. Based on the personality sub-vectors, style sub-vectors, and knowledge sub-vectors, a character query matrix and a character key matrix are constructed and initialized, forming a multi-dimensional character attention calculation matrix set. The personality sub-vectors (64 dimensions), style sub-vectors (128 dimensions), and knowledge sub-vectors (256 dimensions) are each projected into a 768-dimensional vector space via linear transformations to maintain dimensionality consistency with the input text vector sequence, facilitating subsequent matrix operations. This transformation is implemented using independent fully connected networks, while each sub-vector retains its own unique characteristics. These matrices are initialized using Xavier or He initialization strategies to ensure uniform and stable weight distribution, helping to maintain proper gradient propagation during training and preventing vanishing or exploding gradients. To enhance the relevance of personality features to text content, the sub-vectors are initially processed in interaction with the input text vector to generate preliminary attention cues. These signals are continuously adjusted during training through end-to-end learning, thereby improving the fusion of personality and text features. This ensures that the generation process more accurately reflects the character's personality, language style, and knowledge domain characteristics. Multi-head attention weights and semantic attention weights are calculated based on the multi-dimensional character attention calculation matrix group and the input text vector sequence. The multi-head attention mechanism captures the deep semantic relationships in the input text from different feature perspectives by calculating the attention weights of multiple subspaces in parallel. Specifically, standard semantic attention is calculated based on the input text vector sequence, and the semantic associations between words are captured through the self-attention method. At the same time, the character query matrix and key matrix are used to calculate the attention weights driven by the character characteristics, strengthening the model's perception of the character setting. The multi-head attention mechanism uses 12 heads, each of which focuses on a different character dimension or semantic feature. Finally, the output results of each head are spliced and linearly transformed to generate a unified multi-dimensional attention feature representation.After calculating the two types of attention weights, the persona and semantic attention weight matrices are gated and fused to produce an attention weight distribution matrix. This gated fusion mechanism is an effective information integration method that dynamically balances the weights of persona and semantic features in the final generation by introducing a learnable gating vector. The gating vector is generated using a sigmoid activation function, with the output value ranging from 0 to 1. This controls the balance of persona and semantic attention weights. Initially, the ratio of persona to semantic attention weights is set at 6:4. This ratio is automatically adjusted during training based on feedback from model learning to adapt to changing conversational scenarios and user input requirements. The fusion process uses a weighted summation method, dynamically weighting persona and semantic attention according to the gating coefficients to produce a unified attention weight distribution matrix. To prevent loss of fused information or noise interference, residual connections and normalization are introduced during the fusion process to enhance generation stability and training efficiency.
[0094] In one example, the inputting of the attention weight distribution matrix into the multi-objective character constraint loss function for joint optimization to obtain the intermediate dialogue content that satisfies the IP character character constraint conditions includes:
[0095] Constructing a multi-objective personality constraint loss function based on the attention weight distribution matrix, which includes a personality constraint loss component, a style constraint loss component, and a knowledge constraint loss component;
[0096] Calculating the constraint deviation of the generated text's sentiment vector, language style features, and knowledge domain information based on the multi-objective character constraint loss function to obtain a character constraint loss value, a style constraint loss value, and a knowledge constraint loss value;
[0097] Performing weighted linear combination and AdamW optimizer iterative processing based on the personality constraint loss value, the style constraint loss value, and the knowledge constraint loss value to obtain a comprehensive optimization loss value;
[0098] The comprehensive optimization loss value is input into the dialogue generation decoder for gradient backpropagation and constraint guidance generation to obtain the intermediate dialogue content that meets the IP character setting constraints.
[0099] In this example, a multi-objective character constraint loss function is constructed based on the attention weight distribution matrix, covering three dimensions: personality, language style, and knowledge domain. For personality constraints, the loss function extracts the sentiment vector of the generated text and evaluates its similarity with the personality trait vector of the target IP character to assess whether the generated text's emotional tendencies match the character's setting. Personality traits primarily examine 64 fine-grained characteristics such as bravery, gentleness, and sense of humor. The sentiment vector extracts information such as sentiment polarity, intensity, and emotion category from the generated text to ensure comprehensive and accurate matching. For style constraints, the loss function evaluates the matching between the generated text's language style feature vector and the character's language style subvector. Language style features include 128 linguistic characteristic indicators such as sentence complexity, emotional intensity, and frequency of professional terminology, capturing the text's unique expressive style. For knowledge constraints, the loss function identifies entities and concepts mentioned in the generated text and maps these entities to the character's pre-defined knowledge domain. The knowledge features cover 256 dimensions and control the boundaries of multiple professional domains to avoid generating responses that exceed the character's cognitive scope. The above components are designed to form a character constraint loss component, a style constraint loss component, and a knowledge constraint loss component. Together, these three components form a multi-objective character constraint loss function. Based on this multi-objective character constraint loss function, the generated text is subjected to constraint deviation calculations, specifically including the character constraint loss, style constraint loss, and knowledge constraint loss. The character constraint loss is calculated by comparing the degree of match between the generated text's sentiment vector and the IP character's personality trait vector. The lower the match, the higher the character constraint loss, and vice versa. This indicates the degree of consistency between the text's emotional expression and the character's setting. The style constraint loss compares the generated text's sentence complexity, emotional intensity, and the proportion of professional terminology used with the character's language style trait vector to calculate the deviation between the text's language expression and the character's setting. The knowledge constraint loss uses knowledge graph entity linking technology to analyze the overlap between entities in the text and the character's knowledge domain. A higher entity hit rate indicates a better knowledge match and a lower deviation. To ensure the meticulousness of the deviation calculation, all loss values are calculated using standardized values and unified dimensions to ensure that different characteristics are appropriately weighted during the comprehensive evaluation, preventing a single characteristic from dominating the overall optimization direction. A weighted linear combination is performed based on the personality constraint loss value, style constraint loss value, and knowledge constraint loss value to form a comprehensive optimization loss value. In terms of weight setting, the personality constraint loss weight is set to 0.4, the style constraint loss weight is set to 0.3, and the knowledge constraint loss weight is set to 0.3, reflecting that emotional consistency has a slightly higher priority than language style and knowledge accuracy in children's interaction scenarios, which meets the actual application needs of children who have a stronger need for emotional cognition of virtual characters.The weighted linear combination ensures that the comprehensive optimization loss balances the requirements of personality, style, and knowledge, without neglecting the characteristics of any one dimension. The AdamW optimizer is used for iterative updates to optimize the comprehensive optimization loss. The AdamW optimizer features adaptive learning rate adjustment and weight decay mechanisms, effectively preventing model overfitting. The initial learning rate is set to 2e-5, the batch size is 32, and the maximum number of iterations is set to 5000, ensuring good convergence with moderate computing resource consumption. At the same time, gradient clipping technology is used, setting the maximum gradient threshold to 1.0 to avoid model instability caused by excessive gradients, ensuring a smooth and controllable optimization process, and ultimately steadily reducing the comprehensive optimization loss to within the expected range. After the comprehensive optimization loss has effectively converged, it is input into the dialogue generation decoder for gradient backpropagation and constraint-guided generation. The dialogue generation decoder, based on the Transformer decoder architecture, possesses powerful language generation capabilities. By incorporating reverse gradient information from a comprehensive optimization loss, it fine-tunes the generation strategy for each token during the generation process. This allows the model to not only consider contextual semantic coherence for each generated word, but also to verify its conformity to the character setting, language style, and knowledge domain requirements in real time. Through a reverse-guided generation mechanism, the decoder dynamically adjusts attention distribution during the generation process, prioritizing vocabulary and expressions that closely match the character's characteristics. This avoids wording that deviates from the character's personality and answers that exceed knowledge boundaries, thereby improving the character consistency and educational accuracy of the generated text. This results in intermediate dialogue content that meets the constraints of the IP character's personality.
[0100] In one example, the educational content adaptation and role identity consistency verification based on the intermediate dialogue content to obtain a pre-generated dialogue text, and the real-time character deviation correction of the pre-generated dialogue text to obtain a target dialogue generation result include:
[0101] Performing educational knowledge point retrieval and content matching calculation on the intermediate conversation content to obtain a candidate educational content library containing fire safety knowledge, emotional cognition knowledge, and social etiquette knowledge;
[0102] Performing consistency check and time decay processing on the role identity preservation factor and the dialogue history memory matrix according to the candidate education content library to obtain a screening education content set that meets the role identity consistency condition;
[0103] Based on the screened educational content set, layered fusion processing is performed on the content selection layer, the language style is adjusted on the expression conversion layer, and the character emotion is added on the emotion coloring layer to obtain role-based educational fusion content;
[0104] Performing content integration and semantic coherence checking on the role-based education fusion content and the intermediate dialogue content to obtain a pre-generated dialogue text;
[0105] The pre-generated dialogue text is corrected for deviations in the human setting in real time to obtain a target dialogue generation result.
[0106] In this example, educational knowledge points are retrieved and content matching is calculated based on the content of the intermediate conversations. The intermediate conversation content undergoes keyword extraction and semantic embedding, converting it into a high-dimensional semantic vector. This vector is then matched against a pre-defined database of educational knowledge points covering multiple subject areas, including fire safety, emotional cognition, and social etiquette. Each knowledge point within the database is also pre-processed into a corresponding semantic vector. By calculating the similarity between the intermediate conversation content and each knowledge point vector and setting a matching threshold (set to 0.7), knowledge points with a matching score exceeding the threshold are selected to construct a preliminary candidate educational content library. To ensure high relevance between the educational content and the conversation context, matching calculations are based not only on static semantic similarity but also on a contextual topic inference module to analyze the contextual tendencies of the conversation. This further optimizes the relevance screening of knowledge points, ensuring that the selected candidate educational content library is both informative and logically coherent with the conversation context. Based on the candidate educational content library, the role identity preservation factor and the conversation history memory matrix are subjected to consistency verification and time decay processing. The character identity preservation factor is a key parameter that measures the consistency of a character's personality, language style, and knowledge domain across multiple rounds of dialogue. A consistency score is generated by comparing the character feature vector with the vector generated by historical dialogues. Furthermore, the dialogue history memory matrix stores textual and character features from recent rounds of dialogue. A sliding window mechanism is used to maintain the continuity of the dialogue context, and a time decay mechanism is introduced, applying an exponentially decaying weight to earlier dialogue rounds with a decay coefficient set to 0.1, giving higher weight to dialogue history closer to the current moment. Each piece of content in the candidate educational content library is matched against the historical memory matrix, and the similarity between the educational content and the character's characteristics is calculated. Educational knowledge points that significantly deviate from the current character's identity are removed, thus selecting a set of educational content that meets the character identity consistency criteria. Based on this selected set of educational content, a layered fusion process is performed, sequentially: content selection filtering, expression conversion language style adjustment, and emotion toning character emotion addition, to generate character-based educational content. The content selection layer ranks content based on its match with the character's personality, retaining the top 30% of educational content and further removing redundant or low-relevance content. The expression conversion layer adjusts the presentation of educational content based on the character's language style eigenvector. This involves controlling sentence complexity, adjusting emotional intensity, and unifying vocabulary style, ensuring that the educational content aligns with the character's original language style in terms of sentence structure, emotional tendencies, and frequency of professional terminology. The emotional toning layer, based on the character's personality eigenvector, overlays the character's unique emotional expression habits. For example, for a gentle character, gentle and caring vocabulary is added, while for a lively character, positive and encouraging expressions are added. After these three layers of processing, the educational content is highly aligned with the character's setting in terms of subject matter, expression, and emotional tone, forming a role-based educational integration.The role-based educational content is integrated with the intermediate dialogue content and semantic coherence is checked to generate pre-generated dialogue text. During the content integration phase, a context-aware integration mechanism is employed to dynamically select the insertion location for educational content based on the fluidity of the dialogue topic, avoiding abrupt content insertions or logical jumps. After integration, a semantic coherence check module utilizes a natural language inference model to check the logical consistency and semantic cohesion between the pre- and post-integration texts. If semantic discontinuities or reasoning conflicts are detected, adjustments or reorganizations are automatically made to ensure that the generated pre-generated dialogue text meets the standards of natural conversation in terms of logical structure and expressive fluency, avoiding disruptive interactions due to abrupt insertions of educational content. Pre-generated dialogue text is corrected for deviations from the character setting in real time. By introducing a character setting consistency check module, the generated text is monitored in real time for its consistency across personality, style, and knowledge domain dimensions, with evaluation performed on each generated sentence. The check metrics include personality consistency score, style matching score, and knowledge accuracy score. A deviation threshold of 0.2 is set. If the deviation of any metric exceeds the threshold, a local correction mechanism is triggered, rewriting or replacing the deviating segment to restore consistency between the text and the character setting. To improve correction efficiency and naturalness, a dynamic feedback mechanism is introduced during the correction process. This adjusts the correction intensity based on the degree of deviation, with minor adjustments for small deviations and entire sentences rewritten for larger deviations. The correction strategy is fine-tuned based on historical user interaction data, ensuring that the correction process maintains a natural conversation experience while maximally preserving the personality of the IP character. This results in the target dialogue generation result.
[0107] In one example, performing real-time character deviation correction on the pre-generated dialogue text to obtain a target dialogue generation result includes:
[0108] Inputting the pre-generated dialogue text into the candidate generation layer for beam search and diversified expansion to obtain a candidate dialogue text set containing multiple candidate responses;
[0109] Calculating the personality consistency score, style matching score, and knowledge accuracy score based on the candidate dialogue text set to obtain a personality conformity assessment result and a deviation detection value corresponding to each candidate response;
[0110] Comparing and judging the personality deviation threshold and adaptively adjusting the correction strength according to the deviation detection value to obtain a screening candidate response set that meets the personality constraint conditions or regenerate a trigger instruction;
[0111] The screened candidate response set is input into the persona fidelity assessment algorithm for quality assessment and optimal response selection to obtain the target dialogue generation result.
[0112] In this example, pre-generated dialogue text is first input into the candidate generation layer. The candidate generation layer uses beam search technology combined with a diversified expansion strategy to generate multiple candidate responses. Beam search ensures diversity and semantic plausibility in the generated results by retaining the top k highest-scoring sequences at each generation step, with k ranging from 5 to 10. At each step, based on the currently generated portion, the k next words with the highest probability are selected from the vocabulary, expanding into k different paths. This process continues until a terminator is generated or the maximum response length is reached. To avoid repetition or lack of variation in the generated results, a diversified beam search mechanism is introduced on top of the standard beam search. By introducing a penalty coefficient in each generated path, the expansion probability of repeated patterns is controlled, thereby improving the diversity of the generated results in terms of sentence structure, tone, and word choice. Furthermore, a random perturbation strategy is introduced to allow for the selection of suboptimal paths with a lower probability, further increasing the coverage and variability of the candidate responses, resulting in a candidate dialogue text set containing multiple candidate responses. Personality consistency scores, style matching scores, and knowledge accuracy scores are calculated based on the candidate dialogue text set. The personality consistency score is calculated by comparing the candidate text's sentiment vector with the pre-set IP character's personality trait vector. This examines whether the generated text aligns with the character's setting in terms of emotion polarity, intensity, and category. A higher score indicates closer emotional expression to the character's personality. The style match score measures consistency between the text's expression style and the character's setting by statistically analyzing metrics such as sentence complexity, emotional richness, and the frequency of terminology used in the text, and comparing them with the character's language style trait vector. The knowledge accuracy score uses knowledge graph entity recognition technology to extract relevant knowledge points from the text and match them with the character's pre-set knowledge domain. Entity coverage and domain adaptability are calculated to assess whether the text exceeds the character's knowledge boundaries. The calculation of these three metrics yields a personality conformance assessment for each candidate response, which is then used to generate a corresponding deviation test value. A lower deviation indicates a more consistent response with the character's setting and a more suitable candidate for the final response. After obtaining the deviation test values for each candidate response, a comparison is made based on a pre-set personality deviation threshold, and adaptive correction intensity adjustment is performed based on the test results. The persona deviation threshold is set at 0.2, meaning that a candidate response with a deviation of less than or equal to 0.2 is considered to meet the persona constraints. If the deviation exceeds the threshold, appropriate processing is required. For candidate responses whose deviation slightly exceeds the threshold, a lightweight local adjustment mechanism is used to fine-tune them, such as replacing inappropriate vocabulary, adjusting sentence complexity, or deleting entity expressions that do not match the knowledge domain to reduce the deviation. For candidate responses whose deviation far exceeds the threshold, a regeneration mechanism is triggered, returning to the beam search layer to regenerate new candidate responses.During the correction process, an adaptive correction intensity adjustment mechanism is introduced to dynamically adjust the correction amplitude based on the degree of deviation. Smaller deviations result in smaller adjustments, while larger deviations result in larger adjustments. This ensures that character deviations are minimized while maintaining natural conversational flow. This dynamic adjustment strategy improves the consistency and reliability of generated responses, maintaining the fluency and naturalness of the generated responses while also controlling the consistency and stability of the character image. After candidate responses are screened and corrected, those that meet the character constraints are fed into a character fidelity assessment algorithm for quality evaluation and optimal response selection. The character fidelity assessment algorithm comprehensively considers three dimensions: personality consistency score, style match score, and knowledge accuracy score, and performs a weighted average, giving equal weight to each dimension to ensure that the comprehensive evaluation process is not biased towards a single characteristic. The calculated fidelity score is used to rank candidate responses. A minimum passing fidelity score of 0.85 is set; responses below this score are eliminated, ensuring that the final output meets high standards for character consistency. For multiple qualified candidate responses that have passed the screening, they are compared based on auxiliary indicators such as natural language fluency, contextual coherence, and emotional naturalness, and the response with the highest score is selected as the final target dialogue generation result.
[0113] Reference Figure 2 This embodiment provides a children's AI dialogue generation system based on IP character settings, including:
[0114] Acquisition module 1, used to obtain IP character personality data and construct IP character personality constraint matrix;
[0115] Decoupling module 2, for performing triple decoupling on the IP character constraint matrix to obtain a personality feature sub-vector, a style feature sub-vector, and a knowledge feature sub-vector;
[0116] Calculation module 3, configured to perform personality weight attention calculation on the user input text based on the personality feature sub-vector, the style feature sub-vector, and the knowledge feature sub-vector to obtain an attention weight distribution matrix;
[0117] Joint optimization module 4, used to input the attention weight distribution matrix into the multi-objective character constraint loss function for joint optimization, to obtain the intermediate dialogue content that meets the IP character character constraint conditions;
[0118] The real-time correction module 5 is used to perform educational content adaptation and role identity consistency verification based on the intermediate dialogue content to obtain a pre-generated dialogue text, and perform real-time character deviation correction on the pre-generated dialogue text to obtain a target dialogue generation result.
[0119] In this embodiment, for the specific implementation of each unit in the above system embodiment, please refer to the above method embodiment, which will not be repeated here.
[0120] The technical solution provided by the present invention constructs a multi-dimensional IP character personality constraint matrix including a personality constraint layer, a style constraint layer and a knowledge constraint layer. The present invention can realize all-round and accurate quantitative modeling of the IP character personality, and significantly improves the expression integrity and constraint accuracy of the personality characteristics compared to the simple label description method of the prior art. The triple decoupling mechanism is used to decompose the personality constraint matrix into independent personality feature sub-vectors, style feature sub-vectors and knowledge feature sub-vectors, avoiding the problem of mutual interference between different personality dimensions in the prior art, and ensuring the linear independence and independent controllability of the features of each dimension. Based on the personality weight attention modulation mechanism, the present invention can simultaneously consider semantic relevance and personality constraint requirements during the dialogue generation process. Compared with the traditional attention mechanism that only focuses on semantics, it realizes the deep integration and precise matching of dialogue content and IP character personality. Through the joint optimization mechanism of the multi-objective personality constraint loss function, the present invention can simultaneously meet the personality constraint conditions of the IP character in the three dimensions of personality expression, language style and knowledge field, avoiding the personality deviation problem caused by a single optimization target. By adopting a layered fusion strategy and a role identity preservation mechanism, the present invention can naturally integrate educational guidance content into the conversation in a manner that is consistent with the IP character setting. Compared with the blunt insertion method of the prior art, it significantly improves the acceptance of educational content and the learning effect. Through a real-time character deviation correction mechanism and multi-level quality assessment, the present invention can continuously monitor and correct character deviations during the dialogue generation process, ensuring that the generated content always complies with the IP character setting, and effectively solves the technical problem of character identity drift in multiple rounds of dialogue. Through optimized matrix operations and attention calculation mechanisms, the present invention achieves low computational complexity and good real-time performance while ensuring the accuracy of character constraints, meeting the immediate response requirements of children's AI dialogue systems.
[0121] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which implements the above-described method when executed by a processor. It is understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.
[0122] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware using a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the above-described method embodiments. Any reference to memory, storage, database, or other media provided herein and used in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double-speed SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct RAMbus dynamic RAM (DRDRAM), and RAMbus dynamic RAM.
[0123] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, system, article, or method comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, system, article, or method. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, system, article, or method comprising the element.
[0124] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A method for generating children's AI dialogue based on IP character settings, characterized by: include: Obtain IP character personality data and construct IP character personality constraint matrix; Perform triple decoupling on the IP character constraint matrix to obtain a personality feature sub-vector, a style feature sub-vector, and a knowledge feature sub-vector; Performing personality weighted attention calculation on the user input text based on the personality feature sub-vector, the style feature sub-vector, and the knowledge feature sub-vector to obtain an attention weight distribution matrix; Inputting the attention weight distribution matrix into a multi-objective character constraint loss function for joint optimization to obtain intermediate dialogue content that meets the IP character character constraint conditions, including: constructing a multi-objective character constraint loss function containing a personality constraint loss component, a style constraint loss component, and a knowledge constraint loss component based on the attention weight distribution matrix; calculating the constraint deviation of the emotional vector, language style characteristics, and knowledge domain information of the generated text according to the multi-objective character constraint loss function to obtain a personality constraint loss value, a style constraint loss value, and a knowledge constraint loss value; performing weighted linear combination and AdamW optimizer iterative processing based on the personality constraint loss value, the style constraint loss value, and the knowledge constraint loss value to obtain a comprehensive optimization loss value; inputting the comprehensive optimization loss value into a dialogue generation decoder for gradient backpropagation and constraint-guided generation to obtain intermediate dialogue content that meets the IP character character character constraint conditions; Based on the intermediate dialogue content, educational content adaptation and role identity consistency verification are performed to obtain a pre-generated dialogue text, and the pre-generated dialogue text is corrected for human deviation in real time to obtain a target dialogue generation result, including: inputting the pre-generated dialogue text into the candidate generation layer for beam search and diversified expansion to obtain a candidate dialogue text set containing multiple candidate replies; calculating the personality consistency score, style matching score and knowledge accuracy score based on the candidate dialogue text set to obtain a human conformity assessment result and a deviation detection value corresponding to each candidate reply; comparing and judging the human deviation threshold and adaptively adjusting the correction strength according to the deviation detection value to obtain a screened candidate reply set that meets the human constraint conditions or a regeneration trigger instruction; inputting the screened candidate reply set into the human fidelity assessment algorithm for quality assessment and optimal reply selection to obtain a target dialogue generation result.
2. The method for generating children's AI dialogue based on IP character setting according to claim 1 is characterized in that: The step of obtaining IP character personality data and constructing an IP character personality constraint matrix includes: Parse and standardize the input IP character data to obtain a structured character data set containing character personality descriptions, language style characteristics, and professional knowledge areas; Performing feature extraction and numerical quantification mapping on the character attribute information according to the structured character data set to obtain a character feature parameter set including personality constraint parameters, style constraint parameters, and knowledge constraint parameters; Performing multi-dimensional matrix structure organization and dimension allocation based on the personality feature parameter set to obtain a multi-dimensional constraint matrix framework having a personality constraint layer, a style constraint layer, and a knowledge constraint layer; Based on the multi-dimensional constraint matrix framework, matrix element filling and index construction are performed to obtain the IP character personality constraint matrix.
3. The method for generating children's AI dialogue based on IP character setting according to claim 1 is characterized in that: The triple decoupling of the IP character constraint matrix to obtain a personality feature sub-vector, a style feature sub-vector, and a knowledge feature sub-vector includes: Perform dimensional analysis and subspace target determination based on the IP character personality constraint matrix to obtain a decoupled target parameter set; Performing non-negative matrix decomposition on the IP character constraint matrix according to the decoupling target parameter set to obtain a personality subspace matrix, a style subspace matrix, and a knowledge subspace matrix; Perform orthogonal constraint optimization and linear independence check based on the personality subspace matrix, the style subspace matrix, and the knowledge subspace matrix to obtain a decoupled subspace matrix set; Feature vector extraction and vectorized dimension reduction processing are performed on the decoupled subspace matrix set to obtain personality feature subvectors, style feature subvectors, and knowledge feature subvectors.
4. The method for generating children's AI dialogue based on IP character setting according to claim 3 is characterized in that: The non-negative matrix decomposition of the IP character personality constraint matrix according to the decoupling target parameter set is performed to obtain a personality subspace matrix, a style subspace matrix, and a knowledge subspace matrix, including: Initializing the decomposition parameters based on the decoupling target parameter set to obtain a decomposition initialization parameter set including a gradient descent learning rate, an iteration number threshold, and a convergence error threshold; Performing gradient descent iterative decomposition on the IP character personality constraint matrix according to the decomposition initialization parameter set to obtain a decomposition intermediate result matrix set; Dimension allocation and subspace extraction are performed based on the decomposition intermediate result matrix set to obtain a personality subspace matrix, a style subspace matrix and a knowledge subspace matrix.
5. The method for generating children's AI dialogue based on IP character setting according to claim 1 is characterized in that: The performing of personality weighted attention calculation on the user input text based on the personality feature sub-vector, the style feature sub-vector, and the knowledge feature sub-vector to obtain an attention weight distribution matrix includes: Perform word embedding encoding and sequence vectorization on the user input text to obtain the input text vector sequence; Constructing and initializing a character query matrix and a character key matrix according to the personality feature sub-vector, the style feature sub-vector, and the knowledge feature sub-vector to obtain a multi-dimensional character attention calculation matrix group; Perform multi-head attention weight calculation and semantic attention weight calculation based on the multi-dimensional human attention calculation matrix group and the input text vector sequence to obtain a human attention weight matrix and a semantic attention weight matrix; The human-setting attention weight matrix and the semantic attention weight matrix are gated and fused to obtain an attention weight distribution matrix.
6. The method for generating children's AI dialogue based on IP character setting according to claim 1, characterized in that: The educational content adaptation and role identity consistency verification based on the intermediate dialogue content are performed to obtain a pre-generated dialogue text, and the pre-generated dialogue text is corrected for deviations in the character setting in real time to obtain a target dialogue generation result, including: Performing educational knowledge point retrieval and content matching calculation on the intermediate conversation content to obtain a candidate educational content library containing fire safety knowledge, emotional cognition knowledge, and social etiquette knowledge; Performing consistency check and time decay processing on the role identity preservation factor and the dialogue history memory matrix according to the candidate education content library to obtain a screening education content set that meets the role identity consistency condition; Based on the screened educational content set, layered fusion processing is performed on the content selection layer, the language style is adjusted on the expression conversion layer, and the character emotion is added on the emotion coloring layer to obtain role-based educational fusion content; Performing content integration and semantic coherence checking on the role-based education fusion content and the intermediate dialogue content to obtain a pre-generated dialogue text; The pre-generated dialogue text is corrected for deviations in the human setting in real time to obtain a target dialogue generation result.
7. A children's AI dialogue generation system based on IP character settings, characterized by: The steps for implementing the method for generating children's AI dialogue based on IP character settings as described in any one of claims 1 to 6 include: The acquisition module is used to obtain IP character personality data and build the IP character personality constraint matrix; A decoupling module, configured to perform triple decoupling on the IP character constraint matrix to obtain a personality feature sub-vector, a style feature sub-vector, and a knowledge feature sub-vector; a calculation module, configured to perform personality weight attention calculation on the user input text based on the personality feature sub-vector, the style feature sub-vector, and the knowledge feature sub-vector to obtain an attention weight distribution matrix; A joint optimization module is used to input the attention weight distribution matrix into a multi-objective character constraint loss function for joint optimization to obtain intermediate dialogue content that meets the IP character character constraint conditions; A real-time correction module is used to perform educational content adaptation and role identity consistency verification based on the intermediate dialogue content to obtain a pre-generated dialogue text, and to perform real-time character deviation correction on the pre-generated dialogue text to obtain a target dialogue generation result.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for generating children's AI dialogue based on IP character settings according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Man-machine conversation method and system based on self-learning conversation model
CN112541063A
Personnel setting consistency method and device based on dialogue system, electronic equipment and medium
CN116881410A