Children user portrait construction method, system and equipment
By constructing a three-dimensional portrait of child users and using multi-dimensional data collection and fusion technology, the problem that traditional equipment cannot adapt to children's development needs is solved, and precise personalized education is achieved.
Patent Information
- Application Number
- CN202510612100.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-08-15
AI Technical Summary
Traditional children's smart devices push fixed content based on single-dimensional data, which cannot adapt to children's ever-evolving learning needs, making it difficult to achieve accurate personalized education.
By collecting multi-dimensional data, a three-dimensional portrait of children's users is constructed, including voice information, behavioral information and environmental information, multi-modal data is generated, and a model is constructed based on user portraits to push personalized educational content.
It has achieved accurate and comprehensive capture of children's interest preferences, cognitive levels and emotional states, and can reflect children's growing learning needs in real time and realize accurate and personalized educational content push.
Smart Images

Figure CN120499451A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of children's smart devices, and in particular to a method, system and device for constructing a child user portrait. Background Art
[0002] Children's smart devices are a type of intelligent hardware product tailored specifically for children, leveraging cutting-edge technologies like artificial intelligence and the Internet of Things. Their core goal is to provide children with an entertaining and educational learning experience, supporting their cognitive development, language expression, and logical thinking. They also provide companionship and care for children during busy parenting periods. Common children's smart devices on the market, such as reading pens and early childhood education devices, are highly popular among parents and children.
[0003] Most traditional children's smart devices rely on behavior logs or simple voice commands and adopt a fixed content push model. Due to their single data collection dimension, it is difficult to fully capture children's interests, preferences, cognitive levels and emotional states. In this case, if children are in a rapid growth stage, this push model cannot adapt to their evolving learning needs, making it difficult for the device to achieve accurate personalized education.
[0004] Therefore, traditional children's smart devices push fixed content based on single-dimensional data and are unable to adapt to children's evolving learning needs, making it difficult for the devices to achieve accurate personalized education. Summary of the Invention
[0005] In order to overcome the shortcomings of the existing technology, the purpose of the present invention is to provide a method, system and device for constructing a child user portrait, which constructs a child user portrait by collecting multi-dimensional data to adapt to the children's ever-evolving learning needs, thereby enabling the device to achieve accurate personalized education.
[0006] In order to solve the above technical problems, the present invention is implemented according to the following scheme:
[0007] A method for constructing a child user portrait is provided, including:
[0008] Obtain training data for child users;
[0009] Obtaining a user portrait construction model based on the training data;
[0010] Obtain voice information, behavior information, and environmental information of child users;
[0011] Fuse voice information, behavioral information, and environmental information to generate multimodal data;
[0012] The multimodal data is used as input data for the user portrait construction model, and the user portrait construction model outputs a three-dimensional portrait of the child user.
[0013] Compared with the existing technology, the beneficial effects of the method for constructing a child user portrait of the present invention are as follows: by acquiring multi-dimensional information such as voice, behavior and environment and fusing it to generate multimodal data, a three-dimensional portrait of the user is constructed. Compared with traditional single-dimensional data collection, it can more accurately and comprehensively capture children's interest preferences, cognitive levels and emotional states in different scenarios. The constructed three-dimensional portrait can reflect children's ever-evolving learning needs in real time, enabling children's smart devices to push accurate personalized educational content based on the portrait.
[0014] Optionally, the training data includes historical voice information, historical behavior information, and historical environment information of the child user; and obtaining a user portrait construction model based on the training data includes:
[0015] Processing the interference signal of the historical voice information according to the historical environment information to obtain interference-free voice information;
[0016] The user portrait construction model is obtained based on the interference-free voice information and / or the historical behavior information.
[0017] Optionally, the user portrait construction model includes an interest tag model;
[0018] Obtaining the user portrait construction model according to the interference-free voice information and / or the historical behavior information includes:
[0019] Determining, based on the non-interference voice information, the child user's interest information under each interest tag;
[0020] The pre-trained language model is trained according to the interest information to obtain the interest tag model.
[0021] Optionally, the user portrait construction model includes a cognitive assessment model;
[0022] Obtaining the user portrait construction model according to the interference-free voice information and / or the historical behavior information includes:
[0023] determining question-and-answer information of the child user based on the non-interference information and the historical behavior information;
[0024] The hidden Markov model is trained according to the question-answer information to obtain the cognitive evaluation model.
[0025] Optionally, the user portrait construction model includes an emotion recognition model;
[0026] Obtaining the user portrait construction model according to the interference-free voice information and / or the historical behavior information includes:
[0027] Processing the non-interference voice information based on natural language processing technology to determine the emotional information of the child user;
[0028] The emotion recognition model is obtained according to the emotion information.
[0029] Optionally, voice information, behavior information, and environmental information are fused to generate multimodal data, including:
[0030] Based on speech recognition technology, the speech information is converted into speech text data including speech features;
[0031] Extracting time series features of the behavior information to obtain time series data representing the child user's behavior;
[0032] Normalizing the environmental information to obtain scene data representing use by a child user;
[0033] The speech text data, the time series data and the scene data are fused to generate the multimodal data as a feature vector.
[0034] Optionally, fusing the speech text data, the time series data, and the scene data to generate the multimodal data as a feature vector includes:
[0035] Converting the speech text data into speech feature vectors, converting the time series data into behavior feature vectors, and converting the scene data into scene feature vectors;
[0036] Determining correlation parameters between the speech feature vector, the behavior feature vector, and the scene feature vector based on a neural network;
[0037] Determining weights corresponding to the speech feature vector, the behavior feature vector, and the scene feature vector, respectively, based on the attention mechanism and the associated parameters;
[0038] The multimodal data is generated according to the speech feature vector, the behavior feature vector, the scene feature vector and their corresponding weights.
[0039] Optionally, also include:
[0040] Analyze the time series data based on a time series algorithm to obtain behavioral trend data of child users;
[0041] According to the behavior trend data, weights corresponding to the voice feature vector, the behavior feature vector, and the scene feature vector are adjusted respectively to update the multimodal data.
[0042] A child user portrait construction system is also provided, which is applied to the above-mentioned child user portrait construction method, including:
[0043] Training modules for:
[0044] Obtain training data for child users;
[0045] Obtaining a user portrait construction model based on the training data;
[0046] Data fusion module for:
[0047] Obtain voice information, behavior information, and environmental information of child users;
[0048] Fuse voice information, behavioral information, and environmental information to generate multimodal data;
[0049] A portrait construction module is used to use the multimodal data as input data for the user portrait construction model, and the user portrait construction model outputs a three-dimensional portrait of the child user.
[0050] An electronic device is also provided, comprising a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by the processor to implement the method for constructing a child user portrait. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 A flow chart of the method for constructing the present invention. DETAILED DESCRIPTION
[0052] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0053] When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims. In the description of the present application, it should be understood that the terms "first", "second", "third", etc. are only used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence, nor can they be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to the specific circumstances.
[0054] See also Figure 1As shown, a method for constructing a child user portrait of the present invention includes:
[0055] S1: Obtain training data of the child user; in one embodiment of the present invention, the training data includes the child user's historical voice information, historical behavior information and historical environmental information; the training data can be obtained from the smart devices previously used by the child, such as the operation logs stored in the early childhood education device, voice conversation records, and environmental information when the device is used, wherein the operation logs include but are not limited to the behavior records of the child operating the buttons and touching the screen on the device, and the environmental information includes but is not limited to the network status, power level, and usage scenario information detected by the device.
[0056] Historical voice information includes the content of children's conversations with smart devices. Analyzing the most frequently asked questions and discussion topics in historical voice information can identify children's interests within each interest tag. For example, frequently asking questions about the universe indicates an interest in astronomy. Analyzing the question-and-answer content in historical voice information can also determine the child's cognitive level when answering a question. For example, the correctness and complexity of math questions can be used to determine their cognitive level in mathematics. Furthermore, analyzing emotional vocabulary and changes in tone and intonation in historical voice information can identify the child's emotions when communicating with the smart device. For example, a cheerful tone may indicate interest in the content being pushed by the smart device, while a hesitant tone may indicate confusion about the question posed by the smart device.
[0057] Historical behavior information includes records of children's various actions while using smart devices. This information is used to determine children's preferences and abilities. By analyzing children's playback history of stories, Q&A sessions, and other content on their smart devices, their interests can be accurately captured. For example, if a child repeatedly plays dinosaur-themed stories or frequently clicks on science Q&A sessions, this indicates an interest in paleontology; if they frequently watch painting tutorial videos, this indicates an interest in art creation. Furthermore, observing children's behavior while operating smart devices can effectively assess their cognitive abilities. For example, analyzing their speed while completing educational games can determine their cognitive level. If a child can quickly and correctly complete a jigsaw puzzle, this indicates good spatial cognition and hand-eye coordination; if they can quickly and correctly solve puzzles in puzzle games, this indicates strong logical thinking and problem-solving skills. Furthermore, changes in behavior patterns can also indirectly reflect a child's emotional state. For example, a sudden pause in operation or frequent content switching may indicate a loss of interest in the content being pushed by the smart device.
[0058] Historical environmental information includes the environmental conditions in which children use smart devices, and is used to eliminate the impact of environmental factors on the information collected by the device. In different usage environments, the interaction between children and smart devices is also different. Environmental information will interfere with the interaction between children and smart devices, affecting the accuracy of the data. For example, when children are in a noisy public place, the collected voice information may be mixed with a lot of background noise, resulting in a decrease in the accuracy of voice recognition; at this time, by analyzing historical environmental information, identifying the interference factors in the environment, and performing noise reduction and enhancement on the voice information containing interference signals, it is possible to effectively remove the impact of environmental noise and restore clear and accurate voice information, thereby providing a reliable data foundation for subsequent analysis and judgment based on voice information, ensuring that accurate information is extracted from the voice information.
[0059] S2: Based on the training data, the user profile is obtained and the model is constructed, including:
[0060] First, based on historical environmental information, the interference signals in the historical voice information are processed to obtain interference-free voice information. This historical environmental information includes the various environmental conditions in which children used smart devices, such as indoor and outdoor environments and noise levels. In real-world scenarios, environmental factors can significantly impact children's interactions with smart devices, particularly the collection of children's voice information. For example, when children are in noisy environments like shopping malls and school playgrounds, the collected historical voice information will be mixed with a large amount of background noise. This noise acts as interference, severely affecting the quality and accuracy of the voice information. Therefore, we need to identify the interference signal characteristics contained in the historical environmental information. Based on this interference signal characteristic, we can perform noise reduction and filtering on the historical voice information to remove the interference caused by environmental factors, thereby obtaining clear and accurate interference-free voice information. This avoids using voice information containing interference signals to identify children's emotions, cognitive levels, and other information used to construct user profiles.
[0061] Next, a user portrait construction model is obtained based on the interference-free voice information and / or historical behavior information.
[0062] In one embodiment of the present invention, the user portrait construction model includes an interest tag model; in this embodiment, the user portrait construction model is obtained based on the interference-free voice information and / or historical behavior information, including: determining the interest information of the child user under each interest tag based on the interference-free voice information; and training the pre-trained language model based on the interest information to obtain the interest tag model.
[0063] By removing interference signals from historical voice information through historical environmental information, interference-free voice information is obtained, avoiding the situation where voice information with interference signals cannot accurately identify children's emotions towards various interest tags, thereby leading to inaccurate user portraits.
[0064] By training the pre-trained language model (BERT model), an interest tag model is obtained. The input data of the interest tag model is the conversation voice information when children interact with smart devices, and its output data is the children's interest information under each interest tag, specifically the theme clustering results of children's areas of concern, such as "space exploration: 85%", "animal stories: 70%", etc., including interest tags and children's interest level in the interest tags.
[0065] In one embodiment of the present invention, the user portrait construction model includes a cognitive evaluation model; in this embodiment, the user portrait construction model is obtained based on the non-interference voice information and / or historical behavior information, including: determining the question and answer information of the child user based on the non-interference information and historical behavior information; training the hidden Markov model based on the question and answer information to obtain the cognitive evaluation model.
[0066] By removing interference signals from historical voice information through historical environmental information, interference-free voice information is obtained, avoiding the situation where voice information with interference signals cannot accurately identify children's answers to various questions and other information, thereby leading to inaccurate user portraits being constructed.
[0067] The question and answer information includes the content of the child's answer and the complexity of the question. The complexity of the question includes but is not limited to whether the question has a causal relationship, whether there is a logical reasoning level, etc.; before training the hidden Markov model (HMM model) to obtain the cognitive evaluation model, it is necessary to use the question and answer information of the child user to train the decision tree model, divide the question and answer information into multiple levels, and obtain a level determination model for determining the child's cognitive level. The input data of the level determination model is the question and answer information of the child user, and the output data is the child's cognitive level; finally, the output data (cognitive level) of the level determination model is used to train the hidden Markov model (HMM model) to obtain a cognitive evaluation model for predicting the child's degree of knowledge mastery, so that the smart device can dynamically adjust the push content, the child's learning path, etc. according to the child's degree of knowledge mastery.
[0068] In one embodiment of the present invention, the user portrait construction model includes an emotion recognition model; in this embodiment, the user portrait construction model is obtained based on the non-interference voice information and / or historical behavior information, including: processing the non-interference voice information based on natural language processing technology to determine the emotional information of the child user; and obtaining the emotion recognition model based on the emotional information.
[0069] By removing interference signals from historical voice information through historical environmental information, interference-free voice information is obtained, avoiding the situation where voice information with interference signals cannot accurately identify the emotional information of children when interacting with smart devices, thereby leading to inaccurate user portraits being constructed.
[0070] Based on natural language processing technology (NLP technology), the interference-free speech information is processed. Specifically, the speech features such as Mel-frequency cepstral coefficients (MFCC coefficients) and the corresponding text content in the interference-free speech information are first extracted. Then, the natural speech processing technology is used to emotionally label the speech features and the corresponding text content to obtain the emotional information of children when interacting with smart devices. Finally, the support vector machine model (SVM classifier model) is trained. Specifically, the support vector machine model (SVM classifier model) is used to classify the emotional information to obtain an emotion recognition model. The input data of the emotion recognition model is the speech features and the corresponding text content, and the output data is the emotional state of the child, for example, happiness, confusion, frustration, etc.; by integrating speech features with text sentiment analysis, the accuracy of identifying children's emotions is improved.
[0071] S3: Obtain the child user's voice information, behavior information and environmental information; this information includes the real-time voice information obtained when the child interacts with the smart device, the behavior information of operating the smart device, and the environmental information when using the smart device.
[0072] S4: Fusion processing of speech information, behavioral information, and environmental information to generate multimodal data, including:
[0073] First, based on speech recognition technology, the speech information is converted into speech text data including speech features; the temporal features of the behavioral information are extracted to obtain time series data used to represent the behavior of child users; the environmental information is normalized to obtain scene data used to represent the use of child users; finally, the speech text data, time series data and scene data are fused to generate multimodal data as feature vectors.
[0074] In one embodiment of the present invention, based on automatic speech recognition (ASR) technology, voice information is converted into text-based voice text data, and voice features such as speaking speed and intonation are annotated on the voice text data; behavioral information includes various behavioral records of children using smart devices, such as clicking on the screen, playing content, switching interfaces, operating games, etc. Each behavior includes corresponding timestamp information, that is, the specific time point when these behaviors occurred, and the temporal features of the behavioral information are extracted, specifically sorting the behaviors in chronological order to ensure the temporal continuity of the data, and obtaining time series data for reflecting the changing patterns of children's behaviors in the time dimension; environmental information includes numerical data such as network signal strength and power level. Such numerical data is normalized to the [0,1] interval, and combined with the network status and power status during use to obtain scene data for representing children's use.
[0075] In one embodiment of the present invention, speech and text data, time series data, and scene data are fused to generate multimodal data as feature vectors, including:
[0076] First, the speech text data is converted into speech feature vectors, the time series data is converted into behavior feature vectors, and the scene data is converted into scene feature vectors; then, based on the neural network, the correlation parameters between the speech feature vectors, behavior feature vectors and scene feature vectors are determined; finally, based on the attention mechanism and the correlation parameters, the weights corresponding to the speech feature vectors, behavior feature vectors and scene feature vectors are determined; multimodal data is generated based on the speech feature vectors, behavior feature vectors and scene feature vectors and their corresponding weights.
[0077] Speech text data, time series data, and scene data are all obtained from different data sources, and the data types of the three are different. By performing feature extraction on speech text data, time series data, and scene data, we can extract features used to represent key data, convert speech text data, time series data, and scene data into corresponding feature vectors, and standardize data of different data types to obtain feature vectors of the same data type.
[0078] Speech feature vectors, behavioral feature vectors, and scene feature vectors of different dimensions are used to represent information of different dimensions, such as voice communication, behavioral performance, and the scene in which the child user is using the smart device. By learning and fitting the speech feature vectors, behavioral feature vectors, and scene feature vectors through a neural network, correlation parameters reflecting the correlation between feature vectors can be obtained, such as the correlation between the child's speech expression and behavioral performance in a specific scene, or whether a certain behavior is related to the scene in which it is located.
[0079] The purpose of determining the weight of each eigenvector and generating multimodal data based on the attention mechanism and association parameters is to dynamically obtain the key information of the three according to different scenarios and task requirements; since the importance of voice eigenvectors, behavior eigenvectors and scenario eigenvectors in representing the emotions and cognitive levels of child users will vary with the specific scenario, the degree of correlation between the three eigenvectors can be determined through association parameters, and then the attention mechanism can be used to adaptively assign weights to different eigenvectors, highlighting information that is more relevant to the current analysis target and weakening the interference of secondary information, so that children can use smart devices in different scenarios. In the generated multimodal data, the weights of voice information, behavior information and environmental information are all different.
[0080] In one embodiment of the present invention, the construction method also includes: analyzing the time series data based on the time series algorithm to obtain the behavioral trend data of the child user; adjusting the weights corresponding to the voice feature vector, the behavioral feature vector and the scene feature vector according to the behavioral trend data, and updating the multimodal data.
[0081] Children's behavior, speech, and the situations they find themselves in will continue to develop and change over time. Therefore, updating weights based on the time dimension can timely capture these dynamic changes and accurately reflect the behavioral patterns and characteristics of children at different stages. Therefore, by analyzing time series data with a time dimension, we can obtain behavioral trend data that reflects children's changes in the time dimension. Based on this behavioral trend data, we can dynamically adjust the weights corresponding to the three feature vectors to avoid information loss or bias caused by ignoring the time factor. For example, as children age, their language expression ability will gradually improve, and the weight of their speech feature vector may need to be adjusted accordingly to more accurately reflect its importance in multimodal data.
[0082] S5: The multimodal data is used as input data for the user portrait construction model, and the user portrait construction model outputs a three-dimensional portrait of the child user.
[0083] The present invention obtains multi-dimensional information such as voice, behavior and environment and fuses it to generate multimodal data to construct a three-dimensional portrait of the user. Compared with traditional single-dimensional data collection, it can more accurately and comprehensively capture children's interest preferences, cognitive levels and emotional states in different scenarios. The constructed three-dimensional portrait can reflect children's ever-evolving learning needs in real time, enabling children's smart devices to push accurate personalized educational content based on the portrait; at the same time, by dynamically adjusting the weights between the three eigenvectors in the multimodal data, the child user portrait can be dynamically updated.
[0084] A child user portrait construction system of the present invention is applied to the above-mentioned child user portrait construction method, comprising:
[0085] Training modules for:
[0086] Obtain training data for child users;
[0087] Based on the training data, the user portrait is obtained to build the model;
[0088] Data fusion module for:
[0089] Obtain voice information, behavior information, and environmental information of child users;
[0090] Fuse voice information, behavioral information, and environmental information to generate multimodal data;
[0091] The portrait construction module is used to use multimodal data as input data for the user portrait construction model, and the user portrait construction model outputs a three-dimensional portrait of the child user.
[0092] The present invention also provides an electronic device, comprising a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by the processor to implement the above-mentioned construction method.
[0093] The processor can be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0094] The memory can be used to store the program or module, and the processor realizes the various functions of the construction method by running or executing the program or module stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created according to the use of the mobile phone, etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (FlashCard), at least one disk storage device, a flash memory device, or other volatile solid-state storage device.
[0095] The above are only preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.
Claims
1. A method for constructing a child user portrait, characterized in that: include: Obtain training data for child users; Obtaining a user portrait construction model based on the training data; Obtain voice information, behavior information, and environmental information of child users; Fuse voice information, behavioral information, and environmental information to generate multimodal data; The multimodal data is used as input data for the user portrait construction model, and the user portrait construction model outputs a three-dimensional portrait of the child user.
2. A method for constructing a child user portrait according to claim 1, characterized in that: The training data includes historical voice information, historical behavior information, and historical environment information of the child user; based on the training data, a user portrait construction model is obtained, including: Processing the interference signal of the historical voice information according to the historical environment information to obtain interference-free voice information; The user portrait construction model is obtained based on the interference-free voice information and / or the historical behavior information.
3. A method for constructing a child user portrait according to claim 2, characterized in that: The user portrait construction model includes an interest tag model; Obtaining the user portrait construction model according to the interference-free voice information and / or the historical behavior information includes: Determining, based on the non-interference voice information, the child user's interest information under each interest tag; The pre-trained language model is trained according to the interest information to obtain the interest tag model.
4. A method for constructing a child user portrait according to claim 2, characterized in that: The user portrait construction model includes a cognitive evaluation model; Obtaining the user portrait construction model according to the interference-free voice information and / or the historical behavior information includes: determining question-and-answer information of the child user based on the non-interference information and the historical behavior information; The hidden Markov model is trained according to the question-answer information to obtain the cognitive evaluation model.
5. A method for constructing a child user portrait according to claim 2, characterized in that: The user portrait construction model includes an emotion recognition model; Obtaining the user portrait construction model according to the interference-free voice information and / or the historical behavior information includes: Processing the non-interference voice information based on natural language processing technology to determine the emotional information of the child user; The emotion recognition model is obtained according to the emotion information.
6. A method for constructing a child user portrait according to claim 1, characterized in that: The voice information, behavior information and environmental information are integrated and processed to generate multimodal data, including: Based on speech recognition technology, the speech information is converted into speech text data including speech features; Extracting time series features of the behavior information to obtain time series data representing the child user's behavior; Normalizing the environmental information to obtain scene data representing use by a child user; The speech text data, the time series data and the scene data are fused to generate the multimodal data as a feature vector.
7. A method for constructing a child user portrait according to claim 6, characterized in that: The speech text data, the time series data, and the scene data are fused to generate the multimodal data as a feature vector, including: Converting the speech text data into speech feature vectors, converting the time series data into behavior feature vectors, and converting the scene data into scene feature vectors; Determining correlation parameters between the speech feature vector, the behavior feature vector, and the scene feature vector based on a neural network; Determining weights corresponding to the speech feature vector, the behavior feature vector, and the scene feature vector, respectively, based on the attention mechanism and the associated parameters; The multimodal data is generated according to the speech feature vector, the behavior feature vector, the scene feature vector and their corresponding weights.
8. A method for constructing a child user portrait according to claim 7, characterized in that: Also includes: Analyze the time series data based on a time series algorithm to obtain behavioral trend data of child users; According to the behavior trend data, weights corresponding to the voice feature vector, the behavior feature vector, and the scene feature vector are adjusted respectively to update the multimodal data.
9. A child user portrait construction system, applied to the child user portrait construction method according to claims 1-8, characterized in that: include: Training modules for: Obtain training data for child users; Obtaining a user portrait construction model based on the training data; Data fusion module for: Obtain voice information, behavior information, and environmental information of child users; Fuse voice information, behavioral information, and environmental information to generate multimodal data; A portrait construction module is used to use the multimodal data as input data for the user portrait construction model, and the user portrait construction model outputs a three-dimensional portrait of the child user.
10. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by the processor to implement a method for constructing a child user portrait as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Dialogue generation method and device
CN110188177A
Watch robot intelligent autonomous operation system
CN110264791A
Voice signal enhancement method, device and equipment
CN111583946A
Reading equipment, server and data processing method
CN111949773A
Child psychological analysis robot and method based on behavioral psychology
CN113693600A
Cited By
Intelligent children's ability collaborative intervention method, device, medium and product
CN121528520A