Personalized language model construction method of intelligent voice assistant

By collecting and processing children's voice, behavior and image data in intelligent voice assistants, training personalized language models and integrating content recommendation systems and multiple interactive methods, the problem of existing intelligent voice assistants lacking personalized and dynamic adjustments is solved, and more efficient children's interaction and growth records are achieved.

CN120220660AInactive Publication Date: 2025-06-27YUANTU ARTIFICIAL INTELLIGENCE (HANGZHOU) CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510647340.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing intelligent voice assistant lacks a personalized language model for children's characteristics, cannot dynamically adjust the content and interaction methods, and lacks the ability to process image data, resulting in a single and highly repetitive interactive content, and is unable to fully record the growth process of children.

Method used

By collecting users' daily usage data, preprocessing and feature extraction, using deep learning algorithms to train personalized language models, and dynamically adjust them according to users' usage. At the same time, it integrates a content recommendation system and a variety of interactive methods to increase fun and attractiveness, and comprehensively record the growth process of children through image data processing.

Benefits of technology

It realizes that intelligent voice assistants can better understand and respond to children's needs, provide personalized services, dynamically adjust content based on children's interests and growth stages, increase the fun and attractiveness of interaction, and comprehensively record the children's growth process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220660A_ABST
    Figure CN120220660A_ABST
Patent Text Reader

Abstract

The invention discloses a personalized language model construction method for an intelligent voice assistant. The personalized language model construction method comprises the following steps: S1, collecting data according to daily use of a user; s2, preprocessing the collected data, and screening and removing interference content; s3, performing feature extraction on the data, and performing division according to features; s4, performing personalized language model training on the feature data by using a deep learning algorithm, and performing dynamic adjustment according to the use condition of the user; s5, verifying the accuracy and effect of the model; s6, deploying the trained model to an intelligent voice assistant to realize real-time interaction; through the personalized language model, the intelligent voice assistant can better understand and respond to demands of children, and more personalized services are provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a method for constructing a personalized language model of an intelligent voice assistant. Background Art

[0002] With the rapid development of artificial intelligence technology, intelligent voice assistants are increasingly widely used in families. Especially in children's education and growth companionship, intelligent voice assistants can provide rich interactive experiences to help children learn new knowledge and skills. However, existing intelligent voice assistants usually adopt general language models and lack personalized designs for children's characteristics, thus unable to fully meet children's needs.

[0003] Existing intelligent voice assistants mainly rely on speech recognition technology, lack the ability to process image data, and cannot comprehensively record children's growth processes; For example, Chinese Patent CN110134623A, "A Children's Education Device Based on Speech Recognition", proposes to provide educational content through voice interaction, but only relies on voice data and lacks comprehensive analysis of multi-dimensional information such as children's behaviors and expressions, resulting in single interactive content and unable to dynamically adapt to changes in children's interests; Chinese Patent CN112487123A, "A Children's Intelligent Dialogue System", uses a pre-trained language model to generate dialogues, but the model parameters are fixed and cannot be optimized in real time according to children's feedback, resulting in high interaction repetition after long-term use and difficulty in maintaining long-term attraction.

[0004] At the same time, most intelligent voice assistants on the current market adopt general language models and fixed content libraries. Although they can provide basic voice interaction functions, they have deficiencies in terms of personalization and adaptability.

[0005] All in all, the existing technologies mainly have the following defects: 1. Lack of a personalized language model for children's characteristics.

[0006] 2. The content is single and cannot be dynamically adjusted according to children's interests and growth stages.

[0007] 3. The interaction method is relatively single and lacks interest and attraction.

[0008] 4. Lack the ability to process image data and cannot comprehensively record children's growth processes.

[0009] Therefore, improvements are needed to address the above deficiencies. Summary of the Invention

[0010] The purpose of the present invention is to provide a method for constructing a personalized language model of an intelligent voice assistant to overcome the deficiencies in the prior art.

[0011] To achieve the above object, the present invention provides the following technical solutions: This application discloses a method for constructing a personalized language model of an intelligent voice assistant, including the following steps: S1: Collect data according to the user's daily usage; S2: Preprocess the collected data, screen and remove interfering content; S3: Extract features from the data and divide them according to the features; S4: Use a deep learning algorithm to train a personalized language model for the feature data and dynamically adjust it according to the user's usage; S5: Verify the accuracy and effectiveness of the model; S6: Deploy the trained model to the intelligent voice assistant to achieve real-time interaction.

[0012] Preferably, the S1 includes the following: Through daily interaction with the user, collect the user's voice data, behavior data and feedback data, and collect matching image data.

[0013] Preferably, the S2 includes the following sub-steps: S21: Clean the data collected in S1 and remove interfering content; S22: Identify according to the collection information and content of the data, label the data for subsequent processing; S23: Normalize the data to facilitate subsequent data extraction and processing and ensure data quality.

[0014] Preferably, the S3 includes the following: According to the content of the data and combined with the preprocessing information, extract features from the data, including extracting voice features, emotional features, behavior features and image features.

[0015] Preferably, the S5 includes the following: Verify the accuracy of the model through cross-validation; and collect the user's real-time feedback, and evaluate the effectiveness of the model according to the user's evaluation.

[0016] Preferably, image acquisition includes the following: A1: Collect image data during the user's use through a camera; A2: Clean, label and normalize the collected image data; A3: Extract information features in the image, including portrait features and object features; A4: Identify the content in the image, judge the identity information of the speaker, and combine it for annotation; A5: Store the processed image data in the database.

[0017] Preferably, it further includes a content recommendation system, which judges the user's interests and usage stages based on the collected data and the content actively input by the user, and conducts intelligent content recommendation and push.

[0018] Preferably, it further includes interactive method recommendation, including voice dialogue, story telling, game interaction, etc., to increase the interest and attraction of interaction.

[0019] Preferably, it includes an intelligent voice assistant, which includes a data acquisition mechanism for collecting data for model training, a speaker for playing voice music, a memory for temporarily storing data, a storage device for long-term data storage, and a processor for data processing and model operation. The data acquisition mechanism includes a microphone for collecting voice data and a camera for collecting image data.

[0020] This application also discloses a personalized language model construction device for an intelligent voice assistant, which includes a memory and one or more processors. An executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the above-mentioned personalized language model construction method for an intelligent voice assistant.

[0021] This application also discloses a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the above-mentioned personalized language model construction method for an intelligent voice assistant.

[0022] Advantages of the present invention: (1) In the embodiment of the present invention, through the personalized language model, the intelligent voice assistant can better understand and respond to the needs of children, and provide more personalized services; (2) In the embodiment of the present invention, through the content recommendation system, it can recommend appropriate learning content and entertainment activities according to the interests and growth stages of children, enriching the lives of children; (3) In the embodiment of the present invention, the adoption of a variety of interactive method designs increases the interest and attraction of interaction, stimulating children's learning interest; (4) In the embodiment of the present invention, the model can be dynamically adjusted according to the growth and development of children, continuously providing high-quality services; (5) In the embodiment of the present invention, through the collection and processing of image data, it can comprehensively record the growth process of children and provide richer growth records.

[0023] The features and advantages of the present invention will be described in detail through embodiments in conjunction with the drawings. Description of the Drawings

[0024] Figure 1It is a flowchart of the steps of a method for constructing a personalized language model of an intelligent voice assistant according to the present invention; Figure 2 It is a schematic diagram of the device of the embodiment of the present invention. Specific embodiments

[0025] To make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. However, it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the scope of the present invention. In addition, in the following description, the description of well-known structures and technologies is omitted to avoid unnecessarily confusing the concepts of the present invention.

[0026] Refer to Figure 1 , the embodiments of the present invention provide a method for constructing a personalized language model of an intelligent voice assistant, an intelligent voice assistant and a method for constructing a personalized language model of an intelligent voice assistant.

[0027] The method for constructing a personalized language model includes the following steps: S1: Collect data according to the user's daily use; S2: Preprocess the collected data, screen and remove interference content; S3: Extract features from the data and divide them according to the features; S4: Use a deep learning algorithm to train a personalized language model for the feature data and dynamically adjust according to the user's usage; S5: Verify the accuracy and effectiveness of the model; S6: Deploy the trained model to the intelligent voice assistant to achieve real-time interaction.

[0028] The said S1 includes the following content: By interacting with the user daily, collect the user's voice data, behavior data and feedback data, and collect the matching image data.

[0029] The said S2 includes the following sub-steps: S21: Clean the data collected in S1 and remove interference content; S22: Identify according to the collection information and content of the data, label the data for subsequent processing; S23: Normalize the data to facilitate subsequent data extraction and processing and ensure the data quality.

[0030] The said S3 includes the following content: According to the content of the data and combined with the preprocessing information, extract features from the data, including extracting voice features, emotional features, behavior features and image features.

[0031] The S5 includes the following: verifying the accuracy of the model through cross-validation; and collecting real-time feedback from users and evaluating the model's effectiveness based on user evaluations.

[0032] Image acquisition includes the following: A1: Collecting image data of the user during use through a camera; A2: Cleaning, annotating, and normalizing the collected image data; A3: Extracting information features from the images, including portrait features and object features; A4: Identifying the content in the images, judging the identity information of the speaker, and combining it for annotation; A5: Storing the processed image data in a database.

[0033] It also includes a content recommendation system that, based on the collected data and the content actively input by the user, judges the user's interests and usage stage and performs intelligent content recommendation and push.

[0034] It also includes recommended interaction methods, including voice conversations, story-telling, game interactions, etc., to increase the fun and attractiveness of the interaction.

[0035] It includes an intelligent voice assistant, which includes a data acquisition mechanism for collecting data for model training, a speaker for playing voice music, a memory for temporarily storing data, a storage device for long-term data storage, and a processor for data processing and model operations. The data acquisition mechanism includes a microphone for collecting voice data and a camera for collecting image data.

[0036] Embodiment: In this embodiment, the intelligent voice assistant is used for children's growth companionship and performs intelligent processing for children's growth and life; The hardware of the intelligent voice assistant includes the following: Microphone: Used to collect voice data.

[0037] Speaker: Used to play voice and music.

[0038] Processor: Used for data processing and model operations.

[0039] Memory: Used to temporarily store data.

[0040] Storage device: Used for long-term data storage.

[0041] Camera: Used to collect image data.

[0042] Personalized language model construction method: Data collection: Collect voice data, behavior data, and feedback data through daily interactions with children; collect image data through a camera.

[0043] Data collection includes the following: 1. When the child interacts with the assistant, voice data is collected through the microphone and uploaded to the cloud in real time; 2. The camera captures a scene image at regular intervals according to a preset time and stores it in the local database after encryption; at the same time, when the child interacts with the assistant, the camera captures images in real time for synchronization with the interaction data.

[0044] Data preprocessing: Clean, annotate, and normalize the collected voice and image data to ensure data quality.

[0045] Feature extraction: Extract voice features, emotional features, behavior features, and image features for subsequent modeling.

[0046] Model training: Use deep learning algorithms (such as LSTM, Transformer, etc.) to train a personalized language model, which can be dynamically adjusted according to the child's age, interests, and growth stage.

[0047] Model evaluation: Evaluate the accuracy and effectiveness of the model through cross-validation and user feedback.

[0048] Model deployment: Deploy the trained model to the intelligent voice assistant to achieve real-time interaction.

[0049] Content recommendation system: Recommend suitable learning content and entertainment activities according to the child's interests and growth stage.

[0050] Interaction method design: Design various interaction methods, including voice conversations, story-telling, game interactions, etc., to increase the fun and attractiveness of the interaction.

[0051] Image data processing module: Image acquisition: Collect image data of the child's daily life through the camera.

[0052] Image preprocessing: Clean, annotate, and normalize the collected image data.

[0053] Image feature extraction: Extract key features in the image, such as faces, objects, etc.

[0054] Image recognition: Use image recognition algorithms to identify the content in the image to assist in judging the identity of the speaker.

[0055] Image recording: Store the processed image data in the database to record the child's growth process.

[0056] In a feasible embodiment, different data collection methods are used, such as parent-assisted collection, third-party data platforms, etc. After data collection, comprehensive calculations are performed and integrated.

[0057] In a feasible embodiment, the feature extraction algorithms adopted include MFCC, CNN, etc.

[0058] In a feasible embodiment, the deep learning models used include GRU, BERT, etc.

[0059] In a feasible embodiment, the content recommendation algorithms adopted include collaborative filtering, deep learning recommendation systems, etc.

[0060] In a feasible embodiment, different interaction methods are designed, such as virtual role-playing, interactive games, etc.

[0061] In a feasible embodiment, the image recognition algorithms used include convolutional neural networks (CNN), deep learning image recognition models, etc.

[0062] An embodiment of the apparatus for constructing a personalized language model of an intelligent voice assistant according to the present invention can be applied to any device with data processing capabilities. The any device with data processing capabilities can be a device or apparatus such as a computer. The apparatus embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful apparatus, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory for operation. From the hardware level, as Figure 2 shown, it is a hardware structure diagram of any device with data processing capabilities where the apparatus for constructing a personalized language model of an intelligent voice assistant according to the present invention is located. In addition to Figure 2 the processor, memory, network interface, and non-volatile memory shown, the any device with data processing capabilities where the apparatus in the embodiment is located usually also includes other hardware according to the actual functions of the any device with data processing capabilities, which will not be elaborated here. The implementation processes of the functions and roles of each unit in the above apparatus are specifically detailed in the implementation processes of the corresponding steps in the above method, which will not be elaborated here.

[0063] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the descriptions of the method embodiments. The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. A person of ordinary skill in the art can understand and implement it without creative work.

[0064] An embodiment of the present invention also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements a personalized language model construction device for an intelligent voice assistant in the above embodiment.

[0065] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be an external storage device of any device with data processing capabilities, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit and an external storage device of any device with data processing capabilities. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and can also be used to temporarily store the data that has been output or will be output.

[0066] The foregoing is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modification, equivalent replacement, or improvement made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for constructing a personalized language model for an intelligent voice assistant, characterized in that: The steps include: S1: Collect data based on users’ daily usage; S2: Preprocess the collected data to filter and remove interference content; S3: Extract features from the data and divide them according to the features; S4: Use deep learning algorithms to train personalized language models based on feature data, and dynamically adjust based on user usage; S5: Verify the accuracy and effectiveness of the model; S6: Deploy the trained model to the intelligent voice assistant to achieve real-time interaction.

2. The method for constructing a personalized language model for an intelligent voice assistant according to claim 1, characterized in that: The S1 includes the following contents: through daily interaction with the user, the user's voice data, behavior data and feedback data are collected, and the image data matching therewith is collected.

3. The method for constructing a personalized language model for an intelligent voice assistant according to claim 1, characterized in that: The S2 comprises the following sub-steps: S21: Clean the data collected by S1 to remove interference content; S22: Identify and label the data based on the collected information and content of the data to facilitate subsequent processing; S23: Normalize the data to facilitate subsequent data extraction and processing and ensure data quality.

4. The method for constructing a personalized language model for an intelligent voice assistant according to claim 1, wherein: The S3 includes the following contents: extracting features of the data according to the content of the data and in combination with the pre-processed information, including extracting the features into speech features, emotion features, behavior features and image features.

5. The method for constructing a personalized language model for an intelligent voice assistant according to claim 1, characterized in that: The S5 includes the following contents: verifying the accuracy of the model through cross-validation; and collecting real-time feedback from users, and evaluating the effect of the model based on the user's evaluation.

6. The method for constructing a personalized language model for an intelligent voice assistant according to claim 1, characterized in that: Image acquisition includes the following: A1: Collect image data during user use through the camera; A2: Clean, label and normalize the collected image data; A3: Extract information features from the image, including portrait features and object features; A4: Identify the content in the image, determine the speaker’s identity information, and annotate it; A5: Store the processed image data into the database.

7. The method for constructing a personalized language model for an intelligent voice assistant according to claim 1, characterized in that: It also includes a content recommendation system, which judges the user's interests and usage stage based on the collected data and the content actively input by the user, and makes intelligent content recommendations and pushes them.

8. The method for constructing a personalized language model for an intelligent voice assistant according to claim 1, characterized in that: It also includes recommendations for interactive methods, including voice dialogue, storytelling, game interaction, etc., to increase the fun and appeal of the interaction.

9. The method for constructing a personalized language model for an intelligent voice assistant according to claim 1, characterized in that: It comprises an intelligent voice assistant, including a data collection mechanism for collecting data for model training, a speaker for playing voice music, a memory for temporarily storing data, a storage device for long-term storage of data and a processor for data processing and model calculation, wherein the data collection mechanism comprises a microphone for collecting voice data and a camera for collecting image data.

Citation Information

Patent Citations

  • Customizable multi queue DMA interface

    CN110134623A

  • Road connectivity test method and system based on large-range high-precision map

    CN112487123A

  • Self-adaptive big language model enhanced personalized content recommendation method and system

    CN118467818A

  • Personalized vehicle-mounted voice assistant learning system and method

    CN118737146A

  • Intelligent voice assistant optimization method and system

    CN119091865A