Intelligent accompanying system and method based on semantic recognition

Through an intelligent companionship system based on semantic recognition, the problem that traditional methods are difficult to capture the emotional changes of the elderly is solved, emotional recognition and real-time response are achieved, and the efficiency and accuracy of psychological counseling are improved.

CN120011571APending Publication Date: 2025-05-16GUANGXI KINGON SOFTWARE
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510077364.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Traditional methods are difficult to fully capture the emotional changes and psychological needs of the elderly, resulting in poor guidance and manual visits and recording, which affects the accuracy of the analysis.

Method used

An intelligent companionship system based on semantic recognition is adopted, through speech collection, text recognition, emotion recognition and strategy generation units, an emotional model is trained for user emotion recognition, and preset strategies and voice output are initiated based on emotions.

Benefits of technology

It realizes accurate identification and real-time response to emotional changes in the elderly, improves the efficiency and accuracy of psychological counseling, reduces manual intervention, and improves the quality of life of the elderly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011571A_ABST
    Figure CN120011571A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent accompanying system and method based on semantic recognition, relates to the technical field of intelligent questions and answers, and solves the problem that an expected dredging effect cannot be achieved due to the fact that emotional changes and psychological needs of old people are often difficult to comprehensively capture in a traditional method. The method comprises the following steps: converting an acquired voice sample into a text form based on Paraformer voice recognition; constructing an emotional model and performing user emotion recognition based on the emotional model; emotional support strategies and voice output verbal skills are set, time, environment data and user emotions are collected in real time, and the corresponding emotional support strategies and voice output verbal skills are started; a sound copying technology of Ali cloud CosyVoice is adopted to simulate sound of relatives of a user to carry out voice output in a user interaction process. In a word, unprecedented care and support are provided for the elderly in the emotional level, and each elderly in a pension community can feel warmth and accompanying brought by science and technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent question-answering technology, and in particular to an intelligent companionship system and method based on semantic recognition. Background Art

[0002] In terms of scene assessment and response to the emotional state and psychological needs of the elderly in the elderly community, such as loneliness, anxiety, depression and other emotional tendencies, traditional methods mainly rely on manual psychological visits and exchanges. In this process, staff need to personally visit the elderly to communicate face-to-face, then manually register the content of the communication, and then a professional psychologist will conduct a detailed analysis and conduct manual counseling based on this. However, this method faces many challenges in actual operation. First, due to the large number of elderly people in the elderly community, the workload of manual visits is extremely large, resulting in limited attention and communication frequency for each elderly person. Secondly, manual registration of communication content is not only time-consuming and laborious, but also easily affects the accuracy of subsequent analysis due to subjective factors or incomplete records. Furthermore, the scarcity of professional psychologist resources also limits the depth and breadth of analysis and counseling, making it impossible for many elderly people to meet their psychological needs in a timely and effective manner.

[0003] More importantly, due to the limited records of communication frequency and content, traditional methods often fail to fully capture the emotional changes and psychological needs of the elderly, and thus fail to achieve the expected counseling effect. Therefore, exploring more efficient, accurate and personalized psychological assessment and counseling methods has become an urgent problem to be solved in current elderly care community services.

[0004] In view of this, there is a need for an intelligent companionship system and method based on semantic recognition. Summary of the invention

[0005] In view of the problem that traditional methods in the prior art often fail to fully capture the emotional changes and psychological needs of the elderly, and thus fail to achieve the expected counseling effect, the present invention provides an intelligent companionship system and method based on semantic recognition, which can train an emotional model to perform user emotion recognition after collecting and annotating user data, and activate a pre-set strategy based on user emotions, and use pre-set words for voice output. The specific technical solution is as follows:

[0006] An intelligent companion system based on semantic recognition, comprising:

[0007] A voice collection unit, used to collect voice signals input during user interaction;

[0008] A text recognition unit, connected to the voice collection unit, is used to convert the voice signal into text for storage;

[0009] The emotion recognition unit is connected to the text recognition unit and is used to train an emotional model to perform user emotion recognition after collecting and annotating user data;

[0010] The strategy generation unit uses the pre-set words for voice output based on the pre-set emotional support strategy. In addition, the strategy generation unit is also connected to the emotion recognition unit, and also activates the pre-set strategy based on the user's emotion and uses the pre-set words for voice output.

[0011] Preferably, the specific operation of the emotion recognition unit is as follows:

[0012] Emotional annotation is performed on the text transcribed from the speech data to provide necessary supervision information for model training;

[0013] Store the text and sentiment annotation results of speech transcription into big data for AI training and analysis;

[0014] Choose the appropriate model architecture according to the specific application scenario, such as Transformer, RNN, etc.;

[0015] Design network structure, including input layer, encoding layer, decoding layer, etc., to achieve speech-to-text conversion;

[0016] Set model parameters, including the number of layers, number of hidden units, and learning rate, to optimize model performance;

[0017] The Adam optimization algorithm is used to train the model.

[0018] Preferably, emotional support strategies include personalized greetings, weather reminders, holiday wishes, empathetic responses, and emergency calls.

[0019] Preferably, the emotional support strategy is selected by defining a rule set, and the process of defining the rule set is as follows:

[0020] Fixed time trigger: set daily, weekly or monthly reminders;

[0021] Situational awareness: Determine whether a strategy needs to be initiated based on information provided by environmental sensors;

[0022] User input response: When a user expresses a specific keyword or a specific emotion is detected, relevant policies are triggered immediately.

[0023] Preferably, it also includes a voice replication unit, which adopts the voice replication technology of Alibaba Cloud CosyVoice to simulate the voice of the user's relatives to perform voice output during the user interaction process.

[0024] Preferably, it also includes an interest identification and recommendation unit, which identifies the user's favorite topics, activity types and potential emerging fields based on a natural language processing model, and recommends related content.

[0025] Preferably, the process of interest identification is as follows:

[0026] Use BERT as a pre-trained model and fine-tune it based on the characteristics and needs of elderly people's communication. Fine-tuning includes adding health care and family life corpora specific to the lives of the elderly.

[0027] Using named entity recognition technology, we can automatically extract entity information such as names, places, and times from conversations. At the same time, we can combine word frequency statistics to find high-frequency words as possible interest markers.

[0028] Clustering a large amount of conversation text into several topic domains through Latent Dirichlet Allocation to find the themes or topics that users repeatedly mention;

[0029] With the help of deep learning sentiment classifier, the emotional tendency in each conversation is evaluated to help determine which topics can make users feel happy or resonate;

[0030] For each identified topic or activity type, assign one or more tags to facilitate subsequent management and retrieval;

[0031] Through association rule mining or other recommendation algorithms, we can explore the intrinsic connections between different interests and predict new areas that users may be interested in but are not yet aware of.

[0032] An intelligent companionship method based on semantic recognition, applied to the system as described above, comprises the following steps:

[0033] Collecting voice signals input during user interaction and performing data cleaning, wherein the voice signals include voice samples with different accents, speaking speeds, and intonations;

[0034] Based on Paraformer speech recognition, the collected speech samples are converted into text;

[0035] Build an emotional model and identify user emotions based on the emotional model;

[0036] Set emotional support strategies and voice output scripts, collect time, environment data and user emotions in real time, and start corresponding emotional support strategies and voice output scripts;

[0037] Adopting the voice replication technology of Alibaba Cloud CosyVoice, the voice of the user's relatives is simulated for voice output during user interaction;

[0038] Based on natural language processing models, we identify users’ favorite topics, activity types, and potential emerging fields, and recommend relevant content.

[0039] A computer-readable storage medium, the computer-readable storage medium comprising a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to implement the intelligent companionship method based on semantic recognition as described above.

[0040] A processor is used to run a program, wherein the program executes the intelligent companionship method based on semantic recognition as described above when running.

[0041] Compared with the prior art, the present invention has the following beneficial effects:

[0042] The present invention is based on Paraformer speech recognition, converting the collected speech samples into text form; constructing an emotional model and performing user emotion recognition based on the emotional model; setting emotional support strategies and voice output scripts, collecting time, environmental data and user emotions in real time, and starting corresponding emotional support strategies and voice output scripts; using the voice replication technology of Alibaba Cloud CosyVoice to simulate the voices of the user's relatives for voice output during user interaction; based on the natural language processing model, identifying the user's favorite topics, activity types and potential emerging fields, and recommending related content. In short, the present invention provides unprecedented care and support to the elderly at the emotional level, allowing every elderly person in the retirement community to feel the warmth and companionship brought by technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for the specific embodiments or the description of the prior art. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn according to the actual scale.

[0044] Figure 1 is an operation flow chart of the emotion recognition unit of the present invention;

[0045] Figure 2 Generate a unit operation flow chart for the strategy of the present invention;

[0046] Figure 3 This is an operation flow chart of the interest identification and recommendation unit of the present invention. DETAILED DESCRIPTION

[0047] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0048] It should be understood that when used in this specification and the appended claims, the terms "include" and "comprises" indicate the presence of described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.

[0049] It should also be understood that the terms used in the present specification are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the present specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.

[0050] It should be further understood that the term "and / or" used in the present description and the appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0051] In one embodiment of the present invention, an intelligent companion system based on semantic recognition is provided, comprising:

[0052] 1. Voice Collection Unit

[0053] The voice collection unit is used to collect the voice signal input during the user interaction process;

[0054] Specifically: Collect a large amount of high-quality voice data, including voice samples with different accents, speaking speeds, and intonations to ensure the diversity and richness of the data. Data cleaning: Preprocess the collected voice data, including removing noise, standardizing volume, and segmenting voice segments, to improve data quality and analyze happy, sad, and angry emotional data.

[0055] 2. Text Recognition Unit

[0056] The text recognition unit is connected to the voice collection unit and is used to convert the voice signal into text for storage;

[0057] Specifically: Using Paraformer speech recognition technology to transcribe audio data into text can efficiently transcribe audio data into text. With its excellent sequence modeling capabilities and efficient computing performance, Paraformer speech recognition technology has successfully achieved accurate recognition of the semantic content of the elderly's voice chat. In the daily operation of the elderly community, whether asking about daily services, expressing personal needs, or sharing daily life, Paraformer technology can quickly capture and accurately understand the elderly's voice content. Through deep learning and complex algorithm models, this technology effectively overcomes the challenges that the elderly may have, such as unclear pronunciation, slow speech speed, or dialect accents, ensuring the accuracy and reliability of speech recognition.

[0058] 3. Emotion Recognition Unit

[0059] The emotion recognition unit is connected to the text recognition unit to collect and annotate user data and then train the emotional model to perform user emotion recognition. The specific operations are as follows:

[0060] 1. Perform sentiment annotation on the text transcribed from the speech data to provide necessary supervision information for model training.

[0061] 2. Store the text and sentiment annotation results after speech transcription into big data for AI training and analysis;

[0062] 3. Choose the appropriate model architecture according to the specific application scenario, such as Transformer, RNN, etc.;

[0063] 4. Set model parameters, such as the number of layers, number of hidden units, learning rate, etc., to optimize model performance;

[0064] 5. Use Adam optimization algorithm to train the model.

[0065] For voice intercom tasks, the data sets are often very large and high-dimensional due to the large amount of audio feature extraction and complex language model training involved. In this case, using the Adam optimization algorithm can bring significant advantages: first, Adam can speed up the training process and help the model reach a stable state more quickly; second, it helps improve the final performance of the model, especially when fine-tuning weights is required to capture the nuances of speech. Therefore, when pursuing an efficient and high-quality voice interaction experience, the Adam optimization algorithm is usually a better choice. In addition, considering the dynamics and complexity of voice intercom tasks, using an optimizer with adaptive capabilities like Adam is also more helpful for the model to adapt to different input conditions.

[0066] In this system voice intercom, the training and optimization of the emotional model is the key to achieving a natural and humanized interactive experience. This process involves extracting valuable emotional information from massive conversation text data and using it to train an AI model that can accurately identify and respond to emotional states. The following will detail the training process, loss function selection, model evaluation, and model optimization of the emotional model training.

[0067] 1) Training process:

[0068] First, the training process begins with the data preprocessing stage. For the emotional label data in the voice intercom of this system, this includes steps such as cleaning the text to remove irrelevant characters, standardizing the format, word segmentation, and removing stop words to ensure that the data input to the model is clean and structured. Next, each conversation or sentence is labeled as a specific emotional category, such as happy, sad, angry, or neutral, through emotional annotation. These labeled data sets are then divided into training sets, validation sets, and test sets, where the training set is used to actually train the model, while the validation set and test set are used to monitor overfitting and ultimately evaluate model performance, respectively.

[0069] At the beginning of training, the preprocessed data is input into the selected model architecture, such as Transformer-based BERT (Bidirectional Encoder Representations from Transformers) or its variant RoBERTa, which are good at capturing complex patterns in context. During training, the model will iteratively adjust parameters according to the back-propagation algorithm to minimize the difference between the predicted output and the true label. In order to improve efficiency and effectiveness, batch gradient descent or more advanced optimizers such as Adam are usually used for parameter updates. One of the most commonly used libraries is Deeplearning4j (DL4J)

[0070] 2) Loss Function

[0071] The choice of loss function is crucial to guide model learning. In sentiment classification tasks, cross entropy loss is one of the most commonly used loss functions because it effectively measures the difference between the predicted probability distribution and the actual label. Specifically, binary cross entropy is suitable for two-category sentiment classification, while multi-category cross entropy is suitable for multi-category sentiment analysis. In addition, when it comes to sequence modeling, the connectionist temporal classification (CTC) loss can be used to handle problems with variable-length inputs, which allows the model to directly map from features to labels without explicitly aligning inputs and outputs.

[0072] 3) Model evaluation

[0073] Model evaluation is performed on the validation set to check whether the model has learned the correct sentiment classification ability. Evaluation indicators include but are not limited to recognition accuracy (Precision), recall (Recall), F1 score, etc. Accuracy measures the proportion of correct predictions made by the model; recall refers to the proportion of all positive examples that are correctly identified; F1 score takes both into account and provides a balanced evaluation standard. In addition to the above statistical indicators, confusion matrices can also be used to intuitively display the predictions of different categories to help discover possible bias or misclassification problems in the model.

[0074] 4) Model optimization

[0075] Once the initial evaluation is completed, it is time to enter the model optimization phase. If the model performs poorly, it can be improved in a variety of ways. On the one hand, you can consider adjusting the model architecture itself, such as increasing the network depth, changing the number of hidden layer units, or introducing residual connections and other techniques to enhance the model's expressiveness. On the other hand, increasing data diversity is also an effective way to improve the generalization ability of the model, which can be achieved by collecting more types of dialogue samples, using data enhancement methods to generate synthetic dialogues, or transfer learning. In addition, hyperparameter tuning is equally important, including learning rate, batch size, regularization coefficient, etc., which can be found through strategies such as grid search, random search, or Bayesian optimization. The best settings can be found.

[0076] In short, the training and optimization of emotional models in voice intercom is an iterative process. It requires continuous attention to the latest research results and technological developments, and the flexible application of different tools and technical means to ensure that the model can achieve optimal performance, thereby providing users with a more humane and empathetic communication experience.

[0077] 4. Strategy Generation Unit

[0078] The strategy generation unit uses the pre-set words for voice output based on the pre-set emotional support strategy. In addition, the strategy generation unit is also connected to the emotion recognition unit, and also activates the pre-set strategy based on the user's emotion and uses the pre-set words for voice output.

[0079] Emotional support strategies include:

[0080] 1) Daily care and communication

[0081] Personalized Greetings: Personalized morning, midday and evening greetings based on the senior’s name and preferences.

[0082] Weather reminder: Reports the weather conditions of the day every morning and gives corresponding suggestions (such as wearing warm clothes). Health reminder: Reminds you to take medicine or have a physical examination on time.

[0083] Activity suggestions: Recommend suitable indoor or outdoor activities.

[0084] Interest topics: Discuss the books, movies or music that the elderly are interested in. Holiday greetings: Send blessings and greetings on important holidays.

[0085] Community Activities: Introducing information about activities within the community.

[0086] 2) Psychological counseling and emotional management Empathic response: Express understanding and sympathy for the elderly’s emotions.

[0087] Positive Feedback: Give positive comments to boost your confidence.

[0088] Stress relief: Recommend appropriate stress relief methods, such as taking a walk or listening to music. Safety building: Ensure that the elderly feel safe and protected.

[0089] Crisis intervention: Identifying emergency situations and directing help seeking.

[0090] Mental health education: popularize mental health knowledge and improve self-awareness.

[0091] 3) Social interaction and entertainment social bridge: promote online communication with other elderly people.

[0092] Musical accompaniment: Play music or songs that the elderly like.

[0093] Poetry Recitation: Reading classic poetry or literary works.

[0094] Fun Trivia: Play a light-hearted trivia game.

[0095] Art Appreciation: Introducing stories about works of art or artists.

[0096] Handicraft: Guide the production of simple handicrafts.

[0097] Cooking Tutorials: Provide simple recipes and explain the production process.

[0098] Pet Simulation: Simulate the sounds of pets to bring fun to the elderly.

[0099] 4) Healthy living and learning exercise reminder: Remind to do moderate physical exercise.

[0100] Dietary advice: Provide healthy diet suggestions.

[0101] Reading recommendations: Recommend suitable books or articles.

[0102] Language Learning: Provides simple foreign language learning resources.

[0103] Upskilling: Teach new skills or hobbies.

[0104] Travel Dreams: Discuss future travel plans or reminisce about past travel experiences.

[0105] Environmental awareness: Cultivate environmentally friendly habits, such as saving water and electricity.

[0106] 5) Safety assurance and emergency response

[0107] Safety Check: Remind you to regularly check for potential safety hazards in your home.

[0108] Emergency call: Set up a one-touch emergency help function.

[0109] Medication management: Assist in managing medication storage and usage.

[0110] Anti-fraud education: warn of common scams and increase vigilance.

[0111] Accident Prevention: Provides safety tips to prevent accidents such as falls.

[0112] Legal Aid: Provide necessary legal consultation channels.

[0113] Medical Appointments: Helps with making appointments for doctor or hospital services.

[0114] Psychological hotline: Provides professional psychological counseling telephone.

[0115] Neighborhood mutual aid: Encourage participation in community mutual aid activities.

[0116] Safety drills: Provide guidance on fire escape and other safety drills.

[0117] The settings include:

[0118] Personalized Greetings: Personalized morning, midday and evening greetings based on the senior’s name and preferences.

[0119] "Good morning, [Name]! I hope you have a great day."

[0120] "Good afternoon, [Name], it's a nice sunny day today, how are you feeling?"

[0121] "Good evening, [Name], have a nice night."

[0122] "Good evening, [Name], are you ready to rest?"

[0123] Weather reminder: Reports the weather conditions every morning and gives corresponding suggestions (such as wearing warm clothes). "It's going to rain today, remember to bring an umbrella, [name]."

[0124] "The temperature has dropped, please wear more clothes when you go out, okay?"

[0125] "The temperature is just right today, it's a good time for a walk."

[0126] "The weather changes a lot, so be careful to add or remove clothes."

[0127] Listen to their hearts: Encourage the elderly to express their feelings.

[0128] "I'm here listening, you can talk to me about your feelings."

[0129] "How are you feeling today? Is there anything you want to share with me?"

[0130] Empathic response: Show understanding and sympathy for the older adult’s emotions.

[0131] "I can understand that this is not easy for you."

[0132] "It sounds like this matter has caused you a lot of trouble."

[0133] "I can imagine how hard this must be for you."

[0134] “It’s really brave of you to tell me this so openly.

[0135] Social bridge: Facilitate online communication with other seniors.

[0136] “Would you like to call [family / friend] today to chat?”

[0137] "I can set you up with a video call. Who would you like to meet?"

[0138] Safety Check: Remind you to regularly check for potential safety hazards in your home.

[0139] "I checked the electrical appliances at home today and everything is normal."

[0140] "Remember to check regularly whether the gas valve is closed properly."

[0141] "Check if there are any damaged wires in the room."

[0142] These strategies can not only improve the quality of life of the elderly, but also effectively alleviate their loneliness and anxiety, while providing necessary psychological support and a sense of security. The design of this system should fully consider the operating habits and technical acceptance of the elderly, ensuring that its interface is friendly and easy to use.

[0143] When communicating with the elderly through voice, it is very important to choose the right time to start a specific emotional support strategy. Reasonable timing can ensure the effectiveness of the strategy and make the elderly feel cared for and supported rather than disturbed. For this system, using a rule-based AI model system (Rulebased System) to select emotional support strategies can be a simple and effective starting point, especially in the early stages of the project or when resources are limited. The rule-based AI model system relies on predefined conditions and logic to decide when to start a specific emotional support strategy. The following are specific methods and precautions for designing and implementing such a system:

[0144] Design a rule-based emotional support strategy AI model system:

[0145] 1) Define the rule set

[0146] Fixed time trigger: Set daily, weekly or monthly reminders, such as morning greetings, weather forecasts, health tips, etc.

[0147] Context awareness: Determine whether a policy needs to be initiated based on information provided by environmental sensors (such as light, temperature, and motion detectors).

[0148] User input response: When the elderly express specific needs or emotions, relevant strategies are triggered immediately. For example, "I feel a little lonely" may trigger social interaction suggestions.

[0149] 2) Build a rule base

[0150] Categorical labels: Assign one or more labels to each policy, such as "Mental Health," "Recreation and Leisure," "Family Connections," and so on.

[0151] Prioritization: Determine which strategies should take priority in a particular situation. For example, in crisis intervention, emergency help should be prioritized first.

[0152] Condition combination: allows you to combine multiple conditions to form more complex rules. For example, "play light music after 8pm and when no one is in the room"

[0153] 3) Implement logical processing

[0154] State Machine: Use a finite state machine (FSM) to manage transitions between different policies. For example, going from “daily greeting” to “weather reminder”.

[0155] Decision tree: Simplify the multi-condition judgment process by building a decision tree. Each path represents a strategy choice.

[0156] 5. Sound Reproduction Unit

[0157] The voice replication unit uses Alibaba Cloud CosyVoice's voice replication technology to simulate the voices of the user's relatives for voice output during user interaction.

[0158] 6. Interest Identification and Recommendation Unit

[0159] The interest identification and recommendation unit identifies users' favorite topics, activity types, and potential emerging fields based on natural language processing models, and recommends relevant content.

[0160] The following are the specific workflow and technical details:

[0161] 1. Selection and application of pre-trained language models

[0162] Model selection: Use pre-trained models like BERT (Bidirectional Encoder Representations from Transformers) and RoBERTa (Robustly Optimized BERT Pretraining Approach) as they perform well on a wide range of NLP tasks and have strong context understanding capabilities.

[0163] Fine-tuning and adaptation: These pre-trained models are fine-tuned according to the characteristics and needs of elderly people's communication. For example, by adding corpora specific to the lives of the elderly (such as health care, family life, etc.), the model can more accurately capture the unique expressions and interests of this group.

[0164] 2. Conversation content analysis

[0165] Keyword extraction: Use named entity recognition (NER) technology to automatically extract entity information such as names, places, and times from the conversation; at the same time, combine word frequency statistics to find high-frequency words as possible interest markers.

[0166] Topic modeling: Through Latent Dirichlet Allocation (LDA) or other topic modeling algorithms, a large amount of conversation text is clustered into several topic domains to discover themes or topics that users repeatedly mention.

[0167] Sentiment analysis: With the help of deep learning sentiment classifiers (such as LSTM / GRU-based models), the emotional tendency in each conversation is evaluated to help determine which topics can make users feel happy or resonate.

[0168] 3. Point of interest identification and expansion

[0169] Interest tag generation: For each identified topic or activity type, one or more tags are assigned to it for easy subsequent management and retrieval, for example, "gardening", "cooking", "music appreciation", etc.

[0170] Potential interest mining: Through association rule mining or other recommendation algorithms, we can explore the intrinsic connections between different interests and predict new areas that users may not yet be aware of but may be interested in. For example, if an elderly person often talks about his travel experiences, the system may suggest that he try photography or write travel notes.

[0171] Example scenario:

[0172] Assume that Grandpa Li often chats with this system, mentions that he likes to grow flowers, listen to Peking Opera, and occasionally expresses his desire to learn new things. This system will handle it like this:

[0173] Keyword extraction: Identify high-frequency words such as "planting flowers" and "Peking opera".

[0174] Topic Modeling: Discovered that “Gardening” and “Traditional Culture” are Grandpa Li’s main areas of interest.

[0175] Sentiment analysis: It is detected that whenever talking about Peking Opera, Grandpa Li's emotions are more positive.

[0176] Interest tag generation: label Grandpa Li’s interests as “gardening enthusiast” and “Peking Opera fan”.

[0177] In this way, the system can not only understand and remember Grandpa Li’s existing interests and hobbies, but also constantly guide him to explore new possibilities and enrich his daily life experience.

[0178] In order to extract key points of interest from each conversation, we can use pre-trained language models such as BERT or RoBERTa. These models are able to understand the context and extract key information. We will use Hugging Face's Transformers library to load and use these models.

[0179] Based on the interaction data analysis accumulated over a long period of time by the interest point analysis AI model, a detailed user interest and hobby portrait is constructed. This user portrait not only includes known interest points, but also combines the group characteristics of other similar users. Through big data analysis, the system can discover more subtle interest patterns and continuously update and improve the portrait over time. This makes the recommended content more accurate and better meet individual needs.

[0180] In order to realize the construction of user portraits based on interest point analysis, the user's interaction data will be collected, and the user's interest points will be identified and expanded through big data analysis techniques (such as clustering algorithms) to form feature vectors.

[0181] Here is a complete Java code example showing how to build such a system:

[0182] 1. Collect user interaction data.

[0183] 2. Use K-Means clustering algorithm to perform cluster analysis on user data. Specifically, use Apache CommonsMath library to perform K-Means cluster analysis.

[0184] 3. Update user portraits based on clustering results.

[0185] 4. Form the feature vector.

[0186] We will use the Apache Commons Math library to perform K-Means clustering analysis. First, make sure you have added the Apache Commons Math library to your project.

[0187] Here is the complete Java code example:

[0188]

[0189]

[0190]

[0191] illustrate:

[0192] 1. UserInteractionData: This is a simple class used to store user interaction data, including user ID and points of interest.

[0193] 2.buildUserInterestMatrix: This method is used to build the user interest matrix, where each user has an interest point and its corresponding count.

[0194] 3.convertToPoints: This method converts the user interest matrix into a data format suitable for clustering (a list of `DoublePoint`).

[0195] 4.updateUserProfile: This method uses the K-Means clustering algorithm to cluster user data, and updates the user profile based on the clustering results and forms a feature vector.

[0196] This system has a huge learning resource database, covering various types of educational resources, such as online courses, e-books, video tutorials, etc. This database is carefully selected and classified, and creating a vector representation describing the content characteristics of each learning resource (such as a book, a class) is one of the key steps in building an efficient personalized recommendation system. This process usually involves natural language processing (NLP), machine learning, and deep learning techniques to ensure that the generated vector can accurately reflect the core characteristics of the resource and support subsequent similarity calculation and recommendation generation.

[0197] In order to create a database of learning resources and generate vector representations for each resource that describe its content characteristics, we can use word embedding models in natural language processing technology (such as Word2Vec, GloVe) or more advanced pre-trained models (such as BERT). In this example, we will use Java and BERT to implement this function. To simplify the example, we will assume that we already have a pre-trained BERT model and can get the vector representation of the text through API calls.

[0198] In order to achieve personalized recommendations based on user interest profiles, the use of a hybrid recommendation system is indeed a very effective strategy. This system combines the advantages of collaborative filtering and content-based recommendations. The following are specific methods and technical details:

[0199] Construction of hybrid recommendation system:

[0200] Collaborative Filtering:

[0201] User-item matrix: In the interest identification and recommendation unit, a rating matrix is ​​established that includes all users and the resources they have interacted with (such as books, courses, and activities).

[0202] Similarity calculation: Use methods such as cosine similarity and Pearson correlation coefficient to calculate the similarity between users or items.

[0203] Neighbor selection: Find a set of users (K nearest neighbors) that are most similar to the target user and predict the target user’s preferences based on the behavior of these neighbors.

[0204] Recommendation generation: Provide users with a personalized list of recommendations based on the choices of similar users or the relevance of similar items.

[0205] Content-based Recommendation:

[0206] Resource feature vector: For each resource (such as a book or a class), create a vector representation that describes the characteristics of its content.

[0207] User preference model: Based on user profiles and historical behaviors, a preference model is established for each user to capture their interests in different types of resources.

[0208] Matching score calculation: Calculate the similarity between the resource feature vector and the user preference model as the recommendation score.

[0209] Recommendation sorting: Sort candidate resources by score and select the top few as the final recommendation results.

[0210] Mixed recommendations:

[0211] Weighted average method: Simply add the results of collaborative filtering and content-based recommendation according to certain weights to get a comprehensive score.

[0212] Stacked model: Train a machine learning model (such as Random Forest, XGBoost) with the output of two recommendation methods as input features to predict the final recommendation score.

[0213] Fusion network: Design a neural network structure that accepts the embedded representation of users and items at the same time, and outputs the recommendation score after multiple layers of nonlinear transformation.

[0214] Examples of actual application scenarios:

[0215] Suppose an old lady named Grandma Wang often communicates with this system, expresses her love for gardening, and occasionally mentions her desire to learn new things. This system will handle it like this:

[0216] 1. Build a user interest profile: By analyzing Grandma Wang’s conversation records, we found that she is interested in topics such as "gardening" and "handmade", and her emotions are more positive whenever she talks about these.

[0217] 2. Hybrid Recommendation System:

[0218] Collaborative filtering: The system found that other elderly people who also like gardening also participated in the painting class, so it suggested that Grandma Wang try it as well.

[0219] Content-based recommendations: Based on the content that Grandma Wang has browsed in the past, some professional books and video tutorials on plant care are recommended.

[0220] Fusion Networking: Combining the two approaches above, this system also suggests activities that fit her existing interests while also expanding her new skills, such as attending craft shows organized by the community.

[0221] In this way, the system can not only understand and remember Grandma Wang’s existing interests and hobbies, but also constantly guide her to explore new possibilities and enrich her daily life experience, while ensuring that the recommended content always remains fresh and attractive.

[0222] Finally, the system uses powerful AI language models to integrate all the above information and generate personalized recommendation lists. These models have strong semantic understanding and generation capabilities, and can synthesize the most suitable activities, social circles or learning resource suggestions for the elderly according to their specific circumstances. For example, if the elderly are interested in history, the system may recommend a documentary about ancient civilizations; if they are calligraphy enthusiasts, they may introduce an online calligraphy class. In this way, the system not only becomes a good partner in the lives of the elderly, but also helps them explore the new world and improve their quality of life.

[0223] In order to realize a conversation content generation system that combines user voice text and recommended resource data, we can use the Tongyi Qianwen AI language model to generate conversation content. We will simulate a simple scenario in which Tongyi Qianwen generates corresponding conversation content based on the user's voice text and recommended resource data.

[0224] Here is a complete Java code example showing how to implement this functionality:

[0225] 1.Simulate user voice text input.

[0226] 2. Simulate recommended resource data.

[0227] 3. Use the Tongyi Qianwen AI language model to generate dialogue content.

[0228] To simplify the example, we will assume that there is an API that can call the Tongyi Qianwen AI language model. In this example, we will use OpenAI's GPT-3 as a proxy to simulate the functionality of Tongyi Qianwen. You need to register and get the OpenAI API key first. First, make sure you have installed the `OkHttp` library to handle HTTP requests. If you are using a Maven project, you can add the following dependencies in `pom.xml`.

[0229] Here is the complete Java code example:

[0230]

[0231]

[0232] illustrate:

[0233] 1.OPENAI_API_KEY: Replace with your own OpenAI API key.

[0234] 2.generateDialogue: This method builds the request body and calls the OpenAI API to generate the dialogue content.

[0235] 3.buildPrompt: This method builds the prompt text, which includes the user's voice text input and recommended resource data. Run steps:

[0236] 1.Save the above code as `HybridDialogueGenerator.java`.

[0237] 2. Replace `OPENAI_API_KEY` with your own OpenAI API key.

[0238] 3. Compile and run the Java program.

[0239] In this way, you can see the conversation content generated based on the user's voice text and recommended resource data. In addition, you can further expand and optimize this system as needed, such as adding more user data, improving the recommendation algorithm, etc.

[0240] In one embodiment of the present invention, there is also provided an intelligent companionship method based on semantic recognition, which is applied to the system described in the above embodiment, and comprises the following steps:

[0241] Step 1: Collect the voice signal input during the user interaction process and perform data cleaning, wherein the voice signal includes voice samples with different accents, speaking speeds, and intonations.

[0242] Step 2: Based on Paraformer speech recognition, the collected speech samples are converted into text;

[0243] Step 3: Build an emotional model and perform user emotion recognition based on the emotional model.

[0244] The construction process of the emotional model is as follows:

[0245] 1. Perform sentiment annotation on the text transcribed from the speech data to provide necessary supervision information for model training.

[0246] 2. Store the text and sentiment annotation results after speech transcription into big data for AI training and analysis;

[0247] 3. Choose the appropriate model architecture according to the specific application scenario, such as Transformer, RNN, etc.;

[0248] 4. Design a reasonable network structure, including input layer, encoding layer, decoding layer, etc., to achieve speech-to-text conversion;

[0249] 5. Set model parameters, such as the number of layers, number of hidden units, learning rate, etc., to optimize model performance;

[0250] 6. Use Adam optimization algorithm to train the model.

[0251] Step 4: Set emotional support strategies and voice output scripts, collect time, environmental data and user emotions in real time, and activate corresponding emotional support strategies and voice output scripts.

[0252] Specifically, a rule-based AI model system is used to select emotional support strategies. The design process of the rule-based emotional support strategy AI model system is as follows:

[0253] 1) Define the rule set

[0254] Fixed time trigger: Set daily, weekly or monthly reminders, such as morning greetings, weather forecasts, health tips, etc.

[0255] Context awareness: Determine whether a policy needs to be initiated based on information provided by environmental sensors (such as light, temperature, and motion detectors).

[0256] User input response: When the elderly express specific needs or emotions, relevant strategies are triggered immediately. For example, "I feel a little lonely" may trigger social interaction suggestions.

[0257] 2) Build a rule base

[0258] Categorical labels: Assign one or more labels to each policy, such as "Mental Health," "Recreation and Leisure," "Family Connections," and so on.

[0259] Prioritization: Determine which strategies should take priority in a particular situation. For example, in crisis intervention, emergency help should be prioritized first.

[0260] Condition combination: allows you to combine multiple conditions to form more complex rules. For example, "play light music after 8pm and when no one is in the room"

[0261] 3) Implement logical processing

[0262] State Machine: Use a finite state machine (FSM) to manage transitions between different policies. For example, going from “daily greeting” to “weather reminder”.

[0263] Decision tree: Simplify the multi-condition judgment process by building a decision tree. Each path represents a strategy choice.

[0264] Step 5: Use Alibaba Cloud CosyVoice's voice replication technology to simulate the voices of the user's relatives for voice output during user interaction.

[0265] Step 6: Based on the natural language processing model, identify the user’s favorite topics, activity types, and potential emerging fields, and recommend relevant content.

[0266] The following are the specific workflow and technical details:

[0267] 1. Selection and application of pre-trained language models

[0268] Model selection: Use pre-trained models like BERT (Bidirectional Encoder Representations from Transformers) and RoBERTa (Robustly Optimized BERT Pretraining Approach) as they perform well on a wide range of NLP tasks and have strong context understanding capabilities.

[0269] Fine-tuning and adaptation: These pre-trained models are fine-tuned according to the characteristics and needs of elderly people's communication. For example, by adding corpora specific to the lives of the elderly (such as health care, family life, etc.), the model can more accurately capture the unique expressions and interests of this group.

[0270] 2. Conversation content analysis

[0271] Keyword extraction: Use named entity recognition (NER) technology to automatically extract entity information such as names, places, and times from the conversation; at the same time, combine word frequency statistics to find high-frequency words as possible interest markers.

[0272] Topic modeling: Through Latent Dirichlet Allocation (LDA) or other topic modeling algorithms, a large amount of conversation text is clustered into several topic domains to discover themes or topics that users repeatedly mention.

[0273] Sentiment analysis: With the help of deep learning sentiment classifiers (such as LSTM / GRU-based models), the emotional tendency in each conversation is evaluated to help determine which topics can make users feel happy or resonate.

[0274] 3. Point of interest identification and expansion

[0275] Interest tag generation: For each identified topic or activity type, one or more tags are assigned to it for easy subsequent management and retrieval, for example, "gardening", "cooking", "music appreciation", etc.

[0276] Potential interest mining: Through association rule mining or other recommendation algorithms, we can explore the intrinsic connections between different interests and predict new areas that users may not yet be aware of but may be interested in. For example, if an elderly person often talks about his travel experiences, the system may suggest that he try photography or write travel notes.

[0277] Example scenario:

[0278] Assume that Grandpa Li often chats with this system, mentions that he likes to grow flowers, listen to Peking Opera, and occasionally expresses his desire to learn new things. This system will handle it like this:

[0279] Keyword extraction: Identify high-frequency words such as "planting flowers" and "Peking opera".

[0280] Topic Modeling: Discovered that “Gardening” and “Traditional Culture” are Grandpa Li’s main areas of interest.

[0281] Sentiment analysis: It is detected that whenever talking about Peking Opera, Grandpa Li's emotions are more positive.

[0282] Interest tag generation: label Grandpa Li’s interests as “gardening enthusiast” and “Peking Opera fan”.

[0283] In this way, the system can not only understand and remember Grandpa Li’s existing interests and hobbies, but also constantly guide him to explore new possibilities and enrich his daily life experience.

[0284] In order to extract key points of interest from each conversation, we can use pre-trained language models such as BERT or RoBERTa. These models are able to understand the context and extract key information. We will use Hugging Face's Transformers library to load and use these models. Since Java itself does not directly support these deep learning models, we can use Python scripts to process the text and call the Python scripts from Java.

[0285] Based on the interaction data analysis accumulated over a long period of time by the interest point analysis AI model, a detailed user interest and hobby portrait is constructed. This user portrait not only includes known interest points, but also combines the group characteristics of other similar users. Through big data analysis, the system can discover more subtle interest patterns and continuously update and improve the portrait over time. This makes the recommended content more accurate and better meet individual needs.

[0286] In order to realize the construction of user portraits based on interest point analysis, we can use Java to write a simple system. This system will collect user interaction data, identify and expand user interest points through big data analysis techniques (such as clustering algorithms), and form feature vectors.

[0287] Here is a complete Java code example showing how to build such a system:

[0288] 1. Collect user interaction data.

[0289] 2. Use K-Means clustering algorithm to perform cluster analysis on user data. Specifically, use Apache CommonsMath library to perform K-Means cluster analysis.

[0290] 3. Update user portraits based on clustering results.

[0291] 4. Form the feature vector.

[0292] This system has a huge learning resource database, covering various types of educational resources, such as online courses, e-books, video tutorials, etc. This database is carefully selected and classified, and creating a vector representation describing the content characteristics of each learning resource (such as a book, a class) is one of the key steps in building an efficient personalized recommendation system. This process usually involves natural language processing (NLP), machine learning, and deep learning techniques to ensure that the generated vector can accurately reflect the core characteristics of the resource and support subsequent similarity calculation and recommendation generation.

[0293] In order to create a database of learning resources and generate vector representations for each resource that describe its content characteristics, we can use word embedding models in natural language processing technology (such as Word2Vec, GloVe) or more advanced pre-trained models (such as BERT). In this example, we will use Java and BERT to implement this function. To simplify the example, we will assume that we already have a pre-trained BERT model and can get the vector representation of the text through API calls.

[0294] In order to achieve personalized recommendations based on user interest profiles, the use of a hybrid recommendation system is indeed a very effective strategy. This system combines the advantages of collaborative filtering and content-based recommendations, and is supplemented by advanced deep learning technology, which can significantly improve the recommendation effect. The following are specific methods and technical details:

[0295] Construction of hybrid recommendation system:

[0296] Collaborative Filtering:

[0297] User-Item Matrix: In step 6, a rating matrix is ​​established that contains all users and the resources (such as books, courses, and activities) they have interacted with.

[0298] Similarity calculation: Use methods such as cosine similarity and Pearson correlation coefficient to calculate the similarity between users or items.

[0299] Neighbor selection: Find a set of users (K nearest neighbors) that are most similar to the target user and predict the target user’s preferences based on the behavior of these neighbors.

[0300] Recommendation generation: Provide users with a personalized list of recommendations based on the choices of similar users or the relevance of similar items.

[0301] Content-based Recommendation:

[0302] Resource feature vector: For each resource (such as a book or a class), create a vector representation that describes the characteristics of its content.

[0303] User preference model: Based on user profiles and historical behaviors, a preference model is established for each user to capture their interests in different types of resources.

[0304] Matching score calculation: Calculate the similarity between the resource feature vector and the user preference model as the recommendation score.

[0305] Recommendation sorting: Sort candidate resources by score and select the top few as the final recommendation results.

[0306] Mixed recommendations:

[0307] Weighted average method: Simply add the results of collaborative filtering and content-based recommendation according to certain weights to get a comprehensive score.

[0308] Stacked model: Train a machine learning model (such as Random Forest, XGBoost) with the output of two recommendation methods as input features to predict the final recommendation score.

[0309] Fusion network: Design a neural network structure that accepts the embedded representation of users and items at the same time, and outputs the recommendation score after multiple layers of nonlinear transformation.

[0310] Examples of actual application scenarios:

[0311] Suppose an old lady named Grandma Wang often communicates with this system, expresses her love for gardening, and occasionally mentions her desire to learn new things. This system will handle it like this:

[0312] 1. Build a user interest profile: By analyzing Grandma Wang’s conversation records, we found that she is interested in topics such as "gardening" and "handmade", and her emotions are more positive whenever she talks about these.

[0313] 2. Hybrid Recommendation System:

[0314] Collaborative filtering: The system found that other elderly people who also like gardening also participated in the painting class, so it suggested that Grandma Wang try it as well.

[0315] Content-based recommendations: Based on the content that Grandma Wang has browsed in the past, some professional books and video tutorials on plant care are recommended.

[0316] Fusion Networking: Combining the two approaches above, this system also suggests activities that fit her existing interests while also expanding her new skills, such as attending craft shows organized by the community.

[0317] In this way, the system can not only understand and remember Grandma Wang’s existing interests and hobbies, but also constantly guide her to explore new possibilities and enrich her daily life experience, while ensuring that the recommended content always remains fresh and attractive.

[0318] Finally, the system uses powerful AI language models to integrate all the above information and generate personalized recommendation lists. These models have strong semantic understanding and generation capabilities, and can synthesize the most suitable activities, social circles or learning resource suggestions for the elderly according to their specific circumstances. For example, if the elderly are interested in history, the system may recommend a documentary about ancient civilizations; if they are calligraphy enthusiasts, they may introduce an online calligraphy class. In this way, the system not only becomes a good partner in the lives of the elderly, but also helps them explore the new world and improve their quality of life.

[0319] In order to realize a conversation content generation system that combines user voice text and recommended resource data, we can use the Tongyi Qianwen AI language model to generate conversation content. We will simulate a simple scenario in which Tongyi Qianwen generates corresponding conversation content based on the user's voice text and recommended resource data.

[0320] Here is a complete Java code example showing how to implement this functionality:

[0321] 1.Simulate user voice text input.

[0322] 2. Simulate recommended resource data.

[0323] 3. Use the Tongyi Qianwen AI language model to generate dialogue content.

[0324] In summary, the present invention is based on Paraformer speech recognition, converting the collected speech samples into text form; constructing an emotional model and performing user emotion recognition based on the emotional model; setting emotional support strategies and voice output scripts, collecting time, environmental data and user emotions in real time, and starting corresponding emotional support strategies and voice output scripts; using Alibaba Cloud CosyVoice's voice replication technology to simulate the voices of users' relatives for voice output during user interaction; based on the natural language processing model, identifying users' favorite topics, activity types and potential emerging fields, and recommending related content. In short, the present invention provides unprecedented care and support to the elderly at the emotional level, allowing every elderly person in the retirement community to feel the warmth and companionship brought by technology.

[0325] Those of ordinary skill in the art will appreciate that the units of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition of each example has been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0326] In the embodiments provided by the present invention, it should be understood that the division of units is only a logical function division, and there may be other division methods in actual implementation, for example, multiple units can be combined into one unit, one unit can be split into multiple units, or some features can be ignored, etc.

[0327] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0328] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-0nlyMemory), random access memory (RAM, RandomAccessMemory), mobile hard disk, magnetic disk or optical disk, etc., which can store program code.

[0329] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. These modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be included in the scope of the claims and specification of the present invention.

Claims

1. An intelligent companion system based on semantic recognition, characterized in that: include: A voice collection unit, used to collect voice signals input during user interaction; A text recognition unit, connected to the voice collection unit, is used to convert the voice signal into text for storage; The emotion recognition unit is connected to the text recognition unit and is used to train an emotional model to perform user emotion recognition after collecting and annotating user data; The strategy generation unit uses the pre-set words for voice output based on the pre-set emotional support strategy. In addition, the strategy generation unit is also connected to the emotion recognition unit, and also activates the pre-set strategy based on the user's emotion and uses the pre-set words for voice output.

2. According to claim 1, the intelligent companionship system based on semantic recognition is characterized in that: The specific operations of the emotion recognition unit are as follows: Emotional annotation is performed on the text transcribed from the speech data to provide necessary supervision information for model training; Store the text and sentiment annotation results of speech transcription into big data for AI training and analysis; Choose the appropriate model architecture according to the specific application scenario, such as Transformer, RNN, etc.; Set model parameters, including the number of layers, number of hidden units, and learning rate, to optimize model performance; The Adam optimization algorithm is used to train the model.

3. According to claim 1, the intelligent companionship system based on semantic recognition is characterized in that: Emotional support strategies include personalized greetings, weather reminders, holiday greetings, empathetic responses, and emergency calls.

4. The intelligent companionship system based on semantic recognition according to claim 1, characterized in that: The emotional support strategy is selected by defining a rule set. The process of defining a rule set is as follows: Fixed time trigger: set daily, weekly or monthly reminders; Situational awareness: Determine whether a strategy needs to be initiated based on information provided by environmental sensors; User input response: When a user expresses a specific keyword or a specific emotion is detected, relevant policies are triggered immediately.

5. The intelligent companionship system based on semantic recognition according to claim 1, characterized in that: It also includes a voice replication unit, which uses the voice replication technology of Alibaba Cloud CosyVoice to simulate the voices of the user's relatives for voice output during user interaction.

6. The intelligent companionship system based on semantic recognition according to claim 1, characterized in that: It also includes an interest identification and recommendation unit, which identifies the user's favorite topics, activity types, and potential emerging fields based on a natural language processing model, and recommends related content.

7. The intelligent companionship system based on semantic recognition according to claim 6, characterized in that: The process of interest identification is as follows: Use BERT as a pre-trained model and fine-tune it based on the characteristics and needs of elderly people's communication. Fine-tuning includes adding health care and family life corpora specific to the lives of the elderly. Using named entity recognition technology, we can automatically extract entity information such as names, places, and times from conversations. At the same time, we can combine word frequency statistics to find high-frequency words as possible interest markers. Clustering a large amount of conversation text into several topic domains through Latent Dirichlet Allocation to find the themes or topics that users repeatedly mention; With the help of deep learning sentiment classifier, the emotional tendency in each conversation is evaluated to help determine which topics can make users feel happy or resonate; For each identified topic or activity type, assign one or more tags to facilitate subsequent management and retrieval; Through association rule mining or other recommendation algorithms, we can explore the intrinsic connections between different interests and predict new areas that users may be interested in but are not yet aware of.

8. An intelligent companionship method based on semantic recognition, characterized in that: The system applied to any one of claims 1 to 7 comprises the following steps: Collecting voice signals input during user interaction and performing data cleaning, wherein the voice signals include voice samples with different accents, speaking speeds, and intonations; Based on Paraformer speech recognition, the collected speech samples are converted into text; Build an emotional model and identify user emotions based on the emotional model; Set emotional support strategies and voice output scripts, collect time, environmental data and user emotions in real time, and start corresponding emotional support strategies and voice output scripts; Adopting the voice replication technology of Alibaba Cloud CosyVoice, the voice of the user's relatives is simulated for voice output during user interaction; Based on natural language processing models, we identify users’ favorite topics, activity types, and potential emerging fields, and recommend relevant content.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the intelligent companionship method based on semantic recognition as described in claim 8.

10. A processor, characterized in that: The processor is used to run a program, wherein the program, when running, executes the intelligent companionship method based on semantic recognition described in claim 8.

Citation Information

Cited By

  • Adaptive interaction method based on multi-modal emotion calculation and robot

    CN120631175A

  • Intelligent accompanying system and method and intelligent terminal

    CN121214938A