Intelligent dialogue method and system for accompanying old people based on user portraits

By using multimodal data processing and prefix tree analysis, a user profile of elderly interaction sequences is generated, which solves the problem of insufficient user feature representation in existing technologies and achieves accurate capture of elderly interaction patterns and efficient recognition of intent.

CN120998202AActive Publication Date: 2025-11-21KUAISHANGYUN (SHANGHAI) NETWORK TECHNOLOGY CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511525378.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2025-11-21
Estimated Expiration
2045-10-23

AI Technical Summary

Technical Problem

Existing intelligent dialogue methods for elderly companionship fail to effectively combine deep indicators such as heart rate variability, walking ratio, and application usage frequency, resulting in insufficient dimensions of user feature representation, making it difficult to capture the regularity of the elderly's daily interaction patterns, and lacking accurate extraction of long-term interaction habits and potential intentions.

Method used

By collecting multimodal data, preprocessing it, calculating user features, generating user profile vectors, constructing interaction sequences and generating prefix trees, extracting candidate pattern sets, calculating similarity and confidence, selecting the optimal intent, and loading a pre-trained language model to generate dialogue sentences.

Benefits of technology

It enhances the ability to capture long-term interaction patterns of the elderly, improves the accuracy of intent inference, and generates dialogue content that better matches the interests and needs of the elderly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120998202A_ABST
    Figure CN120998202A_ABST
Patent Text Reader

Abstract

The invention discloses an elder accompanying intelligent dialogue method and system based on a user portrait, and relates to the technical field of intelligent interaction, and the method comprises the steps: calculating user features based on multi-modal data, splicing the user features into a user portrait vector, generating a frame sequence from voice, calculating an interaction feature vector, constructing an interaction sequence, generating a prefix tree, and extracting all paths. Calculating a support degree, and generating a candidate mode set; extracting an intention label corresponding to each mode in the candidate mode set, generating a frequent mode set, calculating similarity, generating an associated intention label and confidence, extracting frequency features in user portrait vectors, splicing the frequency features into interest feature sub-vectors, calculating joint utility, and selecting a maximum value as an optimal intention. According to the method, through combination of the prefix tree and frequent pattern mining and interaction sequence clustering, the capability of capturing long-term interaction behavior rules of old people is enhanced, and through combination of interest utility and matching utility, the accuracy of intention inference is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent interaction, in particular to an old person accompanying intelligent conversation method and system based on user portrait. BACKGROUND

[0002] With the continuous acceleration of the process of social aging, intelligent accompanying technology has gradually become an important research direction in the academic and industrial circles. The popularity of intelligent wearable devices and intelligent interaction systems makes health monitoring and emotional care based on multi-modal data gradually have technical implementation conditions. In the field of intelligent conversation, early systems mainly rely on keyword matching or template library retrieval, which is difficult to realize personalized interaction. With the development of natural language processing, speech analysis and user portrait technology, intelligent accompanying systems based on semantic understanding gradually appear, which can provide more natural conversation feedback according to the user's historical behavior and interest preferences to a certain extent.

[0003] The existing old person accompanying intelligent conversation method still has deficiencies. Most systems only perform simple normalization or threshold judgment on original sensor data, and fail to construct comprehensive features in combination with heart rate variability, walking proportion, application use frequency and other deep indicators, resulting in insufficient user feature representation dimension. The existing method lacks structured mining of interaction sequences, and it is difficult to capture the regularity of the daily interaction mode of the elderly. There is a lack of deep analysis based on prefix tree or frequent pattern set, resulting in inaccurate extraction of long-term interaction habits and potential intentions. SUMMARY

[0004] In view of the above existing problems, the present application is proposed.

[0005] Therefore, the present application provides an old person accompanying intelligent conversation method and system based on user portrait, which solves the problem that most systems only perform simple normalization or threshold judgment on original sensor data, fail to construct comprehensive features in combination with heart rate variability, walking proportion, application use frequency and other deep indicators, resulting in insufficient user feature representation dimension. The existing method lacks structured mining of interaction sequences, and it is difficult to capture the regularity of the daily interaction mode of the elderly. There is a lack of deep analysis based on prefix tree or frequent pattern set, resulting in inaccurate extraction of long-term interaction habits and potential intentions.

[0006] To solve the above technical problems, the present application provides the following technical solutions: In a first aspect, the present application provides an old person accompanying intelligent conversation method based on user portrait, which comprises, Collecting multi-modal data and performing preprocessing, calculating user features based on multi-modal data, splicing into user portrait vectors, generating speech frames, calculating interaction feature vectors, constructing interaction sequences, generating prefix trees, extracting all paths, calculating support, and generating candidate pattern sets; Extract the corresponding intent label of each mode in the candidate mode set, generate the frequent mode set, and calculate the similarity, generate the associated intent label and the confidence, extract the frequency feature in the user portrait vector, splice into the interest feature subvector, calculate the joint utility, and select the maximum value as the optimal intent; Extract the text corresponding to the optimal intent, convert it into a TF-IDF vector, calculate the semantic score, generate the final comparison segment, load the pre-trained language model to generate a dialogue sentence, and have an intelligent dialogue with the old people.

[0007] As a preferred scheme of the old people accompanying intelligent dialogue method based on user portrait, the method comprises the steps of: Collect multi-modal data of the old people through the smart bracelet and the smart phone, and perform denoising and normalization processing; The multi-modal data comprises heartbeat interval, step count, sleep duration, walking time, wake-up time, device usage frequency, voice and interaction data; Calculate the user features based on the multi-modal data, and splice the user features into a user portrait vector by using a vector splicing method; Sort the voice according to time sequence to generate a frame sequence; Use the delay coordinate method to construct a phase space vector for the frame sequence, and calculate Lyapunov exponent and fractal dimension based on the phase space vector; Calculate short-time energy based on the frame sequence; Splice the Lyapunov exponent, fractal dimension and short-time energy into a frame feature vector, and calculate the mean value to define an interaction feature vector; Collect historical interaction data to construct an interaction sequence; Use K-means clustering to cluster the interaction feature vectors in the interaction sequence to generate state triplets, and construct a quantized interaction sequence; Initialize an empty prefix tree, the root node is empty, insert the state triplets of the interaction sequence into the tree according to time sequence, each node stores a triplet, the edge represents time transition, and the node records the number of occurrences to generate a prefix tree; Use depth-first search to traverse the prefix tree, extract all paths, calculate the support, perform threshold screening on the support, and generate a candidate mode set.

[0008] As a preferred scheme of the old people accompanying intelligent dialogue method based on user portrait, the method comprises the steps of: For each mode in the candidate mode set, the corresponding intent label is extracted from the historical interaction data according to the timestamp of the quantized interaction sequence to generate a frequent mode set. A time window is set using a sliding window method, interaction records in the time window are selected from the interaction sequence, sorted in ascending order of timestamp, and a recent interaction feature set is generated; An interaction feature vector corresponding to the recent interaction feature set is extracted, the recent interaction feature set is added, and a subsequence is generated; The recent interaction feature vector machine in the subsequence is clustered using K-means clustering, a state triple is generated, and a quantized subsequence is constructed; Based on each mode in the frequent pattern set and the quantized subsequence, the similarity is calculated, the mode with the maximum similarity is selected, and the corresponding associated intent label is obtained; Based on the similarity, the confidence is calculated.

[0009] As a preferred scheme of the old person accompanying intelligent dialogue method based on user portrait, wherein: the calculation of joint utility and the selection of maximum value as the optimal intent comprises: The frequency features in the user portrait vector are extracted and spliced into an interest feature subvector, and the interest utility of the intent is calculated; The interaction data corresponding to the associated intent label and the interest label are converted into TF-IDF vectors, and the cosine similarity is calculated, which is defined as the matching utility of the intent; The mean value of the interest utility and the matching utility is calculated to obtain the joint utility, and the maximum value is selected as the optimal intent.

[0010] As a preferred scheme of the old person accompanying intelligent dialogue method based on user portrait, wherein: the extraction of the text corresponding to the optimal intent, the conversion into a TF-IDF vector, the calculation of a semantic score, and the generation of a final comparison segment comprise: The text corresponding to the optimal intent is extracted, converted into a TF-IDF vector, and the semantic similarity with the text segment is calculated; Filtering segments with a similarity greater than a similarity threshold to form a candidate segment subset; Each candidate segment in the candidate segment subset is extracted, and a semantic score is calculated; The dialogue segment with the maximum semantic score is selected as the final comparison segment.

[0011] As a preferred scheme of the old person accompanying intelligent dialogue method based on user portrait, wherein: the loading of a pre-trained language model to generate a dialogue sentence and the intelligent dialogue with the old person comprise: The text corresponding to the final comparison segment is extracted, a pre-trained language model is loaded, and a dialogue sentence is generated; The dialogue sentence is converted into speech and the intelligent dialogue with the old person is carried out.

[0012] As a preferred scheme of the old person accompanying intelligent conversation method based on user portrait provided in the application, wherein the collection of multi-modal data and preprocessing comprises: The multi-modal data of the old people are collected through the smart bracelet and the smart phone, and are subjected to denoising and normalization processing. The multi-modal data comprise heartbeat interval, step count, sleep duration, walking time, wake-up time, device usage frequency, voice and interaction data.

[0013] In a second aspect, the application provides an old person accompanying intelligent conversation method system based on user portrait, comprising, A collection and processing module is configured to collect multi-modal data of the old people and to perform denoising and normalization processing. A portrait clustering module is configured to calculate user features, splice the user features into a user portrait vector, generate a state triple through K-means clustering, construct a prefix tree and generate a candidate mode set. A matching intention module is configured to extract and screen intention labels from the candidate modes, optimize intention inference by calculating interest utility and matching utility, and select an optimal intention for response. A semantic score module is configured to screen candidate text segments, select the most matched conversation segment based on semantic similarity and semantic score, and perform intelligent conversation.

[0014] In a third aspect, the application provides a computer device comprising a memory and a processor, and the memory stores a computer program, wherein the computer program is executed by the processor to implement any step of the old person accompanying intelligent conversation method based on user portrait according to the first aspect of the application.

[0015] In a fourth aspect, the application provides a computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement any step of the old person accompanying intelligent conversation method based on user portrait according to the first aspect of the application.

[0016] The application has the following beneficial effects: the application combines the prefix tree and the frequent pattern mining to cluster the interaction sequences, thereby enhancing the ability to capture the long-term interaction behavior rules of the old people, and the interest utility and the matching utility are combined to improve the accuracy of intention inference. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0018] Figure 1 A flowchart of the user portrait-based old person accompanying intelligent dialogue method in Embodiment 1; Figure 2 A schematic diagram of the user portrait-based old person accompanying intelligent dialogue system in Embodiment 1. DETAILED DESCRIPTION

[0019] In order to make the above objectives, characteristics and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0020] In the following description, a large number of specific details are set forth in order to facilitate a thorough understanding of the present application, but the present application can also be implemented in other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of the present application, therefore the present application is not limited by the specific embodiments disclosed below.

[0021] Secondly, the "one embodiment" or "embodiment" referred to herein means that the specific features, structures or characteristics can be included in at least one implementation of the present application. "In one embodiment" appearing in different places in the specification does not mean the same embodiment, nor is it an embodiment that is separate or alternative to other embodiments.

[0022] Embodiment 1, with reference to Figure 1 and Figure 2 , the first embodiment of the present application provides a user portrait-based old person accompanying intelligent dialogue method, comprising the following steps: S1, collecting multi-modal data and pre-processing, calculating user features based on multi-modal data, splicing into user portrait vector, generating frame sequence of speech, calculating interaction feature vector, constructing interaction sequence, generating prefix tree, extracting all paths, calculating support, and generating candidate mode set; Specifically, collecting multi-modal data and pre-processing includes: Collecting multi-modal data of the elderly through a smart bracelet and a smart phone, and performing denoising and normalization processing; The multi-modal data includes heartbeat interval, step count, sleep duration, walking time, wake-up time, device usage frequency, voice, and interaction data.

[0023] Multi-dimensional data breaks through the limitations of traditional single-modal portrait, so that the user portrait not only has static information, but also can reflect dynamic changes. Data preprocessing eliminates environmental noise and scale difference, enhances the stability of the feature vector, and is beneficial to the reliability improvement in subsequent pattern mining.

[0024] Further, constructing an interaction sequence, generating a prefix tree, extracting all paths, calculating a support, and generating a candidate mode set include: Collecting multi-modal data of the elderly through smart bracelets and smart phones, and carrying out denoising and normalization processing; The multi-modal data includes heartbeat interval, step count, sleep duration, walking time, wake-up time, device usage frequency, voice, and interaction data; Calculate user features based on multi-modal data; Calculate the standard deviation of the heartbeat interval, defined as heart rate variability, and perform normalization processing to generate HRV features, the formula is: , wherein is the HRV feature, is the heart rate variability, is the median of the heart rate variability; Use the limit function to calculate the step count feature and sleep feature respectively, the formula is: , , wherein is the step count feature, is the sleep feature, and are limit functions, and are the step count normalization upper limit and the sleep duration normalization upper limit respectively, and are the step count and sleep duration in the last 24 hours respectively; Based on the walking time and the wake-up time, calculate the walking time proportion, the formula is: , wherein is the walking time proportion, is the walking time, is the wake-up time; Based on the device usage frequency, calculate the application usage frequency, the formula is: , wherein is the othapplication usage frequency, is the othapplication usage frequency, is the usage frequency normalization upper limit; Use the direct assignment method to define the walking time proportion and the application usage frequency as the walking time proportion feature and the application usage frequency feature; Collect historical interaction data with labels, convert to text data and perform text cleaning, and calculate word vectors, manually define initial keyword vector set for each label, calculate word vector intersection of each word vector with initial keyword set, and assign topic label, get annotated corpus, formula is: , , , wherein is the intersection size of the word vector set and the keyword vector set, is the initial keyword vector set of the th interest label, is the word vector set of the th text, is the topic label of the th text, is the annotated corpus, is the number of texts; Calculate the TF-IDF vector of the interaction data using the TF-IDF method, extract the keyword vector of the keyword set in the annotated corpus, calculate the cosine similarity, and assign the interest label, formula is: , wherein is the interest label assigned to the th interaction, and are the TF-IDF vector of the th interaction and the word vector of the th keyword set, respectively, is the cosine similarity; Calculate the frequency of each interest label and normalize it to generate the frequency feature, formula is: , wherein is the frequency feature of the interest label, is the interest label data, is the interest label assigned to the th interaction; Use vector concatenation method to concatenate user features into user portrait vector, formula is: , wherein F is the user portrait vector, and are the walking time proportion feature and the application usage frequency feature, respectively; Sort the voice in chronological order to generate frame sequence; The phase space vector for the frame sequence is constructed using the delayed coordinate method, and the formula is as follows: , in Let be the phase space vector of the i-th frame. For the audio data of the i-th frame, The delay time is determined by the autocorrelation function. is the start time of the i-th frame; Based on the phase space vector, the Lyapunov exponent and fractal dimension are calculated using the following formula: , , , in The Lyapunov index, The total number of trajectory points. Let be the phase space vector of the j-th frame. for The nearest neighbor vector, For correlation integrals, the proportion of phase space vector point pairs with a distance less than r is calculated, reflecting the distance distribution of phase space vector point pairs, where r is the distance threshold. Let be the phase space vector of the k-th frame. For the Heaviside function, Where j is the fractal dimension and k is the vector index; Based on the frame sequence, the short-time energy is calculated using the following formula: , in Let i be the short-time energy of the i-th frame. This represents the number of intra-frame sampling points. This represents the value of the nth sampling point in the i-th frame, corresponding to the speech data, where n is the sampling point index; The Lyapunov exponent, fractal dimension, and short-time energy are concatenated into a frame feature vector, and the mean is calculated and defined as the interaction feature vector, as shown in the formula: , , in Let i be the feature vector of the i-th frame. Here, M represents the interaction feature vector, and M represents the number of frames. Collect historical interaction data and construct an interaction sequence using the following formula: , in It is an interaction sequence, containing interaction feature vectors and timestamp pairs. For the first the second interaction timestamp, is the interaction index, and t is the current timestamp; The interaction feature vectors in the interaction sequence are clustered using K-means clustering to generate state triplets, and a quantitative interaction sequence is constructed, with the formula being: , wherein is the quantitative interaction sequence, , and is the quantitative state label of the ith interaction (Lyapunov exponent, fractal dimension, and short-time energy), indicating the state triplet; The prefix tree is initialized with an empty root node, and the state triplets of the interaction sequence are inserted into the tree in chronological order. Each node stores a triplet, the edge represents the time transition, and the node records the number of occurrences. The prefix tree is generated; The prefix tree is traversed using depth-first search to extract all paths, each path representing a sequence of state triplets. The support is calculated with the formula: , wherein is the support of the path , is the number of occurrences of the path end node, is the set of all paths, is any path in the set of all paths; Threshold screening is performed on the support to generate a candidate mode set, with the formula being: , wherein is the candidate mode set, is the support threshold.

[0025] The application introduces a limiting function for processing, so that the numerical value is mapped to a reasonable interval range, which avoids the interference of extreme values on the user portrait and ensures the balance between different features. Compared with the simple keyword matching method, the TF-IDF combined with the weight distribution in the corpus can avoid the interference of high-frequency but meaningless words, so as to more accurately reflect the user interest. Combined with semantic similarity calculation (cosine similarity), high-precision matching of interactive text and interest labels can be realized, and the accuracy of the interest portrait is improved. By introducing dynamic features (Lyapunov index and fractal dimension), the voice signal is not only described as a spectrum or energy distribution, but also can capture the dynamic evolution characteristics of the voice. Combined with short-time energy calculation, a multi-angle description of the voice signal can be formed, which not only contains the energy information in the time domain, but also contains complex dynamic characteristics, realizing more comprehensive voice modeling. The generated interactive feature vector not only reflects the properties of single voice interaction, but also captures the voice mode preference of the elderly in a statistical sense, which helps to improve the stability of subsequent intent recognition. The prefix tree can clearly show the time sequence law of the user interaction state while ensuring the storage efficiency. The support filtering generates a candidate mode set, which can effectively filter out low-frequency modes with strong randomness and high noise, and retain core behavior patterns that are highly representative of the user portrait. Not only does it improve the efficiency of pattern matching, but it also provides explainability support for the intelligent companion system.

[0026] S2, extracting the intent label corresponding to each mode in the candidate mode set, generating a frequent mode set, calculating the similarity, generating the associated intent label and confidence, extracting the frequency feature in the user portrait vector, splicing into an interest feature sub-vector, calculating the joint utility, and selecting the maximum value as the optimal intent; Specifically, the intent label corresponding to each mode in the candidate mode set is extracted, a frequent mode set is generated, the similarity is calculated, the associated intent label and confidence are generated, and the frequency feature in the user portrait vector is extracted, which is spliced into an interest feature sub-vector, the joint utility is calculated, and the maximum value is selected as the optimal intent. For each mode in the candidate mode set, the corresponding intent label is extracted from the historical interaction data according to the timestamp of the quantized interaction sequence, a frequent mode set is generated, and the formula is: , , Among them is the intent label associated with the mode path in the candidate mode set, is the intent label extracted from the historical interaction data, is the number of occurrences of the mode corresponding intent, is the frequent mode set, is the frequent mode sequence; A sliding window method is used to set a time window, and the interaction records within the time window are selected from the interaction sequence, sorted in ascending order of timestamp, and a recent interaction feature set is generated. Extract the interaction feature vector corresponding to the recent interaction feature set, add the recent interaction feature set, generate a sub-sequence, and the formula is: , Where is the sub-sequence, is the recent interaction feature set, is the recent interaction feature vector, is the recent interaction data timestamp; Use K-means clustering to cluster the recent interaction feature vectors in the sub-sequence to generate state triplets and construct a quantized sub-sequence, and the formula is: , Where is the quantized sub-sequence, , and is the quantized state label of the first interaction (Lyapunov exponent, fractal dimension, and short-time energy), representing the state triplet, and are the first interaction feature vector and time window, respectively; Based on each pattern in the frequent pattern set and the quantized sub-sequence, calculate the similarity, and the formula is: , Where is the sequence similarity, is the first frequent pattern; Select the pattern with the maximum similarity to obtain the corresponding associated intent label, and the formula is: , Where I is the associated intent label, is the pattern index, is the sequence similarity; Based on the similarity, calculate the confidence, and the formula is: , Where is the confidence.

[0027] Intent annotation enables the association of low-level state triple sequences with high-level semantic labels. By using a sliding window, interaction data within a specific time range can be extracted from the interaction sequence, ensuring the system's time sensitivity and dynamic perception of recent changes in the user's state. This avoids the lag caused by relying solely on long-term patterns and avoids relying solely on label statistics. Instead, it improves the accuracy of real-time intent recognition through dynamic matching. Sequence similarity calculation can tolerate certain temporal perturbations, improving robustness to complex interactive behaviors. By selecting the maximum similarity, it ensures that the output results have the highest relevance to the user's recent state, reducing misjudgments.

[0028] Furthermore, the joint utility is calculated, and the maximum value is selected as the optimal intention, including: Extract frequency features from the user profile vector, concatenate them into an interest feature sub-vector, and calculate the interest utility of the intent using the following formula: , in For the first The interest utility of an intention Let u be the interest feature sub-vector. For the first Sub-vectors of secondary interest features; The interaction data corresponding to the associated intent tags and interest tags are converted into TF-IDF vectors, and the cosine similarity is calculated, defined as the intent matching utility, with the formula as follows: , in For the first The effectiveness of matching an intent. and The first The TF-IDF vectors of the i-th intention and the i-th intention; Calculate the mean of interest utility and matching utility to obtain the joint utility, and select the maximum value as the optimal intention.

[0029] Interest utility characterizes users' long-term interest tendencies, while matching utility characterizes immediate semantic needs. The combination of the two forms a comprehensive metric that combines long-term and short-term needs. TF-IDF and cosine similarity methods can capture the semantic proximity between intent and interest, improving cross-modal semantic alignment capabilities. Through the mean mechanism of joint utility, the integration of interest preferences and immediate interaction is achieved.

[0030] S3. Extract the text corresponding to the optimal intent, convert it into a TF-IDF vector, calculate the semantic score, generate the final comparison fragment, load the pre-trained language model to generate dialogue sentences, and conduct intelligent dialogue with the elderly. Specifically, the text corresponding to the optimal intention is extracted, converted into a TF-IDF vector, a semantic score is calculated, and a final comparison segment is generated, including: The text corresponding to the optimal intention is extracted, converted into a TF-IDF vector, and a semantic similarity with the text segment is calculated, and the formula is: , Wherein is the semantic similarity of the segment and the optimal intention, is the i-th dialogue segment, is the optimal intention index, and are the TF-IDF vectors of the optimal intention and the dialogue segment, respectively; Filtering segments with a similarity greater than a similarity threshold to form a candidate segment subset, and the formula is: , Wherein is the candidate segment subset, is the similarity threshold; Extracting each candidate segment in the candidate segment subset, calculating a semantic score, and the formula is: , Wherein is the semantic score of the dialogue segment, is the TF-IDF vector of the current intention label, is the joint utility of the optimal intention; Selecting the dialogue segment with the maximum semantic score as the final comparison segment.

[0031] Through the TF-IDF method, the weight of the keywords in the text can be quantified to ensure the semantic consistency of the candidate segment and the user's intention. Through the similarity screening mechanism, irrelevant segments that do not match the user's needs are automatically excluded, improving the efficiency of corpus matching and ensuring that the final generated dialogue content not only has correct grammar but also matches the needs of the elderly users at the semantic level. The introduction of the joint utility parameter makes the semantic score not only consider semantic similarity but also reflect the user's long-term interest preferences, allowing the differentiation of segments with similar semantics but significantly different user interests, and the preferential selection of content with high interest matching degree.

[0032] Further, a pre-trained language model is loaded to generate dialogue sentences for intelligent dialogue with the elderly, including: Extracting the text corresponding to the final comparison segment, loading a pre-trained language model (such as BERT variants), and generating dialogue sentences; Converting the dialogue sentences into speech and conducting intelligent dialogue with the elderly.

[0033] The pre-trained language model has strong semantic modeling capability and can generate fluent, natural and contextually appropriate dialogue sentences. Compared with fixed templates, the model can generate sentences with different styles and expressions, avoiding repetition and mechanical feeling. The voice output can enhance the authenticity and warmth of the interaction, making the elderly feel accompanied.

[0034] The embodiment also provides an old person accompanying intelligent dialogue method and system based on user portrait, comprising: A collection and processing module is configured to collect multi-modal data of the old person and perform denoising and normalization processing. A portrait clustering module is configured to calculate user features, splice the user features into a user portrait vector, generate a state triple through K-means clustering, construct a prefix tree, and generate a candidate mode set. A matching intent module is configured to extract and filter intent labels from the candidate mode, optimize intent inference by calculating interest utility and matching utility, and select an optimal intent for response. A semantic score module is configured to filter candidate text segments, select the most matched dialogue segment based on semantic similarity and semantic score, and perform intelligent dialogue.

[0035] The embodiment also provides a computer device suitable for the old person accompanying intelligent dialogue method based on user portrait, comprising a memory and a processor. The memory is configured to store computer executable instructions, and the processor is configured to execute the computer executable instructions to implement the old person accompanying intelligent dialogue method based on user portrait as described in the above embodiment.

[0036] The computer device can be a terminal, and the computer device comprises a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is configured to communicate with external terminals in a wired or wireless manner. The wireless manner can be achieved through WIFI, operator network, NFC (near field communication) or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad arranged on the shell of the computer device. In addition, the input device can be an external keyboard, touchpad or mouse, etc.

[0037] The embodiment also provides a storage medium, which stores a computer program, and the computer program is executed by a processor to implement the method for implementing the intelligent conversation of the old person accompanying based on the user portrait as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk or an optical disk.

[0038] In conclusion, the application enhances the ability to capture the long-term interactive behavior of the elderly by combining the prefix tree with the frequent pattern mining and the interactive sequence clustering, and improves the accuracy of the intention inference by combining the interest utility with the matching utility.

[0039] It should be noted that the above embodiments are only used to illustrate the technical solutions of the application but not limit the application, and although the application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the application can be modified or replaced equivalently without departing from the spirit and scope of the technical solutions of the application, and all of them should be covered in the scope of the claims of the application.

Claims

1. A method for intelligent dialogue for elderly companionship based on user profiles, characterized in that: include, Collect and preprocess multimodal data, calculate user features based on multimodal data, concatenate them into user profile vectors, generate frame sequences from speech, calculate interaction feature vectors, construct interaction sequences, generate prefix trees, extract all paths, calculate support, and generate candidate pattern sets. Extract the intent label corresponding to each pattern in the candidate pattern set, generate a frequent pattern set, calculate the similarity, generate associated intent labels and confidence, extract frequency features from the user profile vector, concatenate them into interest feature sub-vectors, calculate the joint utility, and select the maximum value as the optimal intent. Extract the text corresponding to the optimal intent, convert it into a TF-IDF vector, calculate the semantic score, generate the final comparison fragment, load the pre-trained language model to generate dialogue sentences, and conduct intelligent dialogue with the elderly.

2. The intelligent dialogue method for elderly companionship based on user profiles as described in claim 1, characterized in that: The process of constructing the interaction sequence, generating a prefix tree, extracting all paths, calculating support, and generating a candidate pattern set includes: Multimodal data of the elderly is collected through smart bracelets and smartphones, and then denoised and normalized. The multimodal data includes heart rate interval, step count, sleep duration, walking time, awake time, device usage frequency, voice, and interaction data; User features are calculated based on multimodal data, and user features are concatenated into user profile vectors using vector concatenation. The audio is sorted chronologically to generate a frame sequence; The phase space vector is constructed for the frame sequence using the delayed coordinate method, and the Lyapunov exponent and fractal dimension are calculated based on the phase space vector. Calculate short-time energy based on frame sequences; The Lyapunov exponent, fractal dimension, and short-time energy are concatenated into a frame feature vector, and the mean is calculated and defined as the interaction feature vector. Collect historical interaction data and construct interaction sequences; K-means clustering is used to cluster the interaction feature vectors in the interaction sequence to generate state triples and construct a quantized interaction sequence. Initialize an empty prefix tree with an empty root node. Insert the state triples of the interaction sequence into the tree in chronological order. Each node stores the triples, edges represent time transitions, and nodes record the occurrence counts to generate the prefix tree. The prefix tree is traversed using a depth-first search to extract all paths, the support is calculated, and the support is filtered by a threshold to generate a candidate pattern set.

3. The intelligent dialogue method for elderly companionship based on user profiles as described in claim 2, characterized in that: The process of extracting the intent tag corresponding to each pattern in the candidate pattern set, generating a frequent pattern set, calculating similarity, and generating associated intent tags and confidence scores includes: For each pattern in the candidate pattern set, extract the corresponding intent label from historical interaction data based on the timestamp of the quantized interaction sequence to generate a frequent pattern set; The sliding window method is used to set a time window, and the interaction records within the time window are selected from the interaction sequence, sorted in ascending order of timestamp, to generate a recent interaction feature set; Extract the interaction feature vectors corresponding to the recent interaction feature set, add them to the recent interaction feature set, and generate subsequences; K-means clustering is used to cluster the recent interaction feature vectors in the subsequences to generate state triples and construct quantized subsequences; Based on each pattern and quantized subsequence in the frequent pattern set, the similarity is calculated, the pattern with the highest similarity is selected, and the corresponding associated intent tag is obtained. Calculate the confidence score based on similarity.

4. The intelligent dialogue method for elderly companionship based on user profiles as described in claim 3, characterized in that: The calculation of joint utility and selection of the maximum value as the optimal intention includes: Extract frequency features from user profile vectors, concatenate them into interest feature sub-vectors, and calculate the interest utility of intent. The interaction data corresponding to the associated intent tags and interest tags are converted into TF-IDF vectors, and the cosine similarity is calculated and defined as the matching utility of the intent. Calculate the mean of interest utility and matching utility to obtain the joint utility, and select the maximum value as the optimal intention.

5. The intelligent dialogue method for elderly companionship based on user profiles as described in claim 4, characterized in that: The process of extracting the text corresponding to the optimal intent, converting it into a TF-IDF vector, calculating a semantic score, and generating the final comparison segment includes: Extract the text corresponding to the optimal intent, convert it into a TF-IDF vector, and calculate the semantic similarity with the text fragment; Segments with a similarity greater than a similarity threshold are selected to form a subset of candidate segments; Extract each candidate segment from the candidate segment subset and calculate its semantic score; The dialogue segment with the highest semantic score is selected as the final comparison segment.

6. The intelligent dialogue method for elderly companionship based on user profiles as described in claim 5, characterized in that: The process of loading a pre-trained language model to generate dialogue sentences for intelligent dialogue with the elderly includes: Extract the text corresponding to the final comparison segment, load the pre-trained language model, and generate dialogue sentences; Convert dialogue sentences into speech to enable intelligent conversations with the elderly.

7. The intelligent dialogue method for elderly companionship based on user profiles as described in claim 6, characterized in that: The collection and preprocessing of multimodal data includes: Multimodal data of the elderly is collected through smart bracelets and smartphones, and then denoised and normalized. The multimodal data includes heart rate interval, step count, sleep duration, walking time, awake time, device usage frequency, voice, and interaction data.

8. A user profile-based intelligent dialogue system for elderly companionship, based on the user profile-based intelligent dialogue method for elderly companionship as described in any one of claims 1 to 7, characterized in that: include, The data collection and processing module is used to collect multimodal data of the elderly and perform noise reduction and normalization processing. The user profile clustering module is used to calculate user features, concatenate them into user profile vectors, generate state triples through K-means clustering, construct a prefix tree, and generate a candidate pattern set. The intent matching module is used to extract and filter intent tags from candidate patterns, optimize intent inference by calculating interest utility and matching utility, and select the optimal intent for response. The semantic scoring module is used to filter candidate text fragments, select the most matching dialogue fragments based on semantic similarity and semantic score, and conduct intelligent dialogue.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the intelligent dialogue method for elderly companionship based on user profiles as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the intelligent dialogue method for elderly companionship based on user profiles as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Similar speaking recognition method and system using linear and nonlinear feature extraction

    CA2492204A1

  • Intentional scene recognition method and system based on user portrait

    CN106489148A

  • Universal event sequence frequent plot mining method

    CN108563757A

  • Interactive accompanying method and device, electronic equipment and storage medium

    CN120030118A

  • VSA-based approval method, system and device, and storage medium

    CN120634710A