Old people comfort accompanying semantic understanding method and system based on natural language processing

By supplementing the argument structure of the elderly's fragmented spoken language and constructing emotional dependency links, combined with a semantic database of high-frequency scenarios for the elderly and feedback optimization, a deep understanding of the elderly's emotions and personalized responses are achieved. This solves the shortcomings of existing technologies in understanding the elderly's spoken language and improves the pertinence and practicality of companionship services.

CN122113929APending Publication Date: 2026-05-29CORTELCO SHANGHAI INFORMATION TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CORTELCO SHANGHAI INFORMATION TECH CO LTD
Filing Date
2026-02-26
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately understand the core actions, related objects, and scene elements in the fragmented spoken language of the elderly, and they do not delve deeply enough into the elderly's emotional expressions, resulting in insufficient targeting and practicality of companionship services.

Method used

By using semantic role labeling enhancement algorithms to complete the argument structure of elderly people's spoken language, constructing emotional dependency links, and aligning them with a semantic database of high-frequency scenarios for the elderly, a scenario-based emotional semantic adaptation model is generated. Combined with feedback from the elderly to optimize the algorithm and semantic database, self-learning and personalized responses are achieved.

Benefits of technology

By accurately capturing the emotional inclinations in the elderly's speech and providing emotionally resonant responses, the pertinence and practicality of companionship services are improved. By adapting to the elderly's language habits and emotional changes, high-quality and personalized elderly care services are provided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122113929A_ABST
    Figure CN122113929A_ABST
Patent Text Reader

Abstract

The application provides an old people comforting accompanying semantic understanding method and system based on natural language processing, relates to the technical field of natural language processing, and first collects old people's oral expressions and generates a complete semantic role framework through semantic role labeling enhancement processing; then constructs a sentiment dependency link to generate a sentiment semantic dependency set; then aligns the sentiment semantic dependency set with an old people high-frequency scene semantic library to generate a scene sentiment semantic adaptation model; generates a semantic sentiment linkage response text according to the scene sentiment semantic adaptation model; collects old people's feedback expressions, converts the feedback expressions into supplementary semantic units, adds the supplementary semantic units to the semantic library, adjusts the correlation degree calculation parameters, and re-trains the semantic role labeling enhancement algorithm. The application improves the understanding accuracy of old people's oral expressions, enhances the emotional resonance, and can provide more high-quality personalized old people comforting accompanying services.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of natural language processing technology, and more specifically, to a semantic understanding method and system for elderly care and companionship based on natural language processing. Background Technology

[0002] With the increasing aging of society, the need for emotional support and companionship among the elderly is becoming more prominent. Natural language processing (NLP) technology is gaining attention in the field of elderly care and companionship, aiming to provide more attentive services by understanding the language of the elderly. However, existing technologies face many challenges in practical application.

[0003] When elderly people communicate, their spoken language is often fragmented, with incomplete argument structures and missing or vague key information such as core actions, related objects, and scene elements. This makes it difficult for traditional natural language processing methods to accurately understand the true meaning of their words. For example, an elderly person may simply mention a thing or action without fully expressing the underlying intention and emotion.

[0004] Meanwhile, existing technologies lack a deep understanding of emotional factors when processing the language of the elderly. The emotional expressions of the elderly are often subtle and complex, making it difficult for traditional methods to accurately capture the emotional inclinations in their speech, let alone effectively link emotions with semantics, thus failing to generate responses that meet the emotional needs of the elderly.

[0005] Furthermore, existing semantic databases lack precise organization and optimization for high-frequency scenarios for the elderly, making it difficult to quickly and accurately provide appropriate semantic information based on the elderly's actual communication scenarios, resulting in insufficient targeting and practicality of companionship services. Summary of the Invention

[0006] In view of the aforementioned problems, and in conjunction with the first aspect of the present invention, embodiments of the present invention provide a semantic understanding method for elderly companionship based on natural language processing, the method comprising: Collect elderly people's spoken language expressions, and perform semantic role annotation enhancement processing on the elderly people's spoken language expressions based on the semantic role annotation enhancement algorithm to complete the argument structure in their fragmented expressions and generate a complete semantic role framework. The argument structure covers the core actions, related objects and scene elements in the elderly people's spoken language expressions. Based on the complete semantic role framework, an emotional dependency link is constructed. In the same vector space, the vector representation of the core semantic component and the vector representation of the emotional feature vector are correlated. Based on the correlation calculation result, a data set containing semantic components, emotional features and their correlation weights is generated as the emotional semantic dependency set. Alignment processing is performed between the emotional semantic dependency set and the high-frequency scene semantic library for the elderly to generate a scene emotional semantic adaptation model. The high-frequency scene semantic library for the elderly is a data set constructed by clustering and labeling the collected oral expressions of the elderly through natural language processing algorithms. It is organized according to preset scene categories. The semantic units stored under each category have a correlation degree with the preset scene category that exceeds a first preset threshold, and the frequency of the stored sentence structure patterns in the historical expression data of the preset scene category exceeds a second preset threshold. Based on the scene-emotion semantic adaptation model, semantic and emotional linkage response text is generated, which integrates the complete semantic role framework and scene semantic units to echo the emotional tendency corresponding to the emotional feature vector. The feedback statements of the elderly in response to the semantic-emotional linkage response text are collected. The feedback statements are transformed into supplementary semantic units through the semantic role annotation enhancement process. The supplementary semantic units are added to the corresponding scene category of the elderly high-frequency scene semantic library. Based on the supplementary semantic units and their context in the feedback statements, the correlation calculation parameters used to construct the emotional dependency link are adjusted. The semantic role annotation enhancement algorithm is retrained using a new sample set containing the feedback statements.

[0007] Furthermore, embodiments of the present invention also provide a semantic understanding system for elderly care and companionship based on natural language processing, characterized in that it includes: A processor; a machine-readable storage medium for storing machine-executable instructions of the processor; wherein the processor is configured to perform the above-described semantic understanding method for elderly care and companionship based on natural language processing by executing the machine-executable instructions.

[0008] In another aspect, embodiments of the present invention also provide a computer program product, the computer program product including machine-executable instructions, the machine-executable instructions being stored in a computer-readable storage medium, the processor of the natural language processing-based elderly care and companionship semantic understanding system reading the machine-executable instructions from the computer-readable storage medium, the processor executing the machine-executable instructions, causing the natural language processing-based elderly care and companionship semantic understanding system to execute the aforementioned natural language processing-based elderly care and companionship semantic understanding method.

[0009] Based on the above, a semantic role annotation-based enhancement algorithm was used to process the elderly's spoken language, successfully completing the argument structure in fragmented expressions, generating a complete semantic role framework, constructing emotional dependency links, and calculating the correlation between core semantic component vectors and emotional feature vectors to generate an emotional semantic dependency set. This achieves a deep integration of semantics and emotion, accurately capturing the emotional tendencies in the elderly's speech, and generating more emotionally resonant responses based on emotional features, enhancing the interactivity and emotional connection in communication with the elderly. Aligning the emotional semantic dependency set with a high-frequency scenario semantic library for the elderly generates a scenario-based emotional semantic adaptation model, which can quickly provide adapted semantic and emotional responses according to the specific scenario of the elderly, improving the targeting and practicality of companionship services. Simultaneously, by collecting feedback from the elderly and continuously optimizing the semantic library and algorithm parameters, the system achieves self-learning and continuous improvement, better adapting to the elderly's language habits and emotional changes, and providing higher-quality, more personalized semantic understanding services for elderly companionship. Attached Figure Description

[0010] Figure 1 This is a schematic diagram of the execution flow of the semantic understanding method for elderly companionship based on natural language processing provided in an embodiment of the present invention.

[0011] Figure 2 This is a schematic diagram of exemplary hardware and software components of the semantic understanding system for elderly care and companionship based on natural language processing provided in an embodiment of the present invention. Detailed Implementation

[0012] The present invention will now be described in detail with reference to the accompanying drawings. Figure 1 This is a flowchart illustrating a semantic understanding method for elderly companionship based on natural language processing, provided in one embodiment of the present invention. The following is a detailed description of this semantic understanding method for elderly companionship based on natural language processing.

[0013] Step S110: Collect the elderly person's spoken language expression, and perform semantic role annotation enhancement processing on the elderly person's spoken language expression based on the semantic role annotation enhancement algorithm to complete the argument structure in its fragmented expression and generate a complete semantic role framework. The argument structure covers the core actions, related objects and scene elements in the elderly person's spoken language expression.

[0014] This example uses the complex spoken language of an elderly person communicating with a smart companion device in a home setting. The elderly person's speech includes multiple topic shifts, referential corrections, and semantic interruptions. For example: "Oh, that... it's about Old Li from Building 3, oh no, it's Old Wang from Unit 2. A while ago in the park, the one with the big locust tree, he said his back... he put on some kind of plaster... oh right, got it from the community service station. He used it for a while and said it was much better. My shoulder... these past few days it's been uncomfortable, even raising my arm is difficult..." The core actions involved in this speech are "said," "put on," "uncomfortable," and "difficult." The related objects include "Old Wang," "back," "plaster," and "shoulder." The scene elements include "last time," "a while ago," "park (the one with the big locust tree)," "community service station," "a while," and "the past few days." There are referential corrections from "Old Li from Building 3" to "Old Wang from Unit 2," missing arguments for "what kind of plaster," and interruptions in tone such as "oh" and "oh right."

[0015] Step S111: Receive the elderly person's spoken language expression, convert the elderly person's spoken language expression into text data through a speech-to-text algorithm, and simultaneously retain the natural pause intervals and specific pronouns in the elderly person's spoken language expression during the conversion process, and filter environmental interference noise through a text denoising algorithm.

[0016] The audio signal of an elderly person's spoken expression is received through a multi-channel microphone array. This audio signal contains mixed noises from indoor electrical appliances and outdoor ambient sounds. The speech-to-text algorithm employs an end-to-end model based on deep learning. First, the audio signal is preprocessed, including pre-emphasis, framing, and windowing operations, converting continuous audio into a short-time frame sequence. Each frame undergoes a Fourier transform to obtain spectral features, which are then converted to Mel spectral features using a Mel filter bank. The feature sequence is input to the encoder module, which consists of a multi-layer self-attention mechanism and a feedforward neural network. The self-attention mechanism captures long-distance dependencies by calculating the association weights of features at different locations, while the feedforward neural network performs a non-linear transformation on the features at each location. The encoder output is processed by the decoder module. The decoder combines the generated text sequence with the contextual information from the encoder output, focusing on relevant parts of the audio features through the attention mechanism to generate the text sequence.

[0017] During the conversion process, natural pause intervals are identified by detecting segments of the audio signal with energy lower than a set threshold and lasting for a certain duration, and are retained as punctuation marks or spaces in the text. For example, the pause after "Alas" in "Alas, that..." is converted to a comma, and the extended pause after "that" is converted to an ellipsis. Exclusive reference words such as "Lao Li in Building 3", "Lao Wang in Unit 2", "Community Service Station", "Big Sophora Tree", etc. are directly retained without replacement or standardization. The text noise reduction algorithm uses spectral subtraction. First, the power spectral density of the noise is estimated, the noise power spectrum is subtracted from the power spectrum of the speech signal, and then it is converted back to the time-domain signal through inverse Fourier transform to filter out environmental interference noise and generate noise-reduced text data.

[0018] Step S112: Perform word segmentation and词性标注 (this part seems to be a misspelling in Chinese, assuming it should be "词性 tagging") on the converted and noise-reduced text data. The word segmentation uses a dedicated segmentation dictionary trained and expanded on the daily spoken language corpus of the elderly. The词性标注 (tagging) is used to distinguish noun, verb, and adjective词性类别 (categories), generating a word segmentation and tagging result, where each segmentation unit corresponds to a basic semantic component in the text data.

[0019] The process of constructing the dedicated segmentation dictionary is as follows: Collect a large-scale daily spoken language corpus of the elderly, including transcriptions of recordings from scenarios such as family conversations and community exchanges. After cleaning and deduplicating the corpus, perform word frequency statistics and select high-frequency words as the basic vocabulary list. Use a word vector model to train the basic vocabulary list, and by calculating the semantic similarity between words, add words with similar semantics to the dictionary to form an extended vocabulary list. For the colloquial expressions, dialect words, and exclusive reference words commonly used by the elderly, manually organize and add them to the dictionary to form a dedicated segmentation dictionary.

[0020] When performing word segmentation on the converted and noise-reduced text data, match the text data with the words in the dedicated segmentation dictionary. Use the maximum matching method to sequentially find the longest matching word from left to right as a segmentation unit. For example, for the text "Lao Wang in Unit 2", the segmentation is "二单元 / 名词短语的 / 助词老王 / 名词". After word segmentation, perform词性标注 (tagging). Use a词性标注 (tagging) model based on statistical learning. This词性标注 (tagging) model takes the segmented word sequence as input and predicts the词性类别 (category) of the current word by extracting the context features of the word, such as the previous word, the next word, and their词性类别 (categories). The词性类别 (categories) include nouns, verbs, adjectives, pronouns, adverbs, etc. For example, "说" is tagged as a verb, "腰" is tagged as a noun, and "不得劲" is tagged as an adjective. Generate a word segmentation and tagging result, where each segmentation unit and its corresponding词性标签 (tag) form a basic semantic component.

[0021] It should be noted that the "词性标注" in the original Chinese text seems to be a misspelling. It should probably be "词性 tagging" or something more appropriate in English. The translation has been adjusted accordingly while trying to maintain the overall meaning as accurately as possible.Step S113: Call the optimized semantic role labeling model to perform semantic role labeling enhancement processing on the word segmentation labeling results. The optimized semantic role labeling model is equipped with an argument missing detection module for the fragmented characteristics of spoken language. The argument missing detection module locates the missing argument positions corresponding to the core actions in the text data and generates argument missing labeling results.

[0022] The optimized semantic role labeling model is an improvement upon the general semantic role labeling model, specifically tailored to the fragmented nature of spoken language among the elderly. The model's input is the word segmentation and labeling results. An embedding layer converts words and part-of-speech tags into vector representations. These vector representations are then processed by a bidirectional recurrent neural network layer for feature extraction, capturing contextual information. The feature vectors are input to an attention mechanism layer, which calculates the attention weight of each word for the core action. Words with higher weights are considered arguments for the core action. The model's output is a semantic role label for each word, such as core action, agent, patient, time, and location.

[0023] The argument missing detection module is a crucial component of this model, built upon a dependency parsing algorithm. This algorithm analyzes the word segmentation and tagging results to determine dependency relationships between words, such as subject-verb, verb-object, and modifier-head relationships. Based on the type of core action and common argument combination patterns, the argument missing detection module constructs an expected argument set, clarifying the types and quantities of arguments that the core action should contain. The arguments in the word segmentation and tagging results are compared with the expected argument set to detect missing argument types and their corresponding positions, and the semantic attributes of each missing position are labeled, generating the argument missing tagging results.

[0024] Step S1131: Load the optimized semantic role labeling model. The semantic role labeling model is trained and optimized through elderly people's spoken language expression samples. Its probability threshold for identifying semantic roles is set to be lower than that of the standard written language model. The adjustment of the probability threshold is based on the standard written language model achieving the highest F1 score on a validation set containing multiple fragmented spoken language sentences.

[0025] The optimized semantic role annotation model parameters and structural information are loaded from the model storage path. During training, this semantic role annotation model uses a large number of spoken language samples from elderly people as training data, and these samples are manually annotated with semantic roles. By adjusting the model's hyperparameters, such as the learning rate and batch size, the model achieves good performance on the training set. To adapt to the fragmented nature of spoken language, the probability threshold for identifying semantic roles is lowered. The adjustment process is as follows: a validation set containing multiple fragmented spoken language sentences is collected. Different probability thresholds are used to test the standard written language model on the validation set. The F1 score corresponding to each threshold is calculated, and the probability threshold that results in the highest F1 score is selected as the probability threshold for the optimized model.

[0026] Step S1132: Enable the argument missing detection module in the semantic role labeling model. The argument missing detection module is constructed based on the dependency parsing algorithm, and its detection rules integrate the dependency relationship probability model trained for spoken language corpora.

[0027] After enabling the argument missing detection module, the module first performs dependency parsing on the word segmentation and tagging results. The dependency parsing algorithm determines the relationship between the core action and other words by constructing a dependency relationship tree for the words. The dependency relationship probability model is trained on a large-scale spoken language corpus, which statistically calculates the occurrence probabilities of different dependency relationships in the spoken language corpus, such as the probability distributions of subject-predicate relationships, verb-object relationships, etc. The detection rules are based on the dependency relationship probability model. When the occurrence probability of the dependency relationship corresponding to a certain argument of the core action is lower than the set threshold, it is determined that the argument may be missing.

[0028] Step S1133: Screen out the verb components from the word segmentation and tagging results, and lock the core action words in the text data. Each core action word is the core object for argument missing detection.

[0029] Traverse each word segmentation unit in the word segmentation and tagging results, and screen out the verb components according to the词性标签 (word class tags). The verb components include verbs, verb phrases, etc., such as "say", "paste", "uncomfortable", etc. The screened verb components are used as the core action words, and each core action word is the core object for argument missing detection, and it is necessary to analyze whether the corresponding argument structure is complete.

[0030] Step S1134: Based on the typical argument combinations in the elderly's daily semantic argument library, construct an argument expectation set around the core action words. This argument expectation set clarifies the argument types and quantities that each core action should have.

[0031] The elderly's daily semantic argument library stores a large number of typical combinations of core actions and arguments, which are obtained by analyzing and summarizing the daily spoken language expressions of the elderly. For example, the typical argument combinations of the core action "say" include the agent (who is saying), the patient (what is being said), the object (to whom is being said), etc.; the typical argument combinations of the core action "paste" include the agent (who is pasting), the patient (what is being pasted), the part (where is being pasted), etc. For each core action word, retrieve its corresponding typical argument combination from the elderly's daily semantic argument library to construct an argument expectation set. The argument expectation set clarifies the argument types and quantities that each core action should have. For example, the argument expectation set of the core action "paste" includes the patient (1) and the part (1).

[0032] It should be noted that in the above translation, "词性标签" is directly retained in Chinese as there is no clear English equivalent provided in the context. You may need to replace it with the appropriate English term if there is a specific one in the actual situation.Step S1135: Compare the word segmentation and annotation results with the expected set of arguments, detect the missing argument types and corresponding positions, and mark the semantic attributes of each missing position to generate an initial argument missing annotation result.

[0033] Compare the arguments corresponding to the core action words in the word segmentation and annotation results with the expected set of arguments to check for missing arguments. For example, the expected set of arguments for the core action "paste" includes the patient and the part. In the word segmentation and annotation result "pasted some kind of plaster or something", "plaster" is the patient, but the part argument is missing. Determine the corresponding positions of the missing arguments by analyzing the position of the core action in the text and the context information. Mark the semantic attributes of each missing position. For example, the semantic attribute of the part argument is a body part noun. Generate an initial argument missing annotation result, which includes information such as the core action word, the missing argument type, the missing position, and the semantic attribute.

[0034] Step S1136: Perform redundancy detection on the initial argument missing annotation result, remove the false missing marks caused by spoken language repetition, and identify and remove the redundant missing marks generated by repeated annotation of the same or synonymous words based on the preset repetition pattern rules.

[0035] The preset repetition pattern rules are formulated according to the common repetition phenomena in the spoken language of the elderly, including lexical repetition, semantic repetition, etc. For example, an elderly person may say "my waist, my waist has been uncomfortable recently", where "waist" appears twice, which belongs to lexical repetition. Traverse the initial argument missing annotation result to check for repeated missing marks. Identify and remove the redundant missing marks generated by repeated annotation of the same or synonymous words according to the repetition pattern rules to ensure the accuracy of the argument missing annotation result.

[0036] Step S1137: Use a Transformer-based bidirectional encoder model to calculate the semantic coherence scores of the text fragments before and after each missing position in the initial argument missing annotation result, and remove the missing marks with semantic coherence scores higher than the preset semantic coherence threshold.

[0037] The Transformer-based bidirectional encoder model is pre-trained on a large-scale text corpus and can capture the semantic information of text fragments. Input each missing position in the initial argument missing annotation result and its surrounding text fragments into the model, and the model outputs the semantic vector representation of the text fragments. Calculate the similarity of the semantic vectors of the text fragments before and after the missing position as the semantic coherence score. If the semantic coherence score is higher than the preset semantic coherence threshold, it means that the text fragments before and after are semantically coherent, and the missing mark may be false and needs to be removed; otherwise, retain the missing mark.

[0038] Step S1138: Generate the argument missing annotation result, annotate the type, position and semantic attribute of each missing argument in the argument missing annotation result, establish a corresponding relationship between the argument missing annotation result and the word segmentation annotation result and store it to form structured annotation data.

[0039] Based on the processing results of the above steps, generate the final argument missing annotation result. In the annotation result, clearly annotate the type of each missing argument (such as patient, part, time, etc.), its position in the text data (such as the specific character index range), and its semantic attribute (such as body part noun, food noun, etc.). Establish a corresponding relationship between the argument missing annotation result and the word segmentation annotation result through position information, so that each missing argument can correspond to a specific position in the word segmentation annotation result. Store the annotation result after establishing the corresponding relationship as structured data, such as JSON format, for convenient subsequent argument completion processing.

[0040] Step S114: Retrieve the preset elderly daily semantic argument library, and according to the missing position indicated in the argument missing annotation result, match the corresponding argument content from the elderly daily semantic argument library for completion to improve the argument structure. The elderly daily semantic argument library is constructed by collecting the common expressions of the elderly through natural language processing algorithms, and it contains typical argument combinations corresponding to core actions.

[0041] The preset elderly daily semantic argument library is constructed by processing and analyzing a large number of common expressions of the elderly through natural language processing algorithms. The algorithm first performs word segmentation,词性标注 (should be "pos tagging" in English), and semantic role annotation on the common expressions of the elderly, and extracts the core actions and their corresponding arguments. Then, cluster and summarize the arguments, and summarize the typical argument combinations corresponding to each core action and store them in the argument library. The argument content in the argument library includes words, phrases, etc., and each argument content is associated with the core action and argument type.

[0042] According to the missing position and missing argument type indicated in the argument missing annotation result, retrieve the corresponding argument content from the elderly daily semantic argument library. During the retrieval process, consider the context information of the missing position and select the argument content that best matches the context semantics for completion. For example, for the missing part argument of "paste", combined with the context "his waist", match "waist" as the argument content from the argument library, and the completed text is "He said his waist, and pasted some kind of plaster on his waist...". Through argument completion, improve the argument structure of the core action.

[0043] Step S115: Perform dependency syntactic analysis on the completed argument structure, extract the syntactic association relationships among the core action, associated objects, and scene elements, and generate a dependency syntactic tree to represent the hierarchical association system of each semantic component.

[0044] Dependency parsing is a method for analyzing the syntactic structure of sentences. It determines the sentence's syntactic structure by analyzing the dependency relationships between words. To perform dependency parsing on the completed argument structure, a dependency graph is first constructed, with each word as a node and the dependencies between words as edges. Dependency relationships include subject-verb, verb-object, modifier-head, and adverbial-head relationships. For example, in "he said," "he" and "said" have a subject-verb relationship, and in "apply plaster," "apply" and "plaster" have a verb-object relationship.

[0045] A dependency syntax tree is generated based on the dependency graph. The root node of the tree is the core action word of the sentence, and the other nodes are arguments or modifiers of the core action. The dependency syntax tree clearly shows the hierarchical relationships between the core action, the related objects (such as agent and patient), and scene elements (such as time and place). For example, the core action "say" is the root node, and its child nodes include the agent "he," the patient "his old leg pain, he put on some kind of plaster...", etc. Through the dependency syntax tree, the structural relationships between the semantic components of the sentence can be intuitively understood.

[0046] Step S116: Based on the dependency syntax tree, construct an initial semantic role framework and label the semantic role type corresponding to each argument in the initial semantic role framework. The semantic role type covers agent, patient, and scene roles to cover all the core semantics of the text data.

[0047] An initial semantic role framework is constructed based on a dependency syntax tree. The core of this framework is the core action vocabulary. Around these core action vocabulary, their corresponding arguments are categorized and organized according to semantic role types. Semantic role types include agent, patient, time, place, instrument, and cause, with scene roles including time, place, and other scene-related roles. The semantic role type of each argument is determined based on the dependency relationships between words in the dependency syntax tree. For example, "he" is the agent of the core action "speak," "old cold legs" is the patient of the core action "apply," and "in the garden downstairs" is the scene role of the core action "speak." The semantic role type of each argument is labeled in the initial semantic role framework to ensure coverage of all core semantics in the text data.

[0048] Step S117: Compare the semantic role labeling results of the elderly person's past similar expressions, adjust the role labeling logic of the initial semantic role framework, and update the role type label of the arguments in the initial semantic role framework based on the semantic role labeling results of the elderly person's past similar expressions.

[0049] We collected semantic role annotation results from elderly individuals' past expressions of similar types. Similar expressions refer to those involving the same core action or similar scenarios. We compared the current initial semantic role framework with past annotation results to analyze the differences in role annotation logic. For example, in the past, when elderly individuals expressed "knee pain," "knee" was labeled as the recipient of the pain, while in the current initial framework, "knee" might be labeled as a body part. Therefore, we need to adjust the role annotation logic, updating the role type label for "knee" to recipient. Through comparison and adjustment, we made the role annotation of the initial semantic role framework more consistent with the elderly individuals' expression habits and semantic understanding.

[0050] Step S118: Extract the preceding text content related to the scene elements identified in the initial semantic role framework from the current interaction's session history, and use this preceding text content as a new text attribute to supplement the annotation on the corresponding scene element node.

[0051] The current interaction's conversation history refers to the elderly person's previous statements during this interaction. From the conversation history, we extract contextual content related to the identified scene elements in the initial semantic role framework. For example, the scene element "downstairs garden" might contain the statement "I saw many people while walking in the downstairs garden a few days ago." We extract this "I saw many people while walking in the downstairs garden a few days ago" as contextual content. This contextual content is then added as a new text attribute to the scene element node corresponding to "downstairs garden," enriching the scene element's information and making the semantic role framework more complete.

[0052] Step S119: Integrate the optimized argument structure, dependency syntax relations, and supplemented scene semantic components to form a complete semantic role framework, and output the complete semantic role framework.

[0053] The integrated and optimized argument structure (after argument completion), dependency syntactic relations (lexical dependencies obtained through dependency syntactic analysis), and supplemented scene semantic components (scene elements with added contextual attributes) form a complete semantic role framework. This framework presents the core semantic information of the text data in a structured form, including semantic roles such as core actions, agents, patients, time, and location, and the relationships between them. This complete semantic role framework serves as the foundation for subsequent construction of sentiment dependency links and generation of semantic-sentiment-linked response text.

[0054] Step S120: Based on the complete semantic role framework, construct the emotional dependency link. In the same vector space, calculate the correlation between the vector representation of the core semantic component and the vector representation of the emotional feature vector. Based on the correlation calculation result, generate a data set containing semantic components, emotional features and their correlation weights as the emotional semantic dependency set.

[0055] Step S121: Extract core semantic components from the complete semantic role framework. The core semantic components cover the word segmentation units corresponding to core actions, associated objects and scene elements, and generate a core semantic set, which serves as the core semantic carrier of emotional association.

[0056] The process iterates through each semantic role node in the complete semantic role framework, extracting word segments corresponding to core actions, associated objects, and scene elements. Core actions include "say," "stick," and "feel uncomfortable"; associated objects include "Old Wang," "waist," "ointment," and "shoulder"; scene elements include "a while ago," "park," "community service station," and "these days." These word segments are then combined to generate a core semantic set. Each word segment in the core semantic set is a core semantic carrier of the text data, and subsequent emotional dependency links will be built based on these carriers.

[0057] Step S122: Backtrack the original speech data of the elderly person's spoken expression, filter out environmental noise and equipment interference through speech preprocessing algorithm, and retain the pure speech signal.

[0058] The original speech data of the elderly person's spoken expression was retrieved and recorded synchronously during the recording process. A speech preprocessing algorithm was used to process the original speech data to remove environmental noise and equipment interference. The preprocessing process included: first, converting the sampling rate of the speech data to a preset rate; then, noise reduction using methods such as spectral subtraction or wavelet transform to remove environmental noise; next, speech enhancement to improve the clarity of the speech signal; and finally, endpoint detection to determine the start and end positions of the speech signal and remove silent segments. Through these processing steps, a clean speech signal was preserved for subsequent extraction of tone feature parameters.

[0059] Step S123: Extract tone feature parameters from the pure speech signal, and convert the tone feature parameters into an emotion feature vector through a feature quantization algorithm. The tone feature parameters include speech rate change curve, pause duration distribution, and tone word pronunciation intensity.

[0060] Step S1231: Perform speech rate analysis on the pure speech signal, split the speech segments according to the time axis, calculate the pronunciation duration and word count of each segment, and generate a speech rate change curve, which is used to reflect the dynamic changes in speech rate.

[0061] The pure voice signal is evenly split into multiple voice segments along the time axis, and the duration of each segment can be set according to the actual situation. Perform voice recognition on each voice segment, convert it into text, and count the number of words in the text. At the same time, calculate the pronunciation duration of each voice segment, that is, the end time of the voice segment minus the start time. Calculate the speaking speed of each segment according to the number of words and the pronunciation duration, and the speaking speed is equal to the number of words divided by the pronunciation duration. Arrange the speaking speeds of each segment in chronological order to generate a speaking speed change curve. The horizontal axis of the speaking speed change curve is time, and the vertical axis is the speaking speed. The ups and downs of the curve can intuitively reflect the dynamic change of the speaking speed. For example, the speaking speed of a certain segment is fast, and the speaking speed of a certain segment is slow.

[0062] Step S1232: Extract the pause features in the pure voice signal, mark the positions of natural pauses and semantic pauses, calculate the duration of each pause, count the distribution law of the pause durations, and generate pause duration distribution data.

[0063] In a pure voice signal, a pause refers to a temporary interruption of the voice signal. Pauses can be divided into natural pauses and semantic pauses. Natural pauses are caused by physiological reasons (such as breathing), and semantic pauses are generated to express semantics or emphasize a certain content. Mark the positions of pauses through an endpoint detection algorithm. The endpoint detection algorithm determines the start and end positions of the voice signal by detecting the energy and zero-crossing rate of the voice signal, thereby finding the positions of pauses. For each pause, calculate its duration, that is, the end time of the pause minus the start time. Count the frequencies of pauses in different duration ranges to generate pause duration distribution data. For example, count the number of pauses with durations in the ranges of 0 - 0.2 seconds, 0.2 - 0.5 seconds, and above 0.5 seconds to obtain the pause duration distribution data.

[0064] Step S1233: Locate the modal particle components in the pure voice signal, distinguish modal particles of the exclamatory, soothing, and interrogative types, and calculate the pronunciation intensity parameters of each modal particle through a pronunciation intensity detection algorithm.

[0065] Modal particles are words that express mood, such as "alas", "oh", "ah", "ne", etc. First, perform voice recognition on the pure voice signal, convert it into text, and locate the modal particle components from the text. According to the semantics and pronunciation characteristics of modal particles, they are divided into different types such as exclamatory, soothing, and interrogative. For example, "alas" belongs to the exclamatory modal particle, "oh" belongs to the soothing modal particle, and "ma" belongs to the interrogative modal particle. The pronunciation intensity detection algorithm determines the pronunciation intensity parameters by calculating the energy of the voice segment corresponding to the modal particle. The greater the energy, the higher the pronunciation intensity. Calculate the pronunciation intensity parameters for each modal particle as part of the mood feature parameters.

[0066] Step S1234: Integrate the speech rate change curve, the pause duration distribution data, and the tone word pronunciation intensity parameters to generate a tone feature set, which includes acoustic feature parameters extracted from the speech signal for subsequent emotion classification.

[0067] By integrating speech rate variation curves, pause duration distribution data, and particle pronunciation intensity parameters, a tone feature set is formed. The speech rate variation curve can be represented by multiple feature values, such as the mean, variance, maximum, and minimum values ​​of the curve; the pause duration distribution data can be represented by the pause frequency within different duration ranges; and the particle pronunciation intensity parameter can be represented by the pronunciation intensity value of each particle. These feature values ​​collectively constitute the tone feature set, which includes various acoustic feature parameters extracted from the speech signal. These acoustic feature parameters can reflect the tone characteristics of elderly people when speaking.

[0068] Step S1235: Call the feature quantization algorithm to convert the non-numerical features in the tone feature set into standardized numerical parameters to unify the feature dimensions and form an initial feature matrix.

[0069] The tone feature set may contain non-numerical features, such as the type of tone words (exclamatory, soothing, interrogative, etc.). The feature quantization algorithm transforms these non-numerical features into standardized numerical parameters for subsequent calculations and processing. For categorical non-numerical features, one-hot encoding is used, mapping each category to a binary vector with a length equal to the number of categories. Only the position corresponding to the category is 1, and the other positions are 0. For numerical features, standardization is performed, mapping feature values ​​to the range [0,1] or [-1,1] to eliminate dimensional differences between different features. Through these processes, the tone feature set is transformed into an initial feature matrix, where rows represent samples and columns represent feature parameters.

[0070] Step S1236: Perform dimensionality reduction processing on the initial feature matrix using principal component analysis algorithm, retain core feature components, and remove redundant features to simplify the feature matrix structure and improve the efficiency of subsequent correlation analysis.

[0071] Principal Component Analysis (PCA) is a commonly used dimensionality reduction method. Its purpose is to transform a high-dimensional feature matrix into a low-dimensional feature matrix while preserving the core information of the data. PCA is performed on the initial feature matrix by first calculating the covariance matrix, which reflects the correlation between different features. Then, eigenvalue decomposition is performed on the covariance matrix to obtain eigenvalues ​​and eigenvectors. Eigenvalues ​​represent the importance of the corresponding eigenvectors; the larger the eigenvalue, the more important the corresponding eigenvector. The top k eigenvectors with the largest eigenvalues ​​are selected as principal components, and the initial feature matrix is ​​projected onto these principal components to obtain the dimensionality-reduced feature matrix. The choice of k can be determined based on the cumulative contribution rate of the eigenvalues; when the cumulative contribution rate reaches a preset threshold, principal component selection stops. Dimensionality reduction preserves core features, eliminates redundant features, and simplifies the feature matrix structure, thereby improving the efficiency of subsequent correlation analysis.

[0072] Step S1237: Convert the dimensionality-reduced feature matrix into an emotional feature vector. Each dimension of the emotional feature vector corresponds to a tone feature parameter, and its vector value reflects the quantification level of the corresponding feature.

[0073] Each row of the dimensionality-reduced feature matrix corresponds to a sample's dimensionality-reduced feature, which is then transformed into an emotional feature vector. Each dimension of the emotional feature vector corresponds to a tone feature parameter, which are the core feature components retained after principal component analysis. The magnitude of the vector value reflects the quantification level of the corresponding feature. For example, if a dimension corresponds to speech rate, a larger vector value indicates a faster speech rate; if a dimension corresponds to pause duration, a larger vector value indicates a longer pause. The emotional feature vector can comprehensively reflect the emotional state of the elderly person during spoken expression.

[0074] Step S1238: Input the emotional feature vector into a regression model trained based on an elderly tone and emotion association sample library. The output of the regression model is a calibration value for each dimension of the input vector. Use the calibration value to update the original vector.

[0075] The elderly tone and emotion association sample database contains a large number of tone feature vectors and their corresponding emotion labels. The emotion labels are obtained through manual annotation or emotion analysis algorithms. The regression model takes the tone feature vectors as input and the emotion labels as output, learning the mapping relationship between tone features and emotions through training. The current emotion feature vector is input into the trained regression model, and the model outputs calibration values ​​for each dimension of the input vector. These calibration values ​​are statistically derived from the data in the sample database and are used to adjust the values ​​of each dimension of the emotion feature vector to more accurately reflect the emotional state. The original emotion feature vector is updated using the calibration values ​​to obtain the corrected emotion feature vector.

[0076] Step S1239: Output the corrected sentiment feature vector as input data for the subsequent construction of the semantic sentiment association matrix, and record the feature extraction and quantization parameters for optimization of the feature quantization algorithm.

[0077] The corrected sentiment feature vector is output, which will serve as input data for constructing the semantic sentiment association matrix. Simultaneously, various parameters used in feature extraction and quantization are recorded, such as the sampling rate, parameters of the denoising algorithm, the number of principal components in principal component analysis, and parameters of the regression model. These parameters can be used to optimize the feature quantization algorithm. By analyzing the impact of parameters on feature extraction and quantization results, parameter values ​​can be adjusted to improve the accuracy and reliability of feature quantization.

[0078] Step S12310: Establish and store the correspondence between the sentiment feature vector and the core semantic set. The stored correspondence data is used for parameter calibration when constructing the semantic sentiment association matrix.

[0079] A correspondence is established between the sentiment feature vector and each core semantic component in the core semantic set. For example, a certain dimension of the sentiment feature vector is associated with the core semantic component "feeling uncomfortable". This correspondence can be determined through semantic similarity calculation or manual annotation. The established correspondence data is stored and used to calibrate the parameters of the semantic-sentiment association matrix when constructing it, so that the matrix can more accurately reflect the relationship between semantic components and sentiment features.

[0080] Step S124: Map each word segmentation unit in the core semantic set to a semantic vector and the sentiment feature vector to a sentiment vector. The semantic vector and the sentiment vector are mapped to the same shared semantic-sentiment latent space, with the same vector dimension and each dimension being a dimensionless latent feature representation. In this latent space, the correlation between each semantic vector and the sentiment vector is calculated using a similarity calculation algorithm to generate a correlation matrix.

[0081] A pre-trained word vector model is used to map each word segmentation unit in the core semantic set to a semantic vector. The word vector model is trained on a large-scale text corpus and can convert words into low-dimensional vector representations; the similarity between vectors reflects the semantic similarity of the words. Sentiment feature vectors are then mapped to sentiment vectors through a linear transformation. The parameters of this linear transformation are determined through training, ensuring that the sentiment vectors and semantic vectors have the same dimension and are both mapped to the same shared semantic-sentiment latent space. In this latent space, each dimension of the semantic and sentiment vectors is a dimensionless latent feature representation used to measure some latent attribute of semantics and sentiment.

[0082] Similarity calculation algorithms are used to calculate the correlation between semantic vectors and sentiment vectors. Commonly used similarity calculation algorithms include cosine similarity and Euclidean distance. Cosine similarity measures the similarity between two vectors by calculating the cosine of the angle between them; a larger value indicates higher similarity. Euclidean distance measures the similarity by calculating the geometric distance between two vectors; a smaller distance indicates higher similarity. The correlation calculation results of all semantic vectors and sentiment vectors are arranged into a matrix to generate a correlation matrix. The rows of the matrix represent semantic vectors, the columns represent sentiment vectors, and the matrix elements represent the correlation between the corresponding semantic vector and sentiment vector.

[0083] Step S125: Based on the correlation matrix, filter the combinations where the correlation between semantic vectors and sentiment vectors meets the set criteria, filter the elements in the correlation matrix whose values ​​exceed the preset correlation threshold, and determine the core semantic component and sentiment feature vector corresponding to the element as the correlation node.

[0084] A preset correlation threshold is set, determined based on experience or experimental data, to determine whether the correlation between semantic vectors and sentiment vectors is significant. Each element in the correlation matrix is ​​traversed, and elements with values ​​exceeding the preset correlation threshold are selected. These elements, whose corresponding core semantic components have a strong correlation with the sentiment feature vector, are identified as correlation nodes. Correlation nodes are the core of constructing the sentiment dependency chain, and the subsequent chain structure will be built upon these nodes.

[0085] Step S126: Using the associated node as the core, construct an emotional dependency link structure that connects other semantic components and emotional features, forming a technical link where semantics and emotions are interconnected, so as to define the association logic system of each node.

[0086] Centered on interconnected nodes, this approach connects other semantic components and emotional features through relationships, constructing an emotional dependency link structure. Each node in the link structure represents a core semantic component or emotional feature, and the lines connecting nodes represent their relationships. The direction and strength of these relationships are determined by the correlation values ​​in the correlation matrix and the semantic relationships between semantic components. For example, nodes with higher correlation values ​​are represented by thicker lines, and nodes with semantically causal relationships are represented by directed lines. By constructing this emotional dependency link structure, a technically linked system of semantic and emotional connections is formed, defining the logical system of connections between each node, enabling people to clearly understand the intrinsic connections between semantic components and emotional features.

[0087] Step S127: For the semantic component or sentiment feature in the sentiment dependency link that is directly associated with one associated node and not directly associated with another associated node, calculate the average correlation degree between it and the two end nodes, use it as the transitional correlation weight of the semantic component or sentiment feature, and add it to the link.

[0088] In an affective dependency chain, there may be semantic components or sentiment features that are directly associated with only one related node, but not with another. For these nodes, the average degree of association between them and the two related nodes needs to be calculated as a transitional association weight. The calculation method is as follows: first, find the directly associated node and the other unrelated related node; then, calculate the degree of association between the node and these two related nodes; finally, average the two degree values ​​to obtain the transitional association weight. Adding the transitional association weight to the affective dependency chain makes the chain structure more complete and reflects the relationships between all nodes.

[0089] Step S128: Retrieve the elderly emotional semantic association database, calculate the difference between the association pattern of the current emotional dependency link and the typical association patterns stored in the database, and adjust the association weight of the current link by a weighted average based on the difference. The elderly emotional semantic association database is constructed through the emotional analysis results of historical expressions and contains typical association patterns of semantic components and emotional features.

[0090] The elderly emotional semantic association database is constructed through emotional analysis of the elderly's historical statements. The database stores a large number of typical association patterns between semantic components and emotional features. These patterns are obtained through association analysis and summarization of semantic components and emotional features in historical statements. The database is retrieved, and the association pattern of the current emotional dependency link is compared with the typical association patterns in the database, calculating the differences between them. Differences can be measured using pattern matching algorithms or similarity calculation algorithms. Based on the magnitude of the difference, a weighted average adjustment is applied to the association weights of the current link. For example, if the weight of a certain association in the current link differs significantly from the weight in the typical association pattern, the weight is adjusted according to the degree of difference to make it closer to the weight in the typical pattern.

[0091] Step S129: Integrate the sentiment dependency link, the core semantic set, and the sentiment feature vector to generate an initial sentiment semantic dependency set, so that each word segmentation unit in the core semantic set is associated with at least one sentiment feature vector.

[0092] The sentiment dependency chain, core semantic set, and sentiment feature vectors are integrated to generate an initial sentiment semantic dependency set. During integration, it is ensured that each word segmentation unit in the core semantic set is associated with at least one sentiment feature vector; these associations are represented by the association nodes and transitional association weights in the sentiment dependency chain. The initial sentiment semantic dependency set is a dataset containing semantic components, sentiment features, and their association weights, comprehensively reflecting the semantic and sentiment information of the text data.

[0093] Step S1210: According to a predefined weight distribution rule trained on a spoken sentiment corpus, the weights of various associated nodes in the sentiment dependency link are normalized and adjusted. Based on the co-occurrence frequency of sentiment words and semantic components in the historical expressions, the normalized weights are fine-tuned, and the adjusted link and weight data are output as the sentiment semantic dependency set.

[0094] The predefined weight distribution rules are obtained based on a spoken sentiment corpus, which contains a large number of spoken expressions with sentiment tags. By statistically analyzing the co-occurrence of sentiment words and semantic components in the corpus, the weight distribution patterns of different types of associated nodes are obtained. Based on these patterns, the weights of various associated nodes in the sentiment dependency chain are normalized and adjusted, mapping the weight values ​​to the range of [0,1], making the weights of different nodes comparable.

[0095] Then, the normalized weights are fine-tuned based on the co-occurrence frequency of sentiment words and semantic components in historical expressions. A higher co-occurrence frequency indicates a stronger association between sentiment words and semantic components, and the corresponding association weight should be increased appropriately; a lower co-occurrence frequency indicates a lower association weight. This fine-tuning allows the weights to more accurately reflect the actual degree of association between sentiment words and semantic components. Finally, the adjusted sentiment dependency chains and weight data are output as the sentiment semantic dependency set.

[0096] Step S130: Align the emotional semantic dependency set with the elderly high-frequency scene semantic library to generate a scene emotional semantic adaptation model; wherein, the elderly high-frequency scene semantic library is a data set constructed by clustering and labeling the collected elderly spoken expressions through natural language processing algorithms. It is organized according to preset scene categories. The semantic units stored under each category have a correlation degree with the preset scene category that exceeds a first preset threshold, and the frequency of the stored sentence structure patterns in the historical expression data of the preset scene category exceeds a second preset threshold.

[0097] Step S131: Retrieve the high-frequency scene semantic database for the elderly. The high-frequency scene semantic database for the elderly is constructed by batch analysis of the elderly’s daily scene expressions through natural language processing algorithms. It is stored in categories such as daily life, diet, health and memory. Each scene corresponds to a unique semantic unit and expression paradigm.

[0098] The high-frequency scene semantic database for the elderly is constructed through batch analysis of a large number of daily scene expressions of the elderly using natural language processing algorithms. The algorithm first collects and organizes these expressions, then performs word segmentation, part-of-speech tagging, and semantic role labeling. Next, a clustering algorithm is used to group the expressions, grouping semantically similar expressions together. Each category corresponds to a predefined scene category, such as daily living, eating, health, and reminiscing. The expressions within each scene category are analyzed to extract specific semantic units and expression paradigms. Specific semantic units refer to frequently occurring words and phrases within that scene category; expression paradigms refer to common sentence structure patterns within that scene category. These specific semantic units and expression paradigms are then stored in the high-frequency scene semantic database for the elderly, categorized by scene category.

[0099] Step S132: Extract core semantic units and sentiment feature vectors from the sentiment semantic dependency set to generate an aligned core set.

[0100] Each element in the emotional semantic dependency set is traversed to extract core semantic components and emotional feature vectors. Core semantic components refer to semantic components that serve as association nodes in the emotional dependency chain, such as core actions and associated objects; emotional feature vectors are vectors reflecting emotional states. The extracted core semantic components and emotional feature vectors are combined to generate an aligned core set. The aligned core set is the basis for aligning the emotional semantic dependency set with the semantic database of high-frequency scenarios for the elderly, and it contains the core content that needs to be matched.

[0101] Step S133: Establish semantic matching rules, which are based on semantic similarity analysis algorithms in natural language processing, and set similarity thresholds for semantic unit matching.

[0102] Semantic matching rules are based on semantic similarity analysis algorithms in natural language processing, used to determine whether two semantic units match. Semantic similarity analysis algorithms measure the degree of semantic association between semantic units by calculating the similarity between their semantic vectors. A similarity threshold is set for semantic unit matching; when the similarity between two semantic units exceeds this threshold, they are considered a match; otherwise, they are not. The similarity threshold needs to be adjusted according to the actual application scenario and data characteristics to ensure matching accuracy.

[0103] Step S134: Compare the semantic units in the alignment core set with the scene-specific semantic units in the elderly high-frequency scene semantic library one by one, calculate the semantic similarity and record the similarity value of each comparison result, match the corresponding scene category according to the similarity value, and generate scene matching results.

[0104] Step S1341: Extract semantic units from the alignment core set, classify and organize them according to nouns, verbs and adjectives, generate a semantic unit classification set, and label the part of speech and core meaning of each semantic unit.

[0105] Extract all independent semantic units from the alignment core set, such as "knee," "pain," "community hospital," and "applying ointment." Perform part-of-speech (POS) analysis on these semantic units, distinguishing them into different parts of speech categories such as nouns (e.g., "knee," "community hospital," "ointment"), verbs (e.g., "apply," "pain"), and adjectives (e.g., "effective," "serious"), forming a semantic unit classification set. For each semantic unit, clearly label its core meaning; for example, the core meaning of "knee" is "the joint of the lower limb," and the core meaning of "applying ointment" is "the action of applying a paste-like medication to a body part."

[0106] Step S1342: Retrieve scene-specific semantic units from the high-frequency scene semantic library for the elderly, split them according to scene category, so that each scene corresponds to a set of exclusive semantic units, forming a scene semantic classification index.

[0107] The high-frequency semantic database for the elderly pre-stores scene-specific semantic units according to scene categories (such as health care, daily diet, leisure activities, medical consultation, etc.). For example, the scene-specific semantic units for health care include "pain," "swelling," "ointment," "massage," and "physiotherapy"; the scene-specific semantic units for daily diet include "rice," "noodles," "cooking," "flavor," and "tableware." These scene-specific semantic units are then broken down by scene category to construct a scene semantic classification index, which records all scene-specific semantic units and their core meanings under each scene category.

[0108] Step S1343: Compare each semantic unit in the semantic unit classification set with the exclusive semantic unit in the scene semantic classification index one by one, calculate the semantic similarity and record the similarity value of each comparison result.

[0109] The process iterates through each semantic unit in the semantic unit classification set, such as "knee pain," and then compares it sequentially with the specific semantic units under each scene category in the scene semantic classification index. Taking the health care scene as an example, "knee pain" is compared with specific semantic units such as "pain," "joint discomfort," and "knee swelling" under this scene. When calculating semantic similarity, a cosine similarity calculation method based on word vectors is used. The two semantic units to be compared are converted into word vectors, and the similarity value is obtained by calculating the cosine of the angle between the two word vectors. The specific value of each comparison result is recorded.

[0110] Step S1344: Match the corresponding scene category based on the similarity value to generate a scene matching result.

[0111] Set a similarity threshold, for example, 0.6. For each semantic unit in the semantic unit classification set, compare its similarity value with the scene-specific semantic units with this threshold. If a semantic unit's similarity value with multiple scene-specific semantic units exceeds the threshold, or its similarity value with the core scene-specific semantic unit is the highest and exceeds the threshold, then the semantic unit is matched to that scene category. For example, "knee pain" has a similarity value of 0.85 with "pain" in the health care scene, and its similarity value with other scene-specific semantic units is below 0.5, so "knee pain" is matched to the health care scene. Summarize the scene matching results of all semantic units to generate the overall scene matching results and clarify the main scene categories corresponding to the current alignment core set.

[0112] Step S135: Based on the scene matching result, optimize the quantitative representation of the emotional feature vector using the emotional expression paradigm of the corresponding scene, and standardize the emotional feature vector using the emotional feature quantization template of the corresponding scene to generate a scene-based emotional feature vector.

[0113] Each scene category has its corresponding sentiment expression paradigm and sentiment feature quantification template. These paradigms and templates are derived from the analysis and summarization of a large number of sentiment expressions within that scene category. The sentiment expression paradigm specifies the manner and characteristics of sentiment expression in that scene; the sentiment feature quantification template specifies the quantitative representation form and range of the sentiment feature vector. Based on the scene matching results, the scene category to which the alignment core set belongs is determined, and then the sentiment expression paradigm of that scene category is used to optimize the quantitative representation form of the sentiment feature vector. Specifically, the sentiment feature quantification template for the corresponding scene is used to standardize and transform the sentiment feature vector, mapping the values ​​of each dimension of the vector to the range specified by the template, generating a scene-specific sentiment feature vector. Scene-specific sentiment feature vectors can better adapt to the sentiment expression needs of the corresponding scene.

[0114] Step S136: Construct a scene sentiment alignment matrix. Based on scene-specific semantic units, map the core semantic units in the sentiment semantic dependency set and the scene-specific sentiment feature vectors to the corresponding scene dimensions in the scene sentiment alignment matrix to generate a matrix alignment result.

[0115] A scene-specific semantic unit (MSU) alignment matrix is ​​constructed based on scene-specific semantic units. Rows in the matrix represent scene-specific MSUs, and columns represent scene dimensions, such as semantic and sentiment dimensions. Core MSUs in the sentiment semantic dependency set are mapped to corresponding positions in the matrix along with scene-specific sentiment feature vectors. Specifically, for each core MSU, its matching scene-specific MSU is found and mapped to the corresponding row in the matrix; for scene-specific sentiment feature vectors, they are mapped to the corresponding columns based on their sentiment dimensions. This mapping generates a matrix-based alignment result, which visually demonstrates the position and relationship between core MSUs and scene-specific sentiment feature vectors within the scene-specific sentiment alignment matrix.

[0116] Step S137: Using the scene semantic expansion algorithm, combined with the expression paradigm of the corresponding scene, supplement the matrix alignment result with the required contextual semantics. For each semantic unit in the matrix alignment result, retrieve and attach its common contextual collocation words from the expression paradigm of the corresponding scene to generate an extended semantic matrix.

[0117] The scene semantic expansion algorithm is an algorithm that supplements the semantic units in the matrix alignment result with contextual semantics based on the representation paradigm of the corresponding scene. The representation paradigm contains common sentence structures and contextual collocations in that scene. For each semantic unit in the matrix alignment result, common contextual words that accompany it are retrieved from the representation paradigm of the corresponding scene, and these words are appended to the semantic unit to supplement the contextual semantics. By supplementing the contextual semantics, the meaning of the semantic unit becomes richer and clearer. The supplemented matrix alignment result generates an expanded semantic matrix, which contains more complete scene semantic information.

[0118] Step S138: Using a semantic sentiment fusion algorithm, perform fusion processing on the scene alignment matrix supplemented with contextual semantics, integrate scene semantics, core semantics and sentiment features, and generate an initial scene sentiment semantic adaptation model.

[0119] Semantic sentiment fusion algorithms are used to fuse scene semantics, core semantics, and sentiment features. The algorithm first analyzes the scene alignment matrix supplemented with contextual semantics, extracting information from scene semantics, core semantics, and sentiment features. Then, a fusion strategy is employed to integrate this information; the fusion strategy can be weighted averaging, feature concatenation, etc. Through this fusion process, an initial scene sentiment semantic adaptation model is generated. This initial scene sentiment semantic adaptation model comprehensively considers scene semantics, core semantics, and sentiment features.

[0120] Step S139: Retrieve historical data on emotional adaptation for elderly scenarios, calculate the error between the output of the initial scene emotional semantic adaptation model and the optimization results of the corresponding scene in the historical data on emotional adaptation for elderly scenarios, and use gradient descent to update the internal parameters of the initial scene emotional semantic adaptation model to reduce the error.

[0121] The historical data for elderly scene emotion adaptation includes input data, output results, and corresponding optimization results from past use of the scene emotion semantic adaptation model. The optimization results represent the ideal output obtained through manual adjustment or other optimization methods. The historical data is retrieved, and the output results of the initial scene emotion semantic adaptation model on the historical input data are compared with the corresponding optimization results to calculate the error. This error can be measured using metrics such as mean squared error and cross-entropy. Gradient descent is used to update the internal parameters of the initial scene emotion semantic adaptation model, such as weights and biases, based on the direction of the error gradient. By iteratively updating the parameters, the error between the model output and the optimization results is reduced, improving the model's adaptation performance.

[0122] Step S1310: With the goal of maximizing the semantic alignment accuracy of the initial scene emotional semantic adaptation model on the elderly daily scene test set, the grid search method is used to optimize the fusion weight parameters of scene semantics and emotional features in the initial scene emotional semantic adaptation model, and the model with optimized parameters is output as the scene emotional semantic adaptation model.

[0123] The Elderly Daily Scene Test Set is a dataset containing numerous descriptions of daily scenes experienced by the elderly and their corresponding scene categories, used to evaluate the performance of scene-based emotional semantic adaptation models. To maximize the semantic alignment accuracy of the initial scene-based emotional semantic adaptation model on the test set, a grid search method is used to optimize the fusion weight parameters of scene semantics and emotional features in the model. The grid search method exhaustively explores all possible parameter combinations, training and testing the model under each combination, and calculating the semantic alignment accuracy. The parameter combination with the highest semantic alignment accuracy is selected as the optimal parameter combination and applied to the initial scene-based emotional semantic adaptation model, resulting in the optimized model. The output of the optimized model serves as the scene-based emotional semantic adaptation model, which exhibits high semantic alignment accuracy and emotional adaptation performance.

[0124] Step S140: Generate semantic and emotional linkage response text based on the scene emotional semantic adaptation model, integrate the complete semantic role framework and scene semantic unit, and echo the emotional tendency corresponding to the emotional feature vector.

[0125] Step S141: Analyze the scene category, core semantic unit, emotional feature vector and adaptation rules in the scene sentiment semantic adaptation model, and extract the core parameters required for response generation.

[0126] The scene-sentiment semantic adaptation model is analyzed to extract scene categories, core semantic units, sentiment feature vectors, and adaptation rules. Scene categories indicate the scene to which the current text data belongs, such as a health scene; core semantic units are key semantic components in the scene, such as "waist," "plaque," and "shoulder"; sentiment feature vectors reflect the emotional tendency of the text data, such as pain or discomfort; and adaptation rules specify how to generate response text based on scene categories, core semantic units, and sentiment feature vectors. This extracted information serves as the core parameters required for response generation.

[0127] Step S142: Retrieve the expression paradigm library corresponding to the scene category, and determine the sentence structure and vocabulary selection range of the response text based on the expression paradigm library. The expression paradigm library contains commonly used expression sentences and vocabulary combinations for the elderly in the same scene.

[0128] Each scenario category has its corresponding expression paradigm library, which contains commonly used sentence structures and vocabulary combinations for the elderly in that scenario. Sentence structure refers to the grammatical structure of the statement, such as subject-verb-object or subject-verb structure; vocabulary selection refers to the set of words commonly used in that scenario. The expression paradigm library corresponding to the current scenario category is retrieved, and the sentence structure and vocabulary selection range of the response text are determined based on the information in the library. For example, in a health scenario, common sentence structures might be "What's wrong with your [body part]?" or "We suggest you [take some measures]," and the vocabulary selection range might include "body part," "pain," "uncomfortable," and "see a doctor," etc.

[0129] Step S143: Extract core actions, associated objects, and scene elements from the complete semantic role framework, supplement the extracted content into the response text to construct core semantic content, and insert the extracted core actions, associated objects, and scene elements as key components into the preset syntactic slots of the response text to form the propositional content of the response text.

[0130] The complete semantic role framework is traversed to extract core actions, associated objects, and scene elements. Core actions include "apply" and "feel uncomfortable"; associated objects include "waist," "plaque," and "shoulder"; scene elements include "a while ago," "community service station," and "these past few days." The extracted content is then used as key components and inserted into pre-defined syntactic slots in the response text. These pre-defined syntactic slots are reserved positions within the sentence structure to fill key components. Through this filling, the core semantic content of the response text, i.e., the propositional content, is constructed. The propositional content is the main body of the response text, expressing the core meaning of the response.

[0131] Step S144: Based on the sentiment type library, the quantized value of the sentiment feature vector is converted into sentiment tendency, and the sentiment expression of the response text is adjusted by using the corresponding sentiment tendency's word combination and sentence tone.

[0132] The sentiment type library defines different sentiment types and their corresponding quantification ranges and sentiment tendency descriptions. The quantized values ​​of the sentiment feature vectors are compared with the quantification ranges in the sentiment type library to determine the corresponding sentiment type, which is then converted into a sentiment tendency, such as pain, discomfort, or anxiety. Based on the sentiment tendency, corresponding word combinations and sentence structures are selected to adjust the emotional expression of the response text. For example, for the sentiment tendency of pain, words such as "pain" and "uncomfortable" are selected, and a caring and sympathetic sentence structure is used; for the sentiment tendency of anxiety, words such as "worried" and "anxious" are selected, and a comforting and encouraging sentence structure is used.

[0133] Step S145: Retrieve the vocabulary combination library corresponding to the emotional tendency, and prioritize the selection of emotional words frequently used by the elderly from the vocabulary combination library. The vocabulary combination library is constructed by analyzing the elderly's emotional expressions through natural language processing algorithms, and it contains commonly used words and collocations for different emotional tendencies.

[0134] The lexical combination database corresponding to sentiment tendencies is constructed by analyzing the emotional expressions of the elderly using natural language processing algorithms. The algorithm performs word segmentation, part-of-speech tagging, and sentiment analysis on a large number of elderly emotional expressions, extracting commonly used words and collocations under different sentiment tendencies. It then retrieves the lexical combination database corresponding to the current sentiment tendency, prioritizing the selection of frequently used emotional words from the database. Frequently used emotional words are more in line with the language habits of the elderly, improving the friendliness and understandability of the response text. For example, under the sentiment tendency of pain, frequently used emotional words might include "painful" and "unbearable."

[0135] Step S146: Adjust the sentence structure according to the sentence pattern and tone norms corresponding to the said sentiment tendency, integrate the selected word combination with the core semantic content, and generate the initial response text. The sentence pattern and tone norms are predefined rule bases that map sentiment tendency classifications to specific sentence pattern templates and word selection probability distributions.

[0136] The sentence structure and tone specification is a predefined rule base that maps sentiment categories to specific sentence templates and vocabulary selection probability distributions. Sentence templates are fixed-structure statements, such as "You [body part][emotional vocabulary], we suggest you [measures]". The vocabulary selection probability distribution specifies the probability of selecting different words in different positions. Based on the current sentiment, the corresponding sentence template and vocabulary selection probability distribution are selected from the sentence structure and tone specification. The sentence structure is adjusted according to the sentence template, and words are selected from the chosen vocabulary combinations according to the vocabulary selection probability distribution. These words are then integrated with the core semantic content to generate the initial response text.

[0137] Step S147: Establish a mapping function from the scalar value of the sentiment feature vector to the sentiment intensity level; adjust the occurrence density of sentiment polarity words and the intensity level of modifiers in the initial response text according to predefined rules based on the sentiment intensity level calculated from the sentiment feature vector.

[0138] A mapping function is established that maps the scalar values ​​of the sentiment feature vector to sentiment intensity levels. Sentiment intensity levels can be divided into multiple levels, such as mild, moderate, and severe. The mapping function is established by dividing the scalar values ​​of the sentiment feature vector, with each division corresponding to a sentiment intensity level. The sentiment intensity level is calculated based on the sentiment feature vector, and then the density of sentiment polarity words and the intensity level of modifiers in the initial response text are adjusted according to predefined rules. Sentiment polarity words are words that express emotional tendencies, such as "pain" or "uncomfortable"; modifiers are words that modify sentiment polarity words, such as "very," "extremely," or "somewhat." The predefined rules specify the density of sentiment polarity words and the intensity level of modifiers for different sentiment intensity levels. For example, a severe sentiment intensity level corresponds to a higher density of sentiment polarity words and a stronger intensity level of modifiers.

[0139] Step S148: Retrieve the elderly person's past emotional expression feedback data. Based on the elderly person's preferred emotional vocabulary combinations and sentence types extracted from the past emotional expression feedback data, adjust the emotional expression in the initial response text so that the adjusted emotional expression matches the elderly person's emotional acceptance preferences. At the same time, check the adaptability of the emotional expression to the semantics of the scene.

[0140] For example, step S1481: retrieve the elderly person's past emotional expression feedback data, and extract the elderly person's preferred emotional word combinations and sentence types from the elderly person's past emotional expression feedback data. The elderly person's past emotional expression feedback data covers the elderly person's feedback records on different forms of emotional expression.

[0141] The database retrieved the elderly person's past emotional expression feedback data. This data included records of the elderly person's responses to various forms of emotional expression (such as comfort, encouragement, and sympathy) at different times and in different scenarios. For example, the elderly person showed a high level of acceptance for comforting phrases like "Don't worry too much," while their response to phrases like "It's nothing" was relatively lukewarm. Through natural language processing analysis of this feedback data, the elderly person's preferred emotional vocabulary combinations were extracted, such as "It will get better slowly," "Take more rest," and "Don't worry," as well as preferred sentence types, such as declarative sentences ("Your condition will gradually improve") and suggestive sentences ("I suggest you rest more").

[0142] Step S1482: Compare the emotional word combinations and sentence patterns in the initial response text, filter out content that is inconsistent with the elderly’s preferences, and replace it with emotional words and sentence patterns that the elderly frequently accept.

[0143] The emotional vocabulary and sentence structures in the initial response text are compared sentence by sentence with the elderly's preferences extracted in step S1481. For example, if the initial response text contains "Your pain is not serious," but past feedback data from the elderly shows that they have a low acceptance of expressions like "not serious" and prefer "it will gradually subside," then "not serious" is replaced with "it will gradually subside." If the initial response text uses a rhetorical question that the elderly do not prefer ("Don't you think you should rest?"), it is replaced with a suggested statement that the elderly prefer ("I suggest you take a rest").

[0144] Step S1483: Combine the expression paradigm corresponding to the scene category, check the adaptability of the adjusted emotional expression to the scene semantics, and correct the content that conflicts between the emotional expression and the scene expression habits.

[0145] Based on the scene matching results, determine the current scene category (e.g., health consultation scene). Retrieve the corresponding expression paradigm for that scene category. For example, health consultation scenes typically use professional, caring, and concise expressions, avoiding overly casual or vague vocabulary. Verify that the adjusted emotional expression conforms to the expression paradigm of that scene. If the adjusted text contains content that conflicts with the expression habits of health consultation scenes (such as "This illness is nothing serious"), correct it to "Your symptoms need attention; let's work together to find ways to improve them," ensuring that the emotional expression is both in line with the elderly's preferences and meets the semantic requirements of the scene.

[0146] Step S1484: For the adjusted emotional expression content, retrieve the emotional feedback records of the elderly in the same scenario, and further fine-tune the vocabulary selection and sentence structure based on the emotional feedback records.

[0147] Retrieve the elderly person's historical emotional feedback records in the same scenario category (such as a health consultation scenario). For example, in past health consultations, the elderly person readily accepted the suggestion to "seek medical attention promptly," but hesitated when told they might need to see a doctor. If the current adjusted emotional expression includes "may need to see a doctor," it is fine-tuned to "seek medical attention promptly" based on the historical feedback records. Simultaneously, consider the elderly person's preference for long and short sentences in the same scenario. If the elderly person prefers short sentences, the adjusted long sentence is broken down into multiple short sentences. For example, "Considering your knee pain, it is recommended that you apply a cold compress first and then rest" is broken down into "Your knee is painful; it is recommended that you apply a cold compress first, and then rest."

[0148] Step S1485: Call a pre-trained coherence evaluation model to calculate the connection score between sentiment words and their adjacent context segments in the adjusted text. For connection points where the connection score is lower than the set connection score threshold, use text completion technology based on a bidirectional language model to generate transition words or adjust the word order.

[0149] The model calls a pre-trained coherence evaluation model (such as a BERT-based coherence evaluation model) and inputs the adjusted sentiment expression text into the model. The model evaluates the coherence between each sentiment word in the text and its adjacent context segments, outputting a coherence score (range 0-1). A coherence score threshold of 0.7 is set. If the coherence score of a sentiment word and its context is lower than 0.7, for example, in the sentence "knee pain, will gradually subside, pay more attention to rest," the coherence scores of "will gradually subside" and "pay more attention to rest" are low, then a bidirectional language model (such as the GPT model) is used for text completion. This generates transition words (such as "therefore" or "therefore") between the two segments or adjusts the word order to "knee pain, pay more attention to rest, will gradually subside" to improve the coherence of the text.

[0150] Step S1486: Input the adjusted text into a readability analysis model. If the output reading difficulty level is higher than the preset level for the elderly group, the low-frequency complex words in the text will be automatically replaced with high-frequency basic words from the thesaurus, and the complex long sentences will be split into multiple simple short sentences using syntactic analysis tools.

[0151] The coherently adjusted text is input into a readability analysis model (such as a model based on the Flesch-Kincaid reading difficulty test), and the model outputs the reading difficulty level of the text. The preset reading difficulty level for the elderly is level 4 (corresponding to the fourth-grade reading level). If the model outputs a level higher than 4, such as level 5, the text is simplified. For low-frequency complex words (such as replacing "relieve" with "improve" and "suggest" with "advise"), high-frequency basic words are selected from the thesaurus for replacement. Simultaneously, syntactic analysis tools (such as Stanford Parser) are used to identify complex long sentences in the text and break them down into multiple simple short sentences. For example, "If your knee pain persists and is accompanied by swelling, then it is recommended that you go to the community hospital as soon as possible" is broken down into "If your knee continues to hurt and is swollen, then go to the community hospital as soon as possible."

[0152] Step S1487: Integrate the adjusted emotional expression content with the initial response text. While keeping the core actions, related objects, and scene elements extracted from the complete semantic role framework unchanged, merge the adjusted emotional expression content with the proposition content to generate the optimized response text.

[0153] During the integration process, ensure that the core actions (such as "pain" and "applying ointment"), related objects (such as "knee" and "ointment"), and scene elements (such as "community hospital" and "current time") in the complete semantic role framework are fully preserved and semantically unchanged in the optimized response text. The emotional expression content (such as preferred vocabulary, appropriate sentence structure, coherent sentences, and simplified expressions) adjusted through the above steps is organically merged with the propositional content (core semantic content). For example, "For knee pain, it is recommended to apply a cold compress first, and then rest" is merged with "The community hospital has ointment for knee pain" into "Your knee is painful; it is recommended to apply the ointment prescribed at the community hospital first, and then rest; it will gradually relieve the pain," generating the final optimized response text.

[0154] Step S149: Call a pre-trained language model to score the fluency of the initial response text, and for sentences with a fluency score lower than the fluency threshold, adjust the word order and replace words according to the suggestions provided by the language model; when replacing words, prioritize selecting words specific to the corresponding scenarios from the semantic library of high-frequency scenarios for the elderly.

[0155] The pre-trained language model, trained on a large-scale text corpus, is capable of evaluating text fluency. The initial response text is input into the pre-trained language model, which outputs a fluency score. The fluency score reflects the smoothness and naturalness of the text. A fluency threshold is set; when the fluency score is below this threshold, it indicates that the initial response text has grammatical inconsistencies. Based on suggestions provided by the language model, word order adjustments and word substitutions are performed on incoherent sentences. Word order adjustments change the order of words in a sentence to make it more idiomatic; word substitution replaces words in the original sentence with more appropriate words. When substituting words, priority is given to selecting scene-specific vocabulary from a high-frequency semantic database for elderly users to ensure that the replaced words meet the scene requirements and the language habits of the elderly.

[0156] Step S1410: Extract scene-specific semantic units corresponding to the scene category from the high-frequency scene semantic library for the elderly, and supplement them into the adjusted response text so that the supplemented response text is compatible with the high-frequency scene semantic library for the elderly.

[0157] Scene-specific semantic units corresponding to the current scene category are retrieved from a high-frequency scene semantic database for the elderly. These semantic units are commonly used and representative words and phrases in that scene. The extracted scene-specific semantic units are then added to the adjusted response text. The addition can be placed in the middle, at the beginning, or at the end of a sentence, depending on the content and structure of the response text. This addition makes the response text contain more scene-related information, creating a better fit with the high-frequency scene semantic database for the elderly, and improving the scene-specific relevance and accuracy of the response text.

[0158] Step S1411: Integrate all optimized content, generate semantic and emotional linkage response text, output it and send it to the elderly to obtain feedback, and simultaneously record the response generation parameters for subsequent optimization.

[0159] The optimized core semantic content, emotional expression, fluency adjustments, and scene-specific semantic unit additions from the above steps are integrated to generate the final semantic-emotional linked response text. This text integrates a complete semantic role framework with scene-specific semantic units, echoing the emotional tendencies corresponding to the emotional feature vectors, and accurately and naturally responds to the elderly person's spoken expression. The generated semantic-emotional linked response text is then output and sent to the elderly person to obtain their feedback. Simultaneously, various parameters during the response generation process, such as core parameters, sentence structure, and vocabulary selection, are recorded. These parameters will be used for subsequent optimization and improvement of the response generation model.

[0160] Step S150: Collect feedback statements from elderly people in response to the semantic-emotional linkage response text, transform the feedback statements into supplementary semantic units through the semantic role annotation enhancement process, add the supplementary semantic units to the corresponding scene category of the elderly high-frequency scene semantic library, adjust the correlation calculation parameters used to construct the emotional dependency link according to the supplementary semantic units and their context in the feedback statements, and retrain the semantic role annotation enhancement algorithm using a new sample set containing the feedback statements.

[0161] Step S151: Receive the elderly person's feedback on the semantic and emotional response text, convert the feedback into feedback text data using a speech-to-text algorithm, and perform text noise reduction processing on the feedback text data to retain the core semantic content.

[0162] After receiving the semantically and emotionally linked response text, the elderly person will provide feedback. The audio signal of this feedback is received via a microphone or other device and transmitted to a speech-to-text algorithm. The algorithm processes the audio signal, converting it into feedback text data. During this conversion, simultaneous text denoising is performed to remove environmental noise and interference from the feedback, preserving the core semantic content. The core semantic content refers to the key information in the feedback, such as acceptance, supplementation, or correction of the response text.

[0163] Step S152: Perform semantic role annotation enhancement processing on the noise-reduced feedback text data to generate a feedback semantic role framework, and extract core semantic units from the feedback semantic role framework.

[0164] The noise-reduced feedback text data undergoes semantic role annotation enhancement processing similar to step S113. First, the optimized semantic role annotation model is invoked to perform semantic role annotation enhancement processing on the word segmentation annotation results of the feedback text data. The missing argument detection module locates the positions of missing arguments and generates missing argument annotation results. Then, the elderly's daily semantic argument database is retrieved for argument completion to improve the argument structure. Next, dependency parsing is performed to extract syntactic relationships and generate a feedback semantic role framework. Core semantic units are extracted from the feedback semantic role framework. These core semantic units include core actions, associated objects, scene elements, etc., and these units are the key semantic components in the feedback expression.

[0165] Step S153: Transform the core semantic units in the feedback semantic role framework into supplementary semantic units, classify and organize them according to semantic type, and build a data content system that can be used for updates.

[0166] The extracted core semantic units are analyzed and processed, transforming them into supplementary semantic units. Supplementary semantic units refer to new semantic components that enrich the semantic database of high-frequency scenarios for the elderly, such as new body parts, symptoms, or measures mentioned by the elderly in their feedback. The supplementary semantic units are categorized and organized according to semantic types, including nouns, verbs, adjectives, and phrases. Through this categorization and organization, a data content system is built that can be used to update the semantic database of high-frequency scenarios for the elderly, making the storage and management of supplementary semantic units more organized.

[0167] Step S154: Retrieve the high-frequency scene semantic library for the elderly, write the supplementary semantic units into the storage path corresponding to the scene category, and add them to the corresponding scene category of the high-frequency scene semantic library for the elderly. At the same time, according to the preset timeliness rules, delete the semantic units and expression paradigms whose most recent call time is earlier than the set timestamp from the high-frequency scene semantic library for the elderly.

[0168] The system retrieves a high-frequency scenario semantic database for the elderly and determines the storage path for supplementary semantic units based on the scenario category of the feedback statements. These supplementary semantic units are then written to the corresponding scenario category in the high-frequency scenario semantic database according to their storage paths, enriching the semantic units for that scenario category. Simultaneously, the system cleans up the semantic units and expression paradigms in the high-frequency scenario semantic database according to preset timeliness rules. These timeliness rules specify the retention period for semantic units and expression paradigms; if the most recent call time of a semantic unit or expression paradigm is earlier than a set timestamp, it is deleted from the semantic database to ensure its timeliness and effectiveness.

[0169] Step S155: Analyze the emotional association logic in the feedback semantic role framework, extract a new semantic emotional association pattern, and replace the corresponding content in the original emotional dependency link construction rules with the new semantic emotional association pattern.

[0170] This paper analyzes the affective dependency links within the feedback semantic role framework, dissecting the affective association logic. The affective association logic refers to the ways and patterns of association between core semantic units and affective feature vectors. Based on the analysis results, new semantic-affective association patterns are extracted. These patterns represent new semantic and affective relationships reflected in feedback statements. The corresponding content in the original affective dependency link construction rules is replaced with these new semantic-affective association patterns, updating the construction rules to ensure that the construction of affective dependency links reflects these new relationships.

[0171] Step S156: Adjust the correlation calculation parameters used to construct the emotional dependency link based on the supplementary semantic unit and its context in the feedback statement.

[0172] The contextual information of supplementary semantic units in feedback statements reflects their semantic relationships with other semantic components. Based on this contextual information, the degree of association between supplementary semantic units and other core semantic components is analyzed. The parameters for calculating the degree of association are important factors affecting the results, such as weight parameters and thresholds in similarity calculation algorithms. The parameters for calculating the degree of association between supplementary semantic units and other semantic components are adjusted to make the results more accurately reflect the actual semantic relationships. For example, if a supplementary semantic unit has a close contextual association with a core semantic component, the weight parameter for calculating the degree of association between them is increased.

[0173] Step S157: Retrain the semantic role annotation enhancement algorithm using a new sample set containing the feedback representation.

[0174] The feedback statements and their corresponding semantic role annotations are added to the training sample set as new samples, forming a new sample set. The semantic role annotation enhancement algorithm is then retrained using this new sample set. During training, the model parameters are adjusted by learning the semantic role annotation patterns in the new sample set, improving the accuracy of the algorithm's semantic role annotation of the elderly person's new spoken expressions. Retraining can be done incrementally, updating the original model parameters to improve training efficiency and model performance. Through retraining, the semantic role annotation enhancement algorithm can continuously adapt to changes in the elderly person's spoken expressions, improving its semantic understanding ability.

[0175] In the construction and training of the artificial intelligence model involved in this embodiment, taking the optimized semantic role labeling model as an example, its necessary modules include an embedding layer, a bidirectional recurrent neural network layer, an attention mechanism layer, and an argument missing detection module. The embedding layer receives the vocabulary and part-of-speech tags from the word segmentation and labeling results, converting them into fixed-dimensional vector representations. The vocabulary vectors are initialized using a word vector model pre-trained on everyday spoken language corpora of the elderly, while the part-of-speech tag vectors are randomly initialized and updated during model training. The bidirectional recurrent neural network layer contains forward and backward recurrent units. The hidden state dimension of each recurrent unit is set according to the training data scale. Information flow is controlled through a gating mechanism to extract contextual features. The attention mechanism layer calculates the attention weight of each word for the core action. The weight calculation is based on the similarity between the word vector and the core action vector, and after normalization using a softmax function, an attention distribution is obtained, used to focus on core semantic components. The argument missing detection module includes a dependency analysis submodule and a threshold judgment submodule. The dependency analysis submodule uses a graph convolutional network to encode the sentence dependency syntax tree and outputs the dependency strength between words. The threshold judgment submodule compares the dependency strength with a preset threshold to locate the position of the missing argument.

[0176] The model training steps are as follows: First, a training dataset is constructed, collecting 5000 samples of elderly people's spoken language containing missing arguments. Core actions, argument types, and missing positions are manually labeled, and the dataset is divided into training and validation sets in an 8:2 ratio. Model parameters are initialized; word vectors in the embedding layer use pre-trained parameters, while other parameters are initialized using the Xavier initialization method. Training hyperparameters are set; the batch size is determined based on GPU memory capacity, the initial learning rate is set to 0.001, the Adam optimizer is used, the weight decay coefficient is 0.0001, and the training epochs are set to 50. After each epoch, the F1 score is calculated on the validation set. If the F1 score does not improve for 5 consecutive epochs, training stops. During training, the cross-entropy loss function is used, and the loss calculation considers argument recognition loss and role classification loss. Model parameters are updated through backpropagation. After each training epoch, the model performance is evaluated using the validation set, and the threshold parameters of the missing argument detection module are adjusted to achieve the optimal F1 score on the validation set.

[0177] In specific application scenarios, taking a health-related scenario as an example, the model input is the word segmentation and annotation results of the elderly's spoken language, including word sequences and part-of-speech tags. The output is a semantic role framework with argument missing markers. The input data settings must retain elderly-specific pronouns (such as "old cold legs" and "feeling weak") and natural pause intervals, and a dedicated word segmentation dictionary ensures segmentation accuracy. In the output data, argument missing positions are marked using square brackets followed by the argument type (such as "[missing patient (drug name)]"), facilitating subsequent argument completion processing. When the model is combined with a health-related scenario, by loading a health-specific semantic dictionary, the recognition ability for domain-specific words such as "knee," "waist," and "community hospital" is enhanced. Domain-specific word weights are added at the attention mechanism layer to improve the accuracy of calculating the correlation between core actions and health-related arguments.

[0178] The sentiment feature vector extraction model employs a three-layer fully connected neural network. Inputs include tone feature parameters such as speech rate, pauses, and intensity of interjections, which are standardized before being fed into the network. The first hidden layer has twice the number of neurons as the input feature dimension and uses the ReLU activation function; the second hidden layer has half the number of neurons as the first layer and uses the tanh activation function; the output layer has the number of neurons representing the sentiment dimension (e.g., joy, sadness, pain), and uses the softmax activation function to output the sentiment probability distribution. During training, a sentiment-related sample library of elderly speech, containing 10,000 speech samples and their corresponding sentiment labels, is used. The network parameters are optimized using the mean squared error loss function, enabling the model to accurately predict sentiment based on speech features.

[0179] In the semantic-emotional linked response text generation, the pre-trained language model adopts the Transformer architecture, with model parameters finely tuned based on the daily conversational corpus of the elderly. The input consists of the scene category, core semantic units, and emotional feature vectors output by the scene-emotional semantic adaptation model. Through prompting engineering, the input is converted into natural language prompts (e.g., "In a health scenario, the user mentions knee pain, with an emotional tendency of pain; please generate a comforting response"). The model outputs candidate response texts, which are then filtered for fluency and emotional fit to determine the final response. The input and output data are correlated by constructing a semantic-emotional mapping table, mapping the dimension values ​​of the emotional feature vectors to the intensity of emotional words in the response text (e.g., an emotional intensity of 0.8 corresponds to "very painful," and 0.5 corresponds to "somewhat uncomfortable"), ensuring that the emotional expression of the response text is consistent with the elderly person's emotional tendency.

[0180] Regarding data privacy protection, differential privacy technology is used to add Gaussian noise during the speech-to-text conversion process for the collected audio and text data of elderly people. The noise intensity is adjusted according to the data sensitivity level. Personal information (such as name, address, and specific medical condition) in the text data is anonymized and replaced with placeholders such as "[name]" and "[address]". The data is encrypted using the AES-256 encryption algorithm during storage, and access permissions are set to hierarchical authorization, allowing only the model training and inference modules to access the anonymized data to ensure that privacy-sensitive data is not leaked.

[0181] In one exemplary embodiment, a semantic understanding system for elderly care and companionship based on natural language processing is provided. This system can be a terminal, server, etc., and its internal structure diagram can be as follows: Figure 2 As shown, the semantic understanding system for elderly care and companionship based on natural language processing includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computational and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, near-field communication, or other technologies. When the computer program is executed by the processor, it implements a semantic understanding method for elderly care and companionship based on natural language processing. The display unit is used to form a visually visible image and can be a display screen, projection device, or virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device can be a touch layer covering the display screen, or a button, trackball, or touchpad set on the shell of the natural language processing-based elderly care semantic understanding system, or an external keyboard, touchpad, or mouse, etc.

[0182] It should be noted that, in order to simplify the description of the present invention and thus help to understand one or more embodiments of the invention, multiple features may sometimes be grouped into one embodiment, drawing or description thereof in the foregoing description of the embodiments of the present invention.

Claims

1. A semantic understanding method for elderly companionship based on natural language processing, characterized in that, The method includes: Collect elderly people's spoken language expressions, and perform semantic role annotation enhancement processing on the elderly people's spoken language expressions based on the semantic role annotation enhancement algorithm to complete the argument structure in their fragmented expressions and generate a complete semantic role framework. The argument structure covers the core actions, related objects and scene elements in the elderly people's spoken language expressions. Based on the complete semantic role framework, an emotional dependency link is constructed. In the same vector space, the vector representation of the core semantic component and the vector representation of the emotional feature vector are correlated. Based on the correlation calculation result, a data set containing semantic components, emotional features and their correlation weights is generated as the emotional semantic dependency set. Alignment processing is performed between the emotional semantic dependency set and the high-frequency scene semantic library for the elderly to generate a scene emotional semantic adaptation model. The high-frequency scene semantic library for the elderly is a data set constructed by clustering and labeling the collected oral expressions of the elderly through natural language processing algorithms. It is organized according to preset scene categories. The semantic units stored under each category have a correlation degree with the preset scene category that exceeds a first preset threshold, and the frequency of the stored sentence structure patterns in the historical expression data of the preset scene category exceeds a second preset threshold. Based on the scene-emotion semantic adaptation model, semantic and emotional linkage response text is generated, which integrates the complete semantic role framework and scene semantic units to echo the emotional tendency corresponding to the emotional feature vector. The feedback statements of the elderly in response to the semantic-emotional linkage response text are collected. The feedback statements are transformed into supplementary semantic units through the semantic role annotation enhancement process. The supplementary semantic units are added to the corresponding scene category of the elderly high-frequency scene semantic library. Based on the supplementary semantic units and their context in the feedback statements, the correlation calculation parameters used to construct the emotional dependency link are adjusted. The semantic role annotation enhancement algorithm is retrained using a new sample set containing the feedback statements.

2. The semantic understanding method for elderly companionship based on natural language processing according to claim 1, characterized in that, The process of collecting elderly people's spoken language is followed by semantic role annotation enhancement processing based on a semantic role annotation enhancement algorithm. This process aims to complete the argument structure in their fragmented expressions and generate a complete semantic role framework, including: The system receives spoken language from the elderly and converts it into text data using a speech-to-text algorithm. During the conversion process, the system retains the natural pauses and specific pronouns in the elderly’s spoken language and filters out environmental noise using a text denoising algorithm. The transformed and denoised text data is subjected to word segmentation and part-of-speech tagging. The word segmentation uses a dedicated word segmentation dictionary trained and expanded on the daily spoken language of the elderly. The part-of-speech tagging is used to distinguish the part-of-speech categories of nouns, verbs and adjectives and generate word segmentation tagging results. Each word segmentation unit corresponds to a basic semantic component in the text data. The optimized semantic role labeling model is invoked to perform semantic role labeling enhancement processing on the word segmentation labeling results. The optimized semantic role labeling model is equipped with an argument missing detection module for the fragmented characteristics of spoken language. The argument missing detection module locates the missing argument positions corresponding to the core actions in the text data and generates argument missing labeling results. The system retrieves a pre-defined semantic argument database for elderly people's daily life. Based on the missing positions indicated in the missing argument annotation results, it matches the corresponding argument content from the database to complete the argument structure. The database for elderly people's daily life is constructed by collecting commonly used expressions of the elderly through natural language processing algorithms, and it contains typical argument combinations corresponding to core actions. Dependency parsing is performed on the completed argument structure to extract the syntactic relationships between core actions, related objects, and scene elements, and a dependency parsing tree is generated to represent the hierarchical relationship system of each semantic component. Based on the dependency syntax tree, an initial semantic role framework is constructed, and the semantic role type corresponding to each argument is marked in the initial semantic role framework. The semantic role type covers agent, patient, and scene roles to cover all the core semantics of the text data. By comparing the semantic role labeling results of similar expressions made by the elderly in the past, the role labeling logic of the initial semantic role framework is adjusted, and the role type labels of arguments in the initial semantic role framework are updated based on the semantic role labeling results of similar expressions made by the elderly in the past. Extract the preceding text content related to the scene elements identified in the initial semantic role framework from the current interaction's session history, and use this preceding text content as a new text attribute to supplement the annotation on the corresponding scene element node; The optimized argument structure, dependency syntax relations, and supplemented scene semantic components are integrated to form a complete semantic role framework, which is then output.

3. The semantic understanding method for elderly companionship based on natural language processing according to claim 1, characterized in that, Based on the complete semantic role framework, an emotional dependency chain is constructed. Within the same vector space, the vector representations of core semantic components and the vector representations of emotional feature vectors are correlated. Based on the correlation calculation results, a data set containing semantic components, emotional features, and their correlation weights is generated as the emotional semantic dependency set, including: Core semantic components are extracted from the complete semantic role framework. The core semantic components cover the word segmentation units corresponding to core actions, associated objects and scene elements, and generate a core semantic set, which serves as the core semantic carrier of emotional association. By retrieving the original speech data of the elderly person's spoken expression, environmental noise and equipment interference are filtered out through speech preprocessing algorithms to retain the pure speech signal; The tone feature parameters are extracted from the pure speech signal and transformed into an emotion feature vector through a feature quantization algorithm. The tone feature parameters include speech rate change curve, pause duration distribution, and tone word pronunciation intensity. Each word segmentation unit in the core semantic set is mapped to a semantic vector, and the sentiment feature vector is mapped to a sentiment vector. The semantic vector and the sentiment vector are mapped to the same shared semantic-sentiment latent space, with the same vector dimension and each dimension being a dimensionless latent feature representation. In this latent space, the correlation between each semantic vector and the sentiment vector is calculated using a similarity calculation algorithm to generate a correlation matrix. Based on the correlation matrix, combinations in which the correlation between semantic vectors and sentiment vectors meets the set criteria are selected. Elements in the correlation matrix whose values ​​exceed the preset correlation threshold are selected, and the core semantic components and sentiment feature vectors corresponding to the elements are determined as correlation nodes. With the aforementioned associated nodes as the core, an emotional dependency link structure is constructed that connects other semantic components and emotional features, forming a technical link where semantics and emotions are interconnected, in order to define the association logic system of each node. For the semantic component or sentiment feature in the aforementioned sentiment dependency link that is directly associated with one associated node but not directly associated with another associated node, calculate the average correlation degree between it and the nodes at both ends, use it as the transitional correlation weight of the semantic component or sentiment feature, and add it to the link. The system retrieves the elderly emotional semantic association database, calculates the difference between the association pattern of the current emotional dependency link and the typical association pattern stored in the elderly emotional semantic association database, and adjusts the association weight of the current link by weighted average based on the difference; the elderly emotional semantic association database is constructed through the emotional analysis results of historical expressions, and contains typical association patterns of semantic components and emotional features. Integrate the sentiment dependency chain, the core semantic set, and the sentiment feature vector to generate an initial sentiment semantic dependency set, such that each word segmentation unit in the core semantic set is associated with at least one sentiment feature vector; Based on a predefined weight distribution rule trained on a spoken sentiment corpus, the weights of various associated nodes in the sentiment dependency chain are normalized and adjusted. Based on the co-occurrence frequency of sentiment words and semantic components in the historical expressions, the normalized weights are fine-tuned, and the adjusted chain and weight data are output as the sentiment semantic dependency set.

4. The semantic understanding method for elderly companionship based on natural language processing according to claim 1, characterized in that, The step of aligning the emotional semantic dependency set with the semantic library of high-frequency scenarios for the elderly to generate a scenario-based emotional semantic adaptation model includes: The high-frequency scene semantic database for the elderly is retrieved. This database is constructed by batch analysis of the elderly’s daily scene expressions through natural language processing algorithms. It is stored in categories such as daily life, diet, health and memory. Each scene corresponds to a unique semantic unit and expression paradigm. Extract core semantic units and sentiment feature vectors from the sentiment semantic dependency set to generate an aligned core set; Establish semantic matching rules, which are based on semantic similarity analysis algorithms in natural language processing, and set similarity thresholds for semantic unit matching; The semantic units in the alignment core set are compared one by one with the scene-specific semantic units in the elderly high-frequency scene semantic library. The semantic similarity is calculated and the similarity value of each comparison result is recorded. The corresponding scene category is matched according to the similarity value to generate the scene matching result. Based on the scene matching results, the quantitative representation of the emotional feature vector is optimized using the emotional expression paradigm of the corresponding scene, and the emotional feature vector is standardized and transformed using the emotional feature quantification template of the corresponding scene to generate a scene-based emotional feature vector. Construct a scene sentiment alignment matrix. Based on scene-specific semantic units, map the core semantic units in the sentiment semantic dependency set and the scene-specific sentiment feature vectors to the corresponding scene dimensions in the scene sentiment alignment matrix to generate a matrix alignment result. By using a scene semantic extension algorithm, combined with the expression paradigm of the corresponding scene, the required contextual semantics are supplemented to the matrix alignment result. For each semantic unit in the matrix alignment result, common contextual collocation words are retrieved from the expression paradigm of the corresponding scene and attached to it, generating an extended semantic matrix. The semantic sentiment fusion algorithm is used to perform fusion processing on the scene alignment matrix supplemented with contextual semantics, integrate scene semantics, core semantics and sentiment features, and generate an initial scene sentiment semantic adaptation model. Retrieve historical data on emotional adaptation in elderly scenarios, calculate the error between the output of the initial scenario emotional semantic adaptation model and the optimization results of the corresponding scenario in the historical data on emotional adaptation in elderly scenarios, and use gradient descent to update the internal parameters of the initial scenario emotional semantic adaptation model to reduce the error. With the goal of maximizing the semantic alignment accuracy of the initial scene emotion semantic adaptation model on the elderly daily scene test set, a grid search method is used to optimize the fusion weight parameters of scene semantics and emotion features in the initial scene emotion semantic adaptation model, and the model with optimized parameters is output as the scene emotion semantic adaptation model.

5. The semantic understanding method for elderly companionship based on natural language processing according to claim 1, characterized in that, The process involves generating semantic and emotionally linked response text based on a scene-based emotional semantic adaptation model, integrating a complete semantic role framework with scene semantic units, and echoing the emotional tendencies corresponding to the emotional feature vectors, including: The scene category, core semantic unit, emotional feature vector and adaptation rule in the scene sentiment semantic adaptation model are analyzed, and the core parameters required for response generation are extracted. The expression paradigm library corresponding to the scene category is retrieved, and the sentence structure and vocabulary selection range of the response text are determined based on the expression paradigm library. The expression paradigm library contains commonly used expression sentences and vocabulary combinations for the elderly in the same scene. The core actions, associated objects, and scene elements are extracted from the complete semantic role framework. The extracted content is added to the response text to construct the core semantic content. The extracted core actions, associated objects, and scene elements are used as key components and inserted into the preset syntactic slots of the response text to form the propositional content of the response text. Based on the sentiment type library, the quantized values ​​of the sentiment feature vectors are converted into sentiment tendencies, and the sentiment expression of the response text is adjusted by using word combinations and sentence structures with corresponding sentiment tendencies. The vocabulary combination library corresponding to the emotional tendency is retrieved, and emotional words frequently used by the elderly are selected from the vocabulary combination library. The vocabulary combination library is constructed by analyzing the emotional expressions of the elderly through natural language processing algorithms, and it contains commonly used words and collocations of different emotional tendencies. The sentence structure is adjusted according to the sentence pattern and tone norms corresponding to the emotional tendency, and the selected word combination is integrated with the core semantic content to generate the initial response text. The sentence pattern and tone norms are predefined rule bases that map emotional tendency classifications to specific sentence pattern templates and word selection probability distributions. Establish a mapping function from the scalar values ​​of the sentiment feature vector to the sentiment intensity level; based on the sentiment intensity level calculated from the sentiment feature vector, adjust the occurrence density of sentiment polarity words and the intensity level of modifiers in the initial response text according to predefined rules; Retrieve past emotional expression feedback data of the elderly, and based on the emotional word combinations and sentence types preferred by the elderly extracted from the past emotional expression feedback data, adjust the emotional expression in the initial response text so that the adjusted emotional expression fits the elderly's emotional acceptance preferences, and at the same time check the adaptability of the emotional expression to the semantics of the scene; A pre-trained language model is invoked to score the fluency of the initial response text. For sentences with a fluency score lower than the fluency threshold, word order adjustment and word replacement are performed based on the suggestions provided by the language model. When replacing words, priority is given to selecting words specific to the corresponding scenarios from the semantic library of high-frequency scenarios for the elderly. Extract scene-specific semantic units corresponding to the scene category from the high-frequency scene semantic library for the elderly, and supplement them into the adjusted response text so that the supplemented response text is compatible with the high-frequency scene semantic library for the elderly. Integrate all optimized content to generate semantic and emotionally linked response text, output it and send it to the elderly to obtain feedback, and simultaneously record the response generation parameters for subsequent optimization.

6. The semantic understanding method for elderly companionship based on natural language processing according to claim 2, characterized in that, The optimized semantic role labeling model is invoked to perform semantic role labeling enhancement processing on the word segmentation labeling results. This optimized semantic role labeling model is equipped with an argument missing detection module to address the fragmented nature of spoken language. The argument missing detection module locates the missing arguments corresponding to core actions in the text data, generating argument missing labeling results, including: The optimized semantic role labeling model is loaded. The semantic role labeling model is trained and optimized through elderly people's spoken language expression samples. Its probability threshold for identifying semantic roles is set to be lower than that of the standard written language model. The adjustment of the probability threshold is based on the standard written language model achieving the highest F1 score on a validation set containing multiple fragmented spoken language sentences. The argument missing detection module in the semantic role labeling model is enabled. The argument missing detection module is built based on the dependency parsing algorithm, and its detection rules integrate the dependency relation probability model trained on spoken language corpus. Verb components are filtered from the word segmentation and annotation results to identify core action words in the text data. Each core action word is used as the core object for argument missing detection. Based on the typical argument combinations in the elderly's daily semantic argument library, an argument expectation set is constructed around the core action vocabulary. This argument expectation set clarifies the argument type and quantity that each core action should have. By comparing the word segmentation and annotation results with the expected set of arguments, the missing argument types and corresponding positions are detected, and the semantic attributes of each missing position are marked to generate the initial missing argument annotation results; Redundancy detection is performed on the initial argument missing annotation results to remove false missing tags caused by repeated spoken expressions. Based on the preset repetition pattern rules, redundant missing tags generated by repeated annotation of the same or synonyms in the initial argument missing annotation results are identified and removed. Using a Transformer-based bidirectional encoder model, the semantic coherence score of the text segments before and after each missing position in the initial argument missing labeling results is calculated, and missing labels with semantic coherence scores higher than a preset semantic coherence threshold are removed. Generate argument missing annotation results, and annotate the type, position and semantic attributes of each missing argument in the argument missing annotation results. Establish a correspondence between the argument missing annotation results and the word segmentation annotation results and store them to form structured annotation data.

7. The semantic understanding method for elderly companionship based on natural language processing according to claim 3, characterized in that, The step of extracting tone feature parameters from the clean speech signal and converting the tone feature parameters into an emotion feature vector using a feature quantization algorithm includes: Speech rate analysis is performed on the pure speech signal, the speech segments are split according to the time axis, the pronunciation duration and word count of each segment are calculated, and a speech rate change curve is generated. This speech rate change curve is used to reflect the dynamic changes in speech rate. Extract pause features from the clean speech signal, mark the positions of natural pauses and semantic pauses, calculate the duration of each pause, statistically analyze the distribution pattern of pause duration, and generate pause duration distribution data; The modal particles in the pure speech signal are located, and exclamatory, soothing, and interrogative modal particles are distinguished. The pronunciation intensity parameter of each modal particle is calculated by a pronunciation intensity detection algorithm. By integrating the speech rate variation curve, the pause duration distribution data, and the tone word pronunciation intensity parameters, a tone feature set is generated. This tone feature set includes acoustic feature parameters extracted from the speech signal for subsequent emotion classification. The feature quantization algorithm is invoked to convert the non-numerical features in the tone feature set into standardized numerical parameters to unify the feature dimensions and form an initial feature matrix. The initial feature matrix is ​​subjected to dimensionality reduction processing by principal component analysis algorithm, retaining core feature components and eliminating redundant features, so as to simplify the feature matrix structure and improve the efficiency of subsequent correlation analysis. The dimensionality-reduced feature matrix is ​​transformed into an emotion feature vector, where each dimension of the emotion feature vector corresponds to a tone feature parameter, and its vector value reflects the quantification level of the corresponding feature. The emotional feature vector is input into a regression model trained on a sample library of emotional associations in elderly people's tone of voice. The output of the regression model is a calibration value for each dimension of the input vector. The original vector is updated using this calibration value. The corrected sentiment feature vector is output as input data for the subsequent construction of the semantic sentiment association matrix, and the feature extraction and quantization parameters are recorded for the optimization of the feature quantization algorithm. The emotional feature vectors are associated with the core semantic set and stored. The stored association data is used for parameter calibration when constructing the semantic-emotion association matrix.

8. The semantic understanding method for elderly companionship based on natural language processing according to claim 4, characterized in that, The process of comparing the semantic units in the alignment core set with the scene-specific semantic units in the high-frequency scene semantic library for the elderly one by one, calculating semantic similarity and recording the similarity value of each comparison result, and generating scene matching results by matching the corresponding scene categories based on the similarity values, includes: Semantic units are extracted from the alignment core set, categorized by nouns, verbs, and adjectives, and a semantic unit classification set is generated. The part of speech and core meaning of each semantic unit are then labeled. Retrieve scene-specific semantic units from the high-frequency scene semantic database for the elderly, split them according to scene category, so that each scene corresponds to a set of exclusive semantic units, forming a scene semantic classification index; Each semantic unit in the semantic unit classification set is compared one by one with the exclusive semantic units in the scene semantic classification index, the semantic similarity is calculated and the similarity value of each comparison result is recorded; Select semantic unit combinations with similarity values ​​reaching a set threshold, lock in the corresponding scene category, preliminarily determine the scene dimension to which each semantic unit belongs, and generate initial scene matching results; Conflict detection is performed on the initial scene matching results. If a single semantic unit is detected to match multiple scene categories, the most fitting scene category is determined through semantic context analysis to resolve the matching conflict. The initial scene matching results and conflict detection results are integrated to generate scene matching results. The correspondence between the alignment core set and the scene category is marked in the results. The scene category with the highest matching degree and the similarity value are recorded for each semantic unit. Extract the associated words between the semantic unit and the matched scene-specific semantic unit, and use them as a supplement to the matching basis to strengthen the matching logic. The vector representations of all semantic units in the alignment core set are weighted and averaged to obtain the overall semantic vector; the distance between the overall semantic vector and the center vector of each scene category is calculated, and the scene affiliation probability of each semantic unit in the scene matching result is reweighted and reassigned based on this distance; The optimized scene matching result is output as input data for subsequent adjustments to the sentiment feature vector representation based on the scene matching result, and the matching parameters are recorded for optimization of semantic matching rules. The scene matching results are established and stored in a correspondence with the alignment core set. The stored correspondence data is used as a benchmark for calibration when constructing the scene sentiment alignment matrix in the future.

9. The semantic understanding method for elderly companionship based on natural language processing according to claim 1, characterized in that, The steps include adding the supplementary semantic unit to the corresponding scene category of the high-frequency scene semantic library for the elderly, adjusting the correlation calculation parameters used to construct the emotional dependency link based on the supplementary semantic unit and its context in the feedback expression, and retraining the semantic role annotation enhancement algorithm using a new sample set containing the feedback expression. The system receives feedback from elderly users on semantic and emotional response text, converts the feedback into text data using a speech-to-text algorithm, and performs text noise reduction on the text data to preserve the core semantic content. Semantic role annotation enhancement processing is performed on the noise-reduced feedback text data to generate a feedback semantic role framework, and core semantic units are extracted from the feedback semantic role framework. The core semantic units in the feedback semantic role framework are transformed into supplementary semantic units, classified and organized according to semantic type, and a data content system that can be updated is built. The semantic library for high-frequency scenarios of the elderly is retrieved, and the supplementary semantic units are written into the storage path corresponding to the scenario category. At the same time, according to the preset timeliness rules, semantic units and expression paradigms whose most recent call time is earlier than the set timestamp are deleted from the semantic library for high-frequency scenarios of the elderly. Analyze the emotional association logic in the feedback semantic role framework, extract a new semantic emotional association pattern, and replace the corresponding content in the original emotional dependency link construction rule with the new semantic emotional association pattern. Using the supplementary semantic units and the new semantic sentiment association patterns as new training samples, incremental training is performed on the argument missing detection module and the role labeling module in the semantic role labeling enhancement algorithm to update their model parameters. Using a test set containing the feedback text data and its annotations, calculate the annotation precision, recall, and argument completion accuracy of the algorithm after iteration; if any of the indicators is lower than a preset threshold, adjust the hyperparameters of the incremental training and retrain and test.

10. A semantic understanding system for elderly care and companionship based on natural language processing, characterized in that, include: processor; A machine-readable storage medium for storing machine-executable instructions of the processor; The processor is configured to execute the semantic understanding method for elderly care based on natural language processing according to any one of claims 1 to 9 by executing the machine-executable instructions.