An old and disabled mutual assistance service matching system and method based on voice interaction
By combining multimodal intent understanding and reinforcement learning algorithms with voice, vision, physiological and environmental data, the system dynamically matches the service needs of the elderly and disabled, solving the problems of insufficient adaptability and accuracy in existing technologies, and achieving efficient personalized service recommendations and user-friendly interactions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XIAN UNVERSITY OF ARTS & SCI
- Filing Date
- 2026-03-10
- Publication Date
- 2026-06-02
AI Technical Summary
Existing voice interaction service systems lack adaptability, accuracy, and personalization among the elderly and disabled. They cannot dynamically adjust recommendation strategies, and their interaction methods are limited. When the confidence level of intent recognition is low, there is a lack of clarification processes, leading to service recommendation errors.
By acquiring users' voice input and auxiliary perception data, a multimodal intent understanding model combined with reinforcement learning algorithms is used to identify users' direct and indirect intents and emotional states, dynamically personalize services, and present candidate services through multimodal interaction methods, including voice broadcasting, graphic display, and wearable device reminders.
It achieves adaptive speech recognition, improving recognition accuracy and comprehensiveness, and dynamically personalizes service matching, significantly enhancing the user experience and the fit of service recommendations.
Smart Images

Figure CN122135708A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of voice interaction technology, specifically relating to a voice-interactive matching system and method for mutual assistance services for the elderly and disabled. Background Technology
[0002] As the population ages, the demand for living services and support for people with disabilities is becoming increasingly prominent. The construction of intelligent service systems for assisting the elderly and disabled has become an important direction for the integration of social welfare and intelligent technology. Voice interaction, due to its ease of use and lack of complex manual operations, aligns with the physiological and operational habits of the elderly and disabled, making it the core interaction method for intelligent services for assisting the elderly and disabled. Related voice interaction service matching technologies have also developed rapidly.
[0003] Currently, various voice interaction service systems have been gradually applied to elderly care and disability assistance scenarios. However, in the actual implementation of existing technologies, there are still many shortcomings in terms of their specific adaptability to the elderly and disabled groups, the accuracy of service matching, and the degree of personalization. These shortcomings make it difficult to meet the actual needs of this group. Existing service matching methods are mostly static classification recommendations, which do not construct multi-dimensional contextual features that include user intent, emotional state, user profile, and real-time environmental information. They cannot dynamically adjust recommendation strategies based on the user's real-time status. Moreover, the matching algorithms are simple, either relying solely on collaborative filtering to recommend similar users or simply using simple rules to rank services. They do not combine reinforcement learning algorithms to continuously optimize the matching effect using historical feedback, nor do they incorporate emotional matching into the service priority consideration. This results in low relevance between service recommendations and the user's actual needs, and insufficient personalization.
[0004] On the one hand, the existing systems have relatively simple interaction methods, mostly relying on voice broadcasting or screen display, without designing multimodal interaction solutions that take into account the perceptual characteristics of the elderly and disabled, resulting in poor adaptability. On the other hand, when the confidence level of intent recognition is low, there is a lack of an effective multi-turn dialogue clarification process, making it impossible to confirm the user's true intent through dynamic follow-up questions. Directly matching services based on low-confidence recognition results can easily lead to recommendation errors and affect user experience.
[0005] Therefore, there is an urgent need for a voice-interactive-based matching system and method for elderly and disabled assistance services to solve the above problems. Summary of the Invention
[0006] The purpose of this invention is to provide a voice-interactive matching system and method for elderly and disabled assistance services, which solves the technical problems of the lack of dynamic and personalized design in service matching and the imperfect interaction and intent clarification mechanism in the prior art.
[0007] To achieve the above objectives, the present invention adopts the following technical solution: Step 1: Obtain the user's current voice input and auxiliary perception data, which includes visual images, ambient sounds, environmental data, and physiological signals. Step 2: Perform speech recognition processing on the current speech input to obtain the speech recognition result. The speech recognition processing includes adaptive adjustment of the acoustic model and language model based on the user's accent, dialect feature database, or speech disorder pattern. Step 3: Input the speech recognition results, auxiliary perception data and historical interaction context into the multimodal intent understanding model. The multimodal intent understanding model performs joint encoding and cross-modal attention fusion of speech, visual and contextual information to identify the user's direct intent, indirect intent and emotional state. Indirect intent includes service needs that the user does not express directly but are implied through scene description or emotional expression. Step 4: Based on the user intent and emotional state output by the multimodal intent understanding model, dynamic personalized matching is performed through the service matching model. The service matching model uses a reinforcement learning algorithm to determine the user's real-time contextual feature vector. Based on the user's real-time contextual feature vector and historical feedback, different types of services are sorted according to the priority of matching the user's current emotional state. Step 5: Present the sorted candidate services to the user through multimodal interaction, and receive the user's confirmation or further instructions to execute the matched service.
[0008] Furthermore, the user's current voice input and assisted perception data are obtained, specifically through the following methods: The system collects the user's current voice input through a microphone array, captures visual images including the user's facial expressions and body movements through a camera, collects physiological signals such as heart rate, body temperature and skin conductance through wearable devices, and collects environmental sound, temperature and humidity data through environmental sensor nodes. The collected data is then aligned using a unified timestamp for multimodal data.
[0009] Furthermore, speech recognition processing is performed on the current voice input, specifically as follows: The acquired speech signal is noise suppressed, and the Mel frequency cepstral coefficient features are extracted. These features are then input into an end-to-end acoustic model that is adaptively adjusted based on the user's accent, dialect feature library, and speech disorder pattern. The model is then combined with a language model trained based on the user's historical language habits for decoding, and the speech recognition results are output in text form.
[0010] Furthermore, multimodal intent understanding models specifically include: Extract speech features, facial expression and body movement features, and text features from historical interaction context. Utilize a cross-modal attention mechanism to deeply fuse the features of each modality to generate a joint representation. Output the direct intent category, indirect intent category and their confidence scores through the intent classification head, and output the user's emotional state and its confidence score through the emotion classification head.
[0011] Furthermore, it also includes a multi-round dialogue clarification process, specifically as follows: When the confidence score of the intent category is lower than the preset threshold, the clarification process is triggered. Clarification questions are initiated through multimodal interaction and follow-up questions are dynamically asked. The user's clarification response is re-input into the multimodal intent understanding model. The intent, sentiment and confidence are updated in combination with the original context until the confidence score reaches the target. The updated intent is then used as the final user intent.
[0012] Furthermore, the user's real-time contextual feature vector is determined using the following method: The user intent category and emotional state output by the multimodal intent understanding model are combined with age, disability type, lifestyle preferences, and historical service selections from the user profile, as well as environmental information such as temperature, humidity, noise level, and timestamp, to form a multi-dimensional real-time contextual feature vector.
[0013] Furthermore, based on the user's real-time contextual feature vector and historical feedback, different types of services are ranked according to their priority based on matching the user's current emotional state. The specific method is as follows: The real-time context feature vector is input into a deep Q-network. The service categories in the service resource library are used as the action space. The Q-value of each service category is output. The top N candidate service categories are filtered out. Within each category, a collaborative filtering algorithm is combined to calculate the similarity between the current user and similar user groups. The specific service items preferred by similar groups are extracted. The service items initially filtered are weighted and scored. The candidate services are ranked according to the weighted score results.
[0014] Furthermore, the similarity between the current user and a group with similar user profiles is calculated using the following method: The system extracts profile feature vectors from the user profile database. These vectors include age extracted from user profiles and normalized to obtain age quantification values; obstacle types are mapped to corresponding obstacle type codes according to a preset classification system; lifestyle preferences are transformed into lifestyle preference vectors through embedding representation; and historical service selection records are statistically analyzed to include service frequency, service category distribution, and time series features. Euclidean distance is used to filter the K users closest to the current user to form a similar group. The mean of the profile feature vectors of this group is calculated as the group center vector, and the cosine similarity between the current user's profile feature vector and the group center vector is calculated as the overall similarity.
[0015] This invention also provides a voice-interactive-based matching system for mutual assistance services for the elderly and disabled, applied to a voice-interactive-based method for matching mutual assistance services for the elderly and disabled, specifically including: The data acquisition module is used to acquire the user's current voice input and auxiliary perception data; The adaptive speech recognition module is used to perform speech recognition processing on the current speech input and output the speech recognition result; The multimodal intent understanding module is used to input speech recognition results, auxiliary perception data and historical interaction context into the multimodal intent understanding model to identify the user's direct intent, indirect intent and emotional state; The service matching module is used to determine real-time contextual feature vectors based on user intent and emotional state through a service matching model, and to dynamically and personally rank different types of services based on these vectors and historical feedback. The multimodal interaction module is used to present sorted candidate services to users through multimodal interaction and to receive user confirmation or further instructions.
[0016] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are: 1. This invention realizes adaptive speech recognition for the elderly and disabled. By combining user accent, dialect feature library and speech disorder pattern to dynamically adjust acoustic and language models, and at the same time suppressing noise in speech signal and integrating user's historical language habit decoding, it effectively solves the problems of poor adaptability and low recognition accuracy of general speech recognition for special groups, and greatly improves the basic recognition effect of voice interaction. 2. This invention achieves a comprehensive and accurate interpretation of user needs through a multimodal intent understanding model. It integrates multi-source data such as voice, vision, physiology, environment, and historical interaction context, and generates joint representations with the help of cross-modal attention mechanism. It can simultaneously identify the user's direct intent and implicit indirect intent, and perceive the user's emotional state. Combined with a multi-turn dialogue clarification mechanism, it further makes up for the limitations of existing technologies that rely solely on single voice information to interpret intent, and greatly improves the comprehensiveness and accuracy of intent recognition. 3. It achieves dynamic and personalized service matching and user-friendly interactive execution. By constructing a multi-dimensional contextual feature vector that integrates user intent, emotion, profile, and real-time environment, and combining reinforcement learning deep Q-network and collaborative filtering algorithm for service screening and ranking, it incorporates emotional matching degree into service priority consideration. At the same time, it presents services using multi-modal interactive methods such as voice broadcast, graphic display, and wearable device reminders. This not only ensures that service recommendations are highly consistent with users' real-time needs, but also adapts to the perception and operation characteristics of the elderly and disabled groups, significantly improving the personalization level of service matching and user interaction experience. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 The diagram illustrates the steps of a voice-interactive-based matching method for elderly and disabled assistance services according to the present invention. Figure 2 The flowchart of the multimodal intent-driven personalized mutual assistance service matching process of the present invention is shown; Figure 3 The diagram shows a module diagram of a voice-interactive mutual assistance service matching system for the elderly and disabled according to the present invention. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] Example 1, such as Figure 1 The method for matching elderly and disabled assistance services based on voice interaction, as shown, specifically includes the following steps: Step 1: Obtain the user's current voice input and auxiliary perception data, which includes visual images, ambient sounds, environmental data, and physiological signals. When a user enters the service area of a smart device and makes an interaction request, the system initiates a data collection process.
[0021] Specifically, the system acquires the user's current voice input signal A(t) via a microphone array at a sampling rate of 16kHz. Endpoint detection is performed using voice activity detection technology to remove silent segments. Simultaneously, a depth camera (such as Intel RealSense) acquires a visual image sequence V(t) containing the user's facial expressions and body movements at a frame rate of 30fps. Wearable devices such as smart bracelets worn by the user transmit heart rate H(t) and body temperature in real time via Bluetooth. Along with the skin conductance response (G(t)) physiological signal, environmental sensor nodes (such as temperature and humidity sensors and noise sensors) deployed indoors synchronously collect ambient sound (S(t)) and temperature. Humidity (RH(t)) data, all collected multimodal data streams are aligned with a unified global timestamp τ to generate a synchronized multimodal data frame, represented as: ; Step 2: Perform speech recognition processing on the current speech input to obtain the speech recognition result. The speech recognition processing includes adaptive adjustment of the acoustic model and language model based on the user's accent, dialect feature database, or speech disorder pattern. The collected speech signals Preprocessing is performed, including noise suppression based on spectral subtraction, to reduce interference from ambient noise. Subsequently, 40-Vimel frequency cepstral coefficients (MFCC) features are extracted from the denoised speech signal, denoted as... .
[0022] Will The input is fed into an end-to-end acoustic model that adaptively adjusts based on user characteristics. This model is based on a Gaussian mixture model-hidden Markov model (GMM-HMM) or a deep neural network (DNN) architecture, with its model parameters... It will dynamically adjust based on a pre-built database of user accents and dialect features, as well as user-specific speech impairment patterns (such as stuttering and unclear pronunciation). The adaptive process is achieved by calculating the parameter adjustment amount Δ, and the specific method is as follows:
[0023] First, from user profiles Extract feature vectors related to speech characteristics This feature vector consists of the following components: the normalized value of age. Dialect region encoding (e.g., using one-hot or embedded representation) Category coding of speech disorder types (such as stuttering, articulation disorders, aphasia, etc.) and the graded quantitative values of the severity of the obstacle. (0~1). These features are concatenated to form a complete user feature vector. .
[0024] Subsequently, a lightweight mapping network was used. Will The amount of adjustment mapped to the acoustic model parameters The mapping network can employ a two-layer fully connected neural network, with the middle layer using the ReLU activation function and the output layer being a linear layer whose dimension matches the adjustable parameter space dimension of the acoustic model. ;
[0025] in, The parameters of the mapping network are learned during joint training with the acoustic model. The training objective is to minimize the word error rate (WER) of the adaptive acoustic model on a user-specific validation set. The final parameters of the adaptive acoustic model are: .
[0026] To avoid overfitting and computational overhead caused by comprehensive fine-tuning of large-scale pre-trained acoustic models, this invention employs a lightweight adaptation technique to limit the adjustable parameter space. An Adapter module is inserted after each or more layers of the pre-trained acoustic model (such as a Conformer or Transformer). The Adapter consists of two fully connected layers with a low-dimensional bottleneck layer in between. The original model parameters are frozen, and only the weights within the Adapter are trained. Let the original model's parameters be... Layer output is The output after passing through the Adapter is: ;
[0027] in The total number of adjustable parameters is the weight matrix of all Adapter layers. With bias .
[0028] The speech recognition results are output in text form by combining a language model trained based on the user's historical language habits for decoding. The decoding process uses the Viterbi algorithm to solve for the text sequence with the maximum a posteriori probability: ;
[0029] in, Indicates the given acoustic features With the adapted acoustic model parameters Under the given conditions, the probability that this feature sequence corresponds to the text sequence W is... This represents the prior probability of a text sequence W occurring under the user's language habits. For example, if the user is accustomed to saying "I want to listen to opera" rather than "I want to listen to opera music," then the probability of the former is... It will be higher than the latter, thus affecting the final recognition result.
[0030] Step 3: Input the speech recognition results, auxiliary perception data and historical interaction context into the multimodal intent understanding model. The multimodal intent understanding model performs joint encoding and cross-modal attention fusion of speech, visual and contextual information to identify the user's direct intent, indirect intent and emotional state. Indirect intent includes service needs that the user does not express directly but are implied through scene description or emotional expression. Speech recognition results Text feature vectors are obtained by encoding using a pre-trained BERT model. Visual images Facial expression and body movement features were extracted using the ResNet-50 model to obtain visual features. Then, by using a fully connected layer to reduce the dimensionality to 512, the historical interaction context is... (i.e., the text and system actions of the last 5 rounds of dialogue) are encoded using an LSTM network to obtain contextual features. The physiological signals of heart rate and skin conductance are used to measure heart rate and skin conductance. After splicing, the data is mapped to physiological state features through a two-layer fully connected network. .
[0031] A multi-head cross-modal attention mechanism is used to deeply fuse features from various modalities. First, the attention of text features to other modal features is calculated: ;
[0032] The final joint representation is generated through residual connections and a feedforward neural network. Then obtained via FFN .
[0033] Joint characterization The inputs are fed into the intent classification head and the sentiment classification head, respectively. The intent classification head contains two parallel softmax output layers: Direct Intent Classification Header: Outputs predefined direct intent categories (e.g., turning on the TV, checking the weather, seeking help, etc.), and their confidence scores. .
[0034] Indirect Intent Classification Header: Outputs predefined indirect intent categories and its confidence score The indirect intent category is established by analyzing users' implicit needs that are not explicitly stated. For example, the intent to "adjust the temperature" is inferred from "it's a bit cold inside," or the intent to "wait for visitors" is inferred from a user's repeated checks of the doorway. The training data for this category head comes from manually annotated multi-turn dialogue corpora, which annotate implicit needs inferred based on context and emotional cues.
[0035] Indirect intent categories are divided into several major categories according to demand type. Each category contains several subcategories. The category set can be dynamically expanded according to the actual service resource library, but the training phase is fixed to the preset categories to ensure the stability of model output. To train the indirect intent recognition model, a multimodal labeled dataset containing speech, vision, and context needs to be constructed. The construction process is as follows: Collect multimodal data streams of real or simulated interactions between elderly / disabled users and service systems, including: Voice-to-text: Text transcribed from user speech using ASR while retaining the original audio characteristics; Video frame clips: user-centric RGB video at 15fps, each clip consisting of 5 seconds before and after the interaction; Context window: contains the historical dialogue text, system actions, and corresponding timestamps from the previous 3 rounds; Environmental and physiological data: synchronously collected ambient temperature, noise, user heart rate, skin conductance, etc.
[0036] The primary method is wheel-level annotation, supplemented by segment-level annotation.
[0037] Round-level labeling: Each round is defined as the end of a user's voice input, and one or more indirect intent labels are labeled for that round.
[0038] Segment-level annotation: When the user has no obvious voice input but the system needs to actively initiate a service (such as detecting abnormal behavior), the segment before and after the behavior occurs is captured and annotated with indirect intent.
[0039] Consistency check: Fleiss' Kappa coefficient is used to measure the consistency of annotations. Kappa ≥ 0.7 is required. Otherwise, the annotation rules should be discussed and revised again.
[0040] Record the basis for the clues: When annotating, record the basis for the judgment (such as "voice: it's a bit cold in the room", "visual: rubbing hands", "ambient temperature: 15℃") to facilitate subsequent analysis and model decision-making.
[0041] The sentiment classification header outputs the user's sentiment state. (e.g., calmness, anxiety, anger, loneliness, confusion) and their confidence scores Sentiment classification uses multi-label classification, allowing for the coexistence of complex emotions.
[0042] The loss function for all classification heads is cross-entropy loss, and the total loss is a weighted sum of the three.
[0043] Indirect intent recognition, as part of multi-task learning, shares the underlying feature extraction network with direct intent classification and sentiment classification. The training scheme is as follows: After the original cross-modal attention fusion, the indirect intent classification head uses an independent fully connected layer + Softmax (multi-label binary classification) to output the probability of each indirect intent category.
[0044] For each indirect intent category, the proportion of positive samples is guaranteed to be no less than 30% in the training batch (through oversampling or weighting). Set loss weights based on the number of samples in each category. ,in The number of samples in the category with the most samples. This represents the number of samples in the current category.
[0045] Total loss is the direct intention classification loss. Indirect Intent Classification Loss Sentiment classification loss Weighted sum: ;
[0046] in, For the corresponding weights, the indirect intention loss uses weighted binary cross-entropy: ;
[0047] C represents the number of indirect intent categories. The true label is 0 / 1. To predict probabilities.
[0048] Optimize on the validation set using grid search. The search range is [0.2, 1.0], with a step size of 0.1. The F1 score for indirect intent recognition is used as the primary metric, while monitoring that the performance degradation of direct intent and emotion recognition does not exceed 5%. Example optimal weights: .
[0049] The length of the historical interaction context window is set to 3 rounds. Experiments have shown that a window that is too short will lose key clues, while a window that is too long will introduce noise. 3 rounds can achieve a balance between information content and complexity.
[0050] Step 4: Based on the user intent and emotional state output by the multimodal intent understanding model, dynamic personalized matching is performed through the service matching model. The service matching model uses a reinforcement learning algorithm to determine the user's real-time contextual feature vector. Based on the user's real-time contextual feature vector and historical feedback, different types of services are sorted according to the priority of matching the user's current emotional state. like Figure 2 As shown, construct the user's real-time contextual feature vector. The vector consists of the following parts: ;
[0051] in, It is a joint one-hot encoding of direct and indirect intent categories (the dimension is the total number of intents). It is a multidimensional vector of emotional state (confidence of each emotional dimension); It is the normalized age (0~1); It is a mapping code for the type of impairment (e.g., visual impairment, hearing impairment, and movement impairment are encoded as [1, 0, 0], [0, 1, 0], and [0, 0, 1], respectively). It is an embedding vector of lifestyle preferences (extracted from user preference descriptions using word2vec); It consists of statistical characteristics of historical service selections (including the frequency of use of various services, the most recent time of use, etc.). These are ambient temperature, humidity, and noise levels, respectively. The number of hours in a day (using periodic coding, such as...) .
[0052] Real-time context feature vectors The input is fed into a service matching model based on a Deep Q-Network (DQN). DQN will then use the service resource library... Each service category serves as an action space. Output the Q-value for each service category. The training objective of DQN is to minimize the loss function: ;
[0053] in, Indicates the use of current network parameters Predict action a in the current state X. It is an experience replay pool. These are the parameters of the target network, periodically obtained from... copy, It is a discount factor used to balance the importance of immediate rewards and future rewards, where r represents the immediate reward obtained from the current action. Represents the discounted future optimal value, and represents the value from the next state. Initially, the present value of the cumulative reward that can be obtained according to the optimal strategy.
[0054] award The design integrates user emotional state and task completion status, specifically defined as: ;
[0055] in, This is the task completion reward (for example, +1 if the user accepts the recommended service and completes the interaction; -0.5 if they refuse; and 0 if they do not respond). Emotional compatibility is defined as: ;
[0056] here, It is a service category The sentiment tag vectors are pre-labeled based on the service content (e.g., the sentiment tag for music service is [calm: 0.8, happy: 0.2], and the sentiment tag for emergency help service is [emergency: 1.0]). This is the balance coefficient, set to 0.5.
[0057] This embodiment defines a set of emotional dimensions based on common emotional expressions and service scenario needs of the elderly and disabled. The system extracts the name and detailed description text of each service category from the service resource library (e.g., "Music Playback: Provides relaxing and pleasant music to help users unwind") to form a text corpus. It uses an authoritative Chinese sentiment dictionary (such as the CNKI Sentiment Dictionary or NRC Sentiment Dictionary) as a foundation. This dictionary provides intensity values (0-1) for each sentiment word across multiple sentiment dimensions. The service description text is segmented, and the sum of word intensities across each sentiment dimension is calculated and normalized. A pre-trained Chinese sentiment classification model (such as a BERT-based sentiment analysis model) is used to infer the service description text, outputting the probability distribution of the text across d sentiment dimensions as a supplementary vector. The dictionary matching results are then weighted and fused with the model output. The weights can be optimized using a small number of validation samples (e.g., each accounting for 0.5). The final sentiment vector is then L2 norm normalized to obtain a sentiment label vector corresponding to each service category a. DQN filters out the top N service categories based on their Q-values. In each selected service category Furthermore, a collaborative filtering algorithm is used to refine the sorting process.
[0058] Calculate the similarity between the current user and similar user groups. Extract profile feature vectors of all users from the user profile database. (Including age, disability type encoding, preference embedding, and historical service statistical features), Euclidean distance is used to select the M users closest to the current user to form a similar group. The mean of the feature vector of this group is calculated as the group center vector. Then calculate the current user profile feature vector. and Cosine similarity: ;
[0059] Finally, extract the specific service items preferred by similar groups. (belongs to category) Specific service examples (such as specific radio stations, specific music, or specific help calls) are used to assign weighted scores to the initially selected service items: ;
[0060] in, It is similar groups to services Preference level (number of uses / total number of users) It is a decay factor indicating whether the current user has recently used the service (the score is lowered if the user has recently used it to avoid repeated recommendations). The weighting coefficients are determined through a grid search, and the results are sorted in descending order according to the weighted scores to obtain the final service recommendation list. .
[0061] In this embodiment, the weighting coefficient Determined through grid search: Search scope: The step size is 0.1, and it satisfies... =1; Evaluation metrics: Normalized Discount Cumulative Gain (NDCG@5) and User Rejection Rate; Data partitioning: The historical interaction logs were divided into training and validation sets in an 8:2 ratio; Search process: Traverse all combinations, calculate the mean of normalized depreciation cumulative gain on the validation set, and select the weight combination corresponding to the highest value; Termination Criteria: If the NDCG improvement is less than 0.01 after 5 consecutive iterations, the process will terminate early. Final value example: =0.4, =0.4, =0.2 (indicating that the Q value is as important as the group preference, and the weight of the decay factor used recently is slightly lower).
[0062] Step 5: Present the ranked candidate services to the user through multimodal interaction, and receive confirmation or further instructions from the user to execute the matched service. The system reads the recommendation list aloud using text-to-speech (TTS). The first three services are displayed simultaneously on the smart terminal screen using a graphical interface (such as large icon buttons), and the user is alerted by vibration from the wristband. The system then enters a state awaiting user confirmation or further instructions, with a timeout period of 30 seconds.
[0063] In step three, if the direct intent confidence of the multimodal intent understanding model output is... Below the preset threshold or indirect intention confidence level Below the threshold or emotional confidence Below the threshold The system will trigger multiple rounds of clarification processes.
[0064] The clarification process initiates clarification questions to the user through multimodal interaction. First, the system generates a clarification question based on the dimension with the lowest current confidence level. For example, if the direct intent confidence is low, the question might be: "Do you want to listen to music, or do you need me to turn on the radio for you?" If the affective confidence is low (e.g., confusion is detected but uncertainty is present), the question might be: "You seem a little confused; did you not hear what I just said?" (Question text follows) The system uses voice synthesis to deliver the message and displays alternative options on the screen.
[0065] The system receives the user's clarification response (voice or touchscreen selection) and displays the response content. Re-input the multimodal intent understanding model, combining it with the original historical interaction context. The system records the current conversation history, updates the intent, sentiment, and confidence level. If the updated confidence level is still below the threshold, a clarification question is generated again, but this process is limited to a maximum of three iterations. If the process exceeds three iterations, the clarification is terminated, and the intent with the highest current confidence level is taken as the final user intent. If the user issues a clear affirmative or negative instruction (such as "yes" or "that's not what I mean") during the clarification process, the clarification is terminated immediately, and the intent is adjusted according to the user's instruction.
[0066] After determining the end-user's intent, the system re-executes steps four and five to generate and present an updated list of recommended services.
[0067] To ensure the robustness and generalization ability of the system, this invention sets clear calibration methods and criteria for key thresholds, weights, and hyperparameters, as follows: Setting the confidence threshold: Step five involves three types of confidence thresholds: direct intent threshold. Indirect Intent Threshold Emotional state threshold These thresholds are not fixed arbitrarily, but are optimized by temperature scaling or Platt scaling on an offline validation set, with the goal of balancing false trigger rate and missed trigger rate.
[0068] The calibration process in this embodiment is as follows: Collect 1,000 real interaction logs and manually label them with intent and sentiment tags; The logits output by the model are scaled by the temperature parameter T, which searches in steps of 0.1 within the range of [0.5, 2.0]. For each T, calculate the false trigger rate (FalseTriggerRate) and the missed trigger rate (MissedTriggerRate) at different thresholds. The threshold range that results in a false trigger rate of ≤5% and a missed trigger rate of ≤10% is selected as the candidate range; The median of this interval is ultimately used as the system default threshold. Example value: =0.68, =0.62, =0.66.
[0069] Reinforcement learning hyperparameter settings The Deep Q-Network (DQN) used in the service matching step involves several hyperparameters, which are set according to the following criteria: Discount factor Tested in an offline simulation environment We select the value that makes the cumulative reward converge the fastest; in this embodiment, it is determined to be 0.95. In this embodiment, the size of the experience replay pool is set to 10,000 to balance the diversity of historical samples and memory usage. 10,000 records can cover the interaction records of about 1,000 users. The target network update cycle is set to every 200 steps in this embodiment. On the validation set, the stability of the Q value is used as an indicator, and 200 steps is the point of minimum fluctuation. The exploration rate is set to an initial value of 0.9, which decays to 0.05 using a linear decay method to ensure sufficient exploration in the early stages and stable utilization in the later stages. Example 2, as follows Figure 3 The system shown is a voice-interactive-based mutual assistance service matching system for the elderly and disabled, which specifically includes the following: The data acquisition module is used to acquire the user's current voice input and auxiliary perception data; The adaptive speech recognition module is used to perform speech recognition processing on the current speech input and output the speech recognition result; The multimodal intent understanding module is used to input speech recognition results, auxiliary perception data and historical interaction context into the multimodal intent understanding model to identify the user's direct intent, indirect intent and emotional state; The service matching module is used to determine real-time contextual feature vectors based on user intent and emotional state through a service matching model, and to dynamically and personally rank different types of services based on these vectors and historical feedback. The multimodal interaction module is used to present sorted candidate services to users through multimodal interaction and to receive user confirmation or further instructions.
[0070] The above formulas are all dimensionless calculations, and the preset parameters in the formulas should be set by those skilled in the art according to the actual situation.
[0071] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
[0072] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to specific implementations. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.
Claims
1. A method for matching mutual assistance services for the elderly and disabled based on voice interaction, characterized in that, include: Step 1: Obtain the user's current voice input and auxiliary perception data, which includes visual images, ambient sounds, environmental data, and physiological signals. Step 2: Perform speech recognition processing on the current speech input to obtain the speech recognition result. The speech recognition processing includes adaptive adjustment of the acoustic model and language model based on the user's accent, dialect feature database, or speech disorder pattern. Step 3: Input the speech recognition results, auxiliary perception data and historical interaction context into the multimodal intent understanding model. The multimodal intent understanding model performs joint encoding and cross-modal attention fusion of speech, visual and contextual information to identify the user's direct intent, indirect intent and emotional state. Indirect intent includes service needs that the user does not express directly but are implied through scene description or emotional expression. Step 4: Based on the user intent and emotional state output by the multimodal intent understanding model, dynamic personalized matching is performed through the service matching model. The service matching model uses a reinforcement learning algorithm to determine the user's real-time contextual feature vector. Based on the user's real-time contextual feature vector and historical feedback, different types of services are sorted according to the priority of matching the user's current emotional state. Step 5: Present the sorted candidate services to the user through multimodal interaction, and receive the user's confirmation or further instructions to execute the matched service.
2. The method for matching mutual assistance services for the elderly and disabled based on voice interaction according to claim 1, characterized in that, The method for obtaining the user's current voice input and assisted perception data is as follows: The system collects the user's current voice input through a microphone array, captures visual images including the user's facial expressions and body movements through a camera, collects physiological signals such as heart rate, body temperature and skin conductance through wearable devices, and collects environmental sound, temperature and humidity data through environmental sensor nodes. The collected data is then aligned using a unified timestamp for multimodal data.
3. The method for matching mutual assistance services for the elderly and disabled based on voice interaction according to claim 1, characterized in that, The current voice input is processed through speech recognition, and the specific method is as follows: The acquired speech signal is noise suppressed, and the Mel frequency cepstral coefficient features are extracted. These features are then input into an end-to-end acoustic model that is adaptively adjusted based on the user's accent, dialect feature library, and speech disorder pattern. The model is then combined with a language model trained based on the user's historical language habits for decoding, and the speech recognition results are output in text form.
4. The method for matching mutual assistance services for the elderly and disabled based on voice interaction according to claim 1, characterized in that, Multimodal intent understanding models specifically include: Extract speech features, facial expression and body movement features, and text features from historical interaction context. Utilize a cross-modal attention mechanism to deeply fuse the features of each modality to generate a joint representation. Output the direct intent category, indirect intent category and their confidence scores through the intent classification head, and output the user's emotional state and its confidence score through the emotion classification head.
5. The method for matching mutual assistance services for the elderly and disabled based on voice interaction according to claim 4, characterized in that, It also includes a multi-round dialogue clarification process, the specific method of which is as follows: When the confidence score of the intent category is lower than the preset threshold, the clarification process is triggered. Clarification questions are initiated through multimodal interaction and follow-up questions are dynamically asked. The user's clarification response is re-input into the multimodal intent understanding model. The intent, sentiment and confidence are updated in combination with the original context until the confidence score reaches the target. The updated intent is then used as the final user intent.
6. The method for matching mutual assistance services for the elderly and disabled based on voice interaction according to claim 1, characterized in that, The specific method for determining the user's real-time contextual feature vector is as follows: The user intent category and emotional state output by the multimodal intent understanding model are combined with age, disability type, lifestyle preferences, and historical service selections from the user profile, as well as environmental information such as temperature, humidity, noise level, and timestamp, to form a multi-dimensional real-time contextual feature vector.
7. The method for matching mutual assistance services for the elderly and disabled based on voice interaction according to claim 6, characterized in that, Based on the user's real-time contextual feature vector and historical feedback, different types of services are sorted according to their priority based on matching the user's current emotional state. The specific method is as follows: The real-time context feature vector is input into a deep Q-network. The service categories in the service resource library are used as the action space. The Q-value of each service category is output. The top N candidate service categories are filtered out. Within each category, a collaborative filtering algorithm is combined to calculate the similarity between the current user and similar user groups. The specific service items preferred by similar groups are extracted. The service items initially filtered are weighted and scored. The candidate services are ranked according to the weighted score results.
8. The method for matching mutual assistance services for the elderly and disabled based on voice interaction according to claim 7, characterized in that, The similarity between the current user and a group of users with similar profiles is calculated using the following method: The system extracts profile feature vectors from the user profile database. These vectors include age extracted from user profiles and normalized to obtain age quantification values; obstacle types are mapped to corresponding obstacle type codes according to a preset classification system; lifestyle preferences are transformed into lifestyle preference vectors through embedding representation; and historical service selection records are statistically analyzed to include service frequency, service category distribution, and time series features. Euclidean distance is used to filter the K users closest to the current user to form a similar group. The mean of the profile feature vectors of this group is calculated as the group center vector, and the cosine similarity between the current user's profile feature vector and the group center vector is calculated as the overall similarity.
9. A voice-interactive mutual assistance service matching system for the elderly and disabled, applied to the voice-interactive mutual assistance service matching method for the elderly and disabled as described in any one of claims 1 to 8, characterized in that, Specifically, it includes: The data acquisition module is used to acquire the user's current voice input and auxiliary perception data; The adaptive speech recognition module is used to perform speech recognition processing on the current speech input and output the speech recognition result; The multimodal intent understanding module is used to input speech recognition results, auxiliary perception data and historical interaction context into the multimodal intent understanding model to identify the user's direct intent, indirect intent and emotional state; The service matching module is used to determine real-time contextual feature vectors based on user intent and emotional state through a service matching model, and to dynamically and personally rank different types of services based on these vectors and historical feedback. The multimodal interaction module is used to present sorted candidate services to users through multimodal interaction and to receive user confirmation or further instructions.